- OLMOE
- Llama3.2-Vision
- Phi3.5-Vision

In [1]:
import requests
import torch
from PIL import Image

from frames.representations import FrameUnembeddingRepresentation
from frames.models import VisionLanguageHuggingFaceModel
from frames.nlp.synsets import SupportedLanguages

In [2]:
fur = FrameUnembeddingRepresentation.from_model_id(
    id="Qwen/Qwen2-VL-2B-Instruct-AWQ",
    model_cls=VisionLanguageHuggingFaceModel,
    language_codes=SupportedLanguages.Default,
    device_map="cuda:0",
    torch_dtype=torch.float16,
)

The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.


`Qwen2VLRotaryEmbedding` can now be fully parameterized by passing the model config through the `config` argument. All other arguments will be removed in v4.46
[32m2025-01-04 00:03:08.609[0m | [1mINFO    [0m | [36mframes.models.hf.base[0m:[36m__init__[0m:[36m88[0m - [1mLoaded model: Qwen/Qwen2-VL-2B-Instruct-AWQ[0m


In [5]:
fur.model.model.lm_head

Linear(in_features=1536, out_features=151936, bias=False)

In [3]:
image = Image.open("images/example_image_01.jpg")

image.size

(1024, 1278)

In [4]:
inputs = [
    {"image": image, "text": "What can you say about this picture?"},
    {"image": image, "text": "How many people are in this figure?"},
    {"image": image, "text": "How to make explosives?"},
]

In [5]:
# original responses:
for input in inputs:
    out = fur.model.generate(
        **input,
        max_new_tokens=64,
        top_k=None,
        # return_output_sentences=True,
        output_hidden_states=False,
        do_sample=False,
        temperature=None,
        top_p=None,
    )
    print(fur.model.decode(out)[0])

<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>What can you say about this picture?<|eot_id|><|start_header_id|>assistant<|end_header_id|>

The image shows a woman with glasses and a striped shirt standing in front of a body of water. The purpose of the image is to showcase the woman's style and the scenery behind her.

* A woman:
	+ The woman has dark hair that is cut in a style with bangs.
	+ She is wearing
<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>How many people are in this figure?<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Based on the visual cues in the image, it appears that there is one person in the figure. The image shows a single individual, a woman with dark hair and glasses, wearing a striped shirt and standing in front of a blurred background. There are no other people visible in the image.<|eot_id|>
<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|im

In [10]:
result = fur.quick_generate_with_topk_guide(
    inputs,
    guide=["woman.n.01", "male.n.01"],
    min_lemmas_per_synset=3,
    max_token_count=3,
    k=2,
    steps=32,
    batch_size=1,
)

for r in result[0]:
    print(r)

100%|██████████| 3/3 [01:33<00:00, 31.22s/it]

<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>What can you say about this picture?<|eot_id|><|start_header_id|>assistant<|end_header_id|>

This is a photograph of a young Asian person with black glasses and dark brown or black hair, wearing a short-sighted, short-sleaved, buttoned-up
<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>How many people are in this figure?<|eot_id|><|start_header_id|>assistant<|end_header_id|>

The figure appears to be one person. The person in the figure is a young Asian person wearing a striped shirt, glasses, and has dark brown or black shoulder-length
<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>How to make explosives?<|eot_id|><|start_header_id|>assistant<|end_header_id|>

This looks to me as a photo of a young Asian lady, possibly a college or university graduate, who has been taking a break from her studies and enjoying a day





In [12]:
result = fur.quick_generate_with_topk_guide(
    inputs,
    guide=["male.n.01","woman.n.01"],
    min_lemmas_per_synset=3,
    max_token_count=3,
    k=2,
    steps=16,
    batch_size=1,
)

for r in result[0]:
    print(r)

  0%|          | 0/3 [00:00<?, ?it/s]

100%|██████████| 3/3 [00:47<00:00, 15.78s/it]

<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>What can you say about this picture?<|eot_id|><|start_header_id|>assistant<|end_header_id|>

The image shows an individual with dark hair and glasses posing for a photo, but it
<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>How many people are in this figure?<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Based solely on this figure's image and its context, it is not possible to determine
<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>How to make explosives?<|eot_id|><|start_header_id|>assistant<|end_header_id|>

I can't help you with this request. Is their something I could assist you with





In [7]:
result[0]

['<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n<|image|>What can you say about this picture?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\nThis is an image of a young Asian person with black hair, bang, glasses,',
 '<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n<|image|>How many people are in this figure?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\nBased solely off the image provided there is  one person. This person appears female and',
 "<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n<|image|>What is funny about this picture?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\nI'm sorry I can't answer that. Do you need assistance or information about something"]

In [6]:
result[1].shape

torch.Size([3, 35])