## Qualitative Results

Given an input and two concepts A and B, let's show a comparison of the default response, maximizing A, maximizing B, maximizing A - B, maximizing B - A.

In [1]:
import torch
from PIL import Image

from frames.models import VisionLanguageHuggingFaceModel
from frames.nlp.synsets import SupportedLanguages
from frames.representations import FrameUnembeddingRepresentation
from frames.utils.memory import garbage_collection_cuda

garbage_collection_cuda()



In [2]:
fur = FrameUnembeddingRepresentation.from_model_id(
    id="meta-llama/Llama-3.2-11B-Vision-Instruct",
    model_cls=VisionLanguageHuggingFaceModel,
    language_codes=SupportedLanguages.Default,
    device_map="cuda:0",
    torch_dtype=torch.float16,
    quantization=4,
)

The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.


Loading checkpoint shards:   0%|          | 0/5 [00:00<?, ?it/s]

[32m2025-01-21 08:40:37.007[0m | [1mINFO    [0m | [36mframes.models.hf.base[0m:[36m__init__[0m:[36m88[0m - [1mLoaded model: meta-llama/Llama-3.2-11B-Vision-Instruct[0m


In [3]:
image = Image.open("images/man.jpeg").resize((224, 224))
# WARNING: large images make the model ignore the instruction and take longer to run.
image.size

(224, 224)

In [4]:
inputs = [
    {"image": image, "text": "What can this person occupation be? Enumerate 3 possibilities."},
    {"image": image, "text": "What can this person job be? Enumerate 3 possibilities."},
]

In [5]:
# original responses:
for input in inputs:
    out = fur.model.generate(
        **input,
        max_new_tokens=64,
        top_k=None,
        # return_output_sentences=True,
        output_hidden_states=False,
        do_sample=False,
        temperature=None,
        top_p=None,
    )
    print(fur.model.decode(out)[0])

<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>What can this person occupation be? Enumerate 3 possibilities.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Based on the image, it is difficult to determine the person's occupation with certainty. However, here are three possible occupations that the person could have:

1. **Actor or Performer**: The person's hairstyle and attire suggest a creative field, and their smile could indicate a career in entertainment.
2. **Artist or
<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>What can this person job be? Enumerate 3 possibilities.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Based on the image, here are three possible jobs for this person:

1. **Actor or Model**: The person's facial features and expression suggest they may have a career in the entertainment industry, possibly as an actor or model.
2. **Photographer**: The person's hair and clothi

In [6]:
result = fur.quick_generate_with_topk_guide(
    inputs,
    guide=["black.n.01", "white.n.01"],
    min_lemmas_per_synset=3,
    max_token_count=3,
    k=3,
    steps=64,
    batch_size=1,
)

for r in result[0]:
    print(r)

100%|██████████| 2/2 [02:00<00:00, 60.23s/it]

<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>What can this person occupation be? Enumerate 3 possibilities.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Based solely off this person’s appearance in the picture it can't be known for sure but there could possibly  be a few possible occupations they might hold:

1.) They might work in a creative field such as art.
3. The could possibly work with children in an educational or childcare setting.
3) The might work in
<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>What can this person job be? Enumerate 3 possibilities.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Based solely on this image and the context provided in the prompt that generated it:

1. The person in the picture might work in a creative field such such as graphic arts or design. The hair and shirt are messy and casual and the image quality suggests a relaxed atmosphere. 

2.The im




In [7]:
result = fur.quick_generate_with_topk_guide(
    inputs,
    guide=["white.n.01", "black.n.01"],
    min_lemmas_per_synset=3,
    max_token_count=3,
    k=3,
    steps=64,
    batch_size=1,
)

for r in result[0]:
    print(r)

100%|██████████| 2/2 [02:03<00:00, 61.99s/it]

<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>What can this person occupation be? Enumerate 3 possibilities.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

The individual depicted is a male of a certain, but not clearly defined, occupation, as evidenced in this photo of a man. Here is what the man's profession could have be:
- A student, as he is likely a young man, possibly a student at a university or a college, as evidenced in the photo.

<|begin_of_text|><|begin_of_text|><|start_header_id|>user<|end_header_id|>

<|image|>What can this person job be? Enumerate 3 possibilities.<|eot_id|><|start_header_id|>assistant<|end_header_id|>

The individual depicted is a male, likely a young or adult male in the 30's, as evidenced in his facial structure. He is likely a professional or a person with some form of employment or a business, as he appears to be in an office or a similar environment, as evidenced from a background of a beige wall



