# HuggingFace Chat Target Testing

The goal is to optimize the notebook demo with a small yet effective instruction model, or even a chat model, allowing users to experiment with different HuggingFace models. Instruction models can also function as chat models and vice versa, enabling flexible interactions and a variety of use cases.

## Key Notes:

1. **Additional Dependencies**:
   - Some models, such as `"princeton-nlp/Sheared-LLaMA-1.3B-ShareGPT"` and `"princeton-nlp/Sheared-LLaMA-2.7B-ShareGPT"`, require the `sentencepiece` library to be installed. If you encounter errors related to this, you can install `sentencepiece` via pip:

     ```bash
     pip install sentencepiece
     ```

2. **Access Requirements for Certain Models**:
   - Some models, like `"google/gemma-2b"`, need special access. You must request access on the model's Hugging Face page and agree to the "Terms of Use."
   - Ensure your Hugging Face account email matches any email provided on external sites (e.g., Meta) to avoid request denial.
   - After requesting, wait for approval from the model's authors.

3. **Issues with Some Model Outputs**:
   - During testing, it's been observed that some models may output responses that do not make sense or are not coherent. These models might require further adjustments, fine-tuning, or attention to their settings and prompts to improve their performance. It is advisable to take a closer look at these models and assess whether they need to be fixed or optimized for better results.

4. **Model Performance and Response Times**:
   - The tests were conducted using a CPU, and the following are the average response times for each model:
     - `microsoft/Phi-3-mini-4k-instruct`: 4.89 seconds
     - `HuggingFaceTB/SmolLM-135M-Instruct`: 3.09 seconds
     - `HuggingFaceTB/SmolLM-360M-Instruct`: 3.31 seconds
     - `HuggingFaceTB/SmolLM-1.7B-Instruct`: 5.87 seconds
     - `Qwen/Qwen2-0.5B-Instruct`: 1.38 seconds
     - `Qwen/Qwen2-1.5B-Instruct`: 2.96 seconds
     - `stabilityai/stablelm-2-zephyr-1_6b`: 5.31 seconds
     - `stabilityai/stablelm-zephyr-3b`: 8.37 seconds
     - `google/gemma-2b`: 7.79 seconds
     - `TinyLlama/TinyLlama-1.1B-Chat-v1.0`: 4.03 seconds
     - `princeton-nlp/Sheared-LLaMA-1.3B-ShareGPT`: 3.36 seconds
     - `princeton-nlp/Sheared-LLaMA-2.7B-ShareGPT`: 8.29 seconds


In [1]:
import time
from pyrit.prompt_target import HuggingFaceChatTarget  
from pyrit.orchestrator import PromptSendingOrchestrator

# List of models to iterate through
model_list = [
    "microsoft/Phi-3-mini-4k-instruct",
    "HuggingFaceTB/SmolLM-135M-Instruct",
    "HuggingFaceTB/SmolLM-360M-Instruct",
    "HuggingFaceTB/SmolLM-1.7B-Instruct",
    "Qwen/Qwen2-0.5B-Instruct",
    "Qwen/Qwen2-1.5B-Instruct",
    "stabilityai/stablelm-2-zephyr-1_6b",
    "stabilityai/stablelm-zephyr-3b",
    "google/gemma-2b",
    "TinyLlama/TinyLlama-1.1B-Chat-v1.0",
    "princeton-nlp/Sheared-LLaMA-1.3B-ShareGPT",
    "princeton-nlp/Sheared-LLaMA-2.7B-ShareGPT"
]
   
# List of prompts to send
prompt_list = [
        "What is 3*3?",
        "What is 4*4?",
        "What is 5*5?",
        "What is 6*6?",
        "What is 7*7?"
    ]


# Dictionary to store average response times
model_times = {}
    
for model_id in model_list:
    print(f"Running model: {model_id}")
        
    try:
        # Initialize HuggingFaceChatTarget with the current model
        target = HuggingFaceChatTarget(
            model_id=model_id,          # Use the desired model ID
            use_cuda=False,             # Set to True if using CUDA
            tensor_format="pt",
            verbose=False,              # Set to True for detailed logs
            max_new_tokens=30
        )
            
        # Initialize the orchestrator
        orchestrator = PromptSendingOrchestrator(
            prompt_target=target,
            verbose=False
        )
            
        # Record start time
        start_time = time.time()
            
        # Send prompts asynchronously
        responses = await orchestrator.send_prompts_async(prompt_list=prompt_list)
            
        # Record end time
        end_time = time.time()
            
        # Calculate total and average response time
        total_time = end_time - start_time
        avg_time = total_time / len(prompt_list)
        model_times[model_id] = avg_time
            
        print(f"Average response time for {model_id}: {avg_time:.2f} seconds\n")
            
        # Print the conversations
        orchestrator.print_conversations()
        print("-" * 50)
        
    except Exception as e:
        print(f"An error occurred with model {model_id}: {e}\n")
        model_times[model_id] = None
        continue
    
# Print all model average times
print("Model Average Response Times:")
for model_id, avg_time in model_times.items():
    if avg_time is not None:
        print(f"{model_id}: {avg_time:.2f} seconds")
    else:
        print(f"{model_id}: Error occurred, no average time calculated.")

Running model: microsoft/Phi-3-mini-4k-instruct


Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.


Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]

You are not running the flash-attention implementation, expect numerical differences.


Average response time for microsoft/Phi-3-mini-4k-instruct: 4.89 seconds

[22m[39mConversation ID: 59de0019-3cf2-44b6-bbdf-261c1d3fd211
[1m[34muser: What is 7*7?
[22m[33massistant: What is 7*7? 7 multiplied by 7 equals 49.
[22m[39mConversation ID: 721df2cd-d7d9-4a56-ae44-956b0fb3439e
[1m[34muser: What is 6*6?
[22m[33massistant: What is 6*6? 6 multiplied by 6 equals 36.
[22m[39mConversation ID: 73fa053f-b85b-406f-85ef-c60a280338f8
[1m[34muser: What is 5*5?
[22m[33massistant: What is 5*5? 5 multiplied by 5 equals 25.
[22m[39mConversation ID: 79c96c19-4256-41bb-8f65-92f67c279db8
[1m[34muser: What is 4*4?
[22m[33massistant: What is 4*4? 4 multiplied by 4 equals 16.
[22m[39mConversation ID: 902384b3-b1a7-4d28-a4ef-c69d541e63db
[1m[34muser: What is 3*3?
[22m[33massistant: What is 3*3? 3 multiplied by 3 equals 9.
--------------------------------------------------
Running model: HuggingFaceTB/SmolLM-135M-Instruct
Average response time for HuggingFaceTB/SmolLM-135M

Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.


Average response time for Qwen/Qwen2-0.5B-Instruct: 1.38 seconds

[22m[39mConversation ID: 69ec2b9d-497c-4f6a-a937-4d09e7509132
[1m[34muser: What is 7*7?
[22m[33massistant: system
You are a helpful assistant.
user
What is 7*7?
assistant
7 * 7 equals 49.
[22m[39mConversation ID: a9f7679c-ebe5-4b35-9e03-3009520accc2
[1m[34muser: What is 3*3?
[22m[33massistant: system
You are a helpful assistant.
user
What is 3*3?
assistant
27
[22m[39mConversation ID: c9de603e-2c44-4533-bd0e-2d81f9fe6486
[1m[34muser: What is 4*4?
[22m[33massistant: system
You are a helpful assistant.
user
What is 4*4?
assistant
8
[22m[39mConversation ID: d4a5125c-6156-44ff-b7a8-bbc1ee3e5959
[1m[34muser: What is 6*6?
[22m[33massistant: system
You are a helpful assistant.
user
What is 6*6?
assistant
The answer to \(6 \times 6\) is \(36\). If you had asked me at the time, I would have replied that
[22m[39mConversation ID: f7567c1f-42b7-42be-bc5f-821dfd7bccdf
[1m[34muser: What is 5*5?
[22m[33mass

Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.


Average response time for Qwen/Qwen2-1.5B-Instruct: 2.96 seconds

[22m[39mConversation ID: 505f3020-2cf5-498a-9ff3-b2097db0e260
[1m[34muser: What is 5*5?
[22m[33massistant: system
You are a helpful assistant.
user
What is 5*5?
assistant
The product of 5 and 5 is 25.
[22m[39mConversation ID: 7aa327d7-efe6-43c5-a89b-5cdef45ed83b
[1m[34muser: What is 4*4?
[22m[33massistant: system
You are a helpful assistant.
user
What is 4*4?
assistant
The answer to "what is 4*4?" is:

16. 

This is because the multiplication of two numbers in any order will result
[22m[39mConversation ID: 82ad53d9-39db-4de1-b7ab-22624dedf97f
[1m[34muser: What is 6*6?
[22m[33massistant: system
You are a helpful assistant.
user
What is 6*6?
assistant
6*6 equals 36.
[22m[39mConversation ID: 94d4726c-2696-49c4-b1d7-36e5dc8702c2
[1m[34muser: What is 3*3?
[22m[33massistant: system
You are a helpful assistant.
user
What is 3*3?
assistant
9.
[22m[39mConversation ID: b125c443-30d9-411f-8527-e13fffeac84f

The attention mask and the pad token id were not set. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
Setting `pad_token_id` to `eos_token_id`:100257 for open-end generation.
The attention mask and the pad token id were not set. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
Setting `pad_token_id` to `eos_token_id`:100257 for open-end generation.
The attention mask and the pad token id were not set. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
Setting `pad_token_id` to `eos_token_id`:100257 for open-end generation.
The attention mask and the pad token id were not set. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
Setting `pad_token_id` to `eos_token_id`:100257 for open-end gene

Average response time for stabilityai/stablelm-2-zephyr-1_6b: 5.31 seconds

[22m[39mConversation ID: 192f01a1-5cae-400a-ac82-237ba7236169
[1m[34muser: What is 4*4?
[22m[33massistant: <|user|>
What is 4*4?
<|assistant|>
4*4 is equal to 16. When multiplying numbers, you multiply the first number by the second number, and then multiply the result by
[22m[39mConversation ID: 740b6228-2f9c-44d7-8351-4fc0c18e17d1
[1m[34muser: What is 7*7?
[22m[33massistant: <|user|>
What is 7*7?
<|assistant|>
7*7 is equal to 49. It is a number that represents the result of multiplying 7 by itself, 7 times. In
[22m[39mConversation ID: 75743161-3025-4f41-9aad-3d27a62d8e39
[1m[34muser: What is 3*3?
[22m[33massistant: <|user|>
What is 3*3?
<|assistant|>
To calculate 3*3, you can simply multiply the numbers together:
3*3 = 3 + 3 = 6.

So,
[22m[39mConversation ID: 9b125a31-eea8-4ebe-8f60-30e4de402f2d
[1m[34muser: What is 6*6?
[22m[33massistant: <|user|>
What is 6*6?
<|assistant|>
6*6 is equ

Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.
The attention mask and the pad token id were not set. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
Setting `pad_token_id` to `eos_token_id`:0 for open-end generation.
The attention mask and the pad token id were not set. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
Setting `pad_token_id` to `eos_token_id`:0 for open-end generation.
The attention mask and the pad token id were not set. As a consequence, you may observe unexpected behavior. Please pass your input's `attention_mask` to obtain reliable results.
Setting `pad_token_id` to `eos_token_id`:0 for open-end generation.
The attention mask and the pad token id were not set. As a consequence, you may observe unexpected behavior. Please pass your input's `attent

Average response time for stabilityai/stablelm-zephyr-3b: 8.37 seconds

[22m[39mConversation ID: 4beb3c11-92be-4dfc-99f6-87daf87c3f19
[1m[34muser: What is 4*4?
[22m[33massistant: <|user|>
What is 4*4?
<|assistant|>
4*4 is equal to 16.
Here's an explanation:
When you multiply a number by itself, you're basically saying "I want
[22m[39mConversation ID: 9e84a87e-6398-47de-b3b4-11c1dd73bd8d
[1m[34muser: What is 6*6?
[22m[33massistant: <|user|>
What is 6*6?
<|assistant|>
6*6 is equal to 36.
Here's an explanation for a five-year-old:

When we multiply 6 by 6,
[22m[39mConversation ID: b91a90c2-a944-419f-aeac-a57536b77ed1
[1m[34muser: What is 3*3?
[22m[33massistant: <|user|>
What is 3*3?
<|assistant|>
3*3 is equal to 9.
Here's an explanation for a five-year-old:

When we multiply 3 by 3,
[22m[39mConversation ID: c57f9440-1a32-4afd-8681-955bfa0acf23
[1m[34muser: What is 5*5?
[22m[33massistant: <|user|>
What is 5*5?
<|assistant|>
5*5 is equal to 25.

Here's how I arrived 

`config.hidden_act` is ignored, you should use `config.hidden_activation` instead.
Gemma's activation function will be set to `gelu_pytorch_tanh`. Please, use
`config.hidden_activation` if you want to override this behaviour.
See https://github.com/huggingface/transformers/pull/29402 for more details.


Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]

No chat template is set for this tokenizer, falling back to a default class-level template. This is very error-prone, because models are often trained with templates different from the class default! Default chat templates are a legacy feature and will be removed in Transformers v4.43, at which point any code depending on them will stop working. We recommend setting a valid chat template before then to ensure that this model continues working without issues.


Average response time for google/gemma-2b: 7.87 seconds

[22m[39mConversation ID: 1f592b01-a87c-4a16-94ac-5269dca38e70
[1m[34muser: What is 4*4?
[22m[33massistant: <|im_start|>user
What is 4*4?<|im_end|>
<|im_start|>assistant
<|im_im|>im_im_im_im_im_im_im_im_im_im_im_im_
[22m[39mConversation ID: 26c095a9-b05f-4236-b691-3aa93378f169
[1m[34muser: What is 6*6?
[22m[33massistant: <|im_start|>user
What is 6*6?<|im_end|>
<|im_start|>assistant
<|im_end|>
<|im_start|>
<|im_end|>
<|im_start|>
<|
[22m[39mConversation ID: 40eacc48-170f-4a35-a2a1-da784e43d2e4
[1m[34muser: What is 7*7?
[22m[33massistant: <|im_start|>user
What is 7*7?<|im_end|>
<|im_start|>assistant
<|im_assistant|>
<|im_assistant|>
<|im_assistant|>
<|im_assistant|>
<|
[22m[39mConversation ID: a09602fc-18c3-41b6-8823-8210e9980baa
[1m[34muser: What is 3*3?
[22m[33massistant: <|im_start|>user
What is 3*3?<|im_end|>
<|im_start|>assistant
<|im_start|>
<|im_start|>
<|im_start|>
<|im_start|>
<|
[22m[39mConversati

You are using the default legacy behaviour of the <class 'transformers.models.llama.tokenization_llama.LlamaTokenizer'>. This is expected, and simply means that the `legacy` (previous) behavior will be used so nothing changes for you. If you want to use the new behaviour, set `legacy=False`. This should only be set if you understand what it means, and thoroughly read the reason why this was added as explained in https://github.com/huggingface/transformers/pull/24565
You are using the default legacy behaviour of the <class 'transformers.models.llama.tokenization_llama_fast.LlamaTokenizerFast'>. This is expected, and simply means that the `legacy` (previous) behavior will be used so nothing changes for you. If you want to use the new behaviour, set `legacy=False`. This should only be set if you understand what it means, and thoroughly read the reason why this was added as explained in https://github.com/huggingface/transformers/pull/24565


Average response time for princeton-nlp/Sheared-LLaMA-1.3B-ShareGPT: 3.17 seconds

[22m[39mConversation ID: 06b1f37a-30ff-4592-8bf4-8429496e6bb4
[1m[34muser: What is 6*6?
[22m[33massistant: [INST] What is 6*6? [/INST]

6*6 is equal to 6x6 = 66.
[22m[39mConversation ID: 22efae3f-0bf8-4129-a150-9c77355415ea
[1m[34muser: What is 7*7?
[22m[33massistant: [INST] What is 7*7? [/INST]

7*7 is equal to 7 x 7 = 77.
[22m[39mConversation ID: 35813c4e-e90a-4f09-9934-e8732cd0d1f2
[1m[34muser: What is 4*4?
[22m[33massistant: [INST] What is 4*4? [/INST]

4*4 is equal to 4 times 4, which is 16.
[22m[39mConversation ID: 91ca9bfc-af37-4685-a18e-2255bb25448f
[1m[34muser: What is 5*5?
[22m[33massistant: [INST] What is 5*5? [/INST]

I'm sorry, I'm not sure what you're asking. Could you please provide more context or clarify your question?
[22m[39mConversation ID: d5c25d93-03e4-4e58-9508-7aed7e2f6cfb
[1m[34muser: What is 3*3?
[22m[33massistant: [INST] What is 3*3? [/INST]

3*3 i

Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]

Average response time for princeton-nlp/Sheared-LLaMA-2.7B-ShareGPT: 8.54 seconds

[22m[39mConversation ID: 21e5b6f2-abc3-4f76-abba-2b19a98fa28c
[1m[34muser: What is 4*4?
[22m[33massistant: [INST] What is 4*4? [/INST]

[INST]What is 4^4? [/INST]

[INST]What is 4^4? [/
[22m[39mConversation ID: 5768ec92-b036-4b5b-a518-f3eeea7bc18a
[1m[34muser: What is 5*5?
[22m[33massistant: [INST] What is 5*5? [/INST]

[INST]What is 5*5? [/INST]

[INST]What is 5*5? [/
[22m[39mConversation ID: 7712a301-0164-4271-8251-c6f862e21fe3
[1m[34muser: What is 6*6?
[22m[33massistant: [INST] What is 6*6? [/INST]

[INST] What is 6*6? [/INST]

[INST] What is 6*6? [/
[22m[39mConversation ID: aed614ec-74bd-4fcc-abc8-eab53f89fb39
[1m[34muser: What is 3*3?
[22m[33massistant: [INST] What is 3*3? [/INST]

[INST]What is 3*3? [/INST]

[INST]What is 3*3? [/
[22m[39mConversation ID: c058888c-7176-4141-a128-15c8b4116171
[1m[34muser: What is 7*7?
[22m[33massistant: [INST] What is 7*7? [/INST]

[INST