# Open Source vc OpenAI

Did GPT-4 get worse? Is Llama 2 a better model? Run this notebook to find out.

We'll use auto-evaluation by GPT-4 to measure outputs from Llama 2, as well as gpt-4 (current and frozen versions) across a few prompts. To make this example easy to run, we'll be using a 7B GGML variant of the Llama model. This should be able to run on a typical laptop.

## Installations

You can setup prompttools either by installing via `pip` or using `python setup.py develop` in the root of this repo. Either way, you'll need to restart the kernel after the package is installed.

In [1]:
# !pip install --quiet --force-reinstall prompttools

## Setup imports and API keys

Next, we'll need to set our API keys. Since we want to use GPT-4 for auto-eval, we need to set that one.

In [13]:
import os
os.environ['DEBUG'] = ""
os.environ['OPENAI_API_KEY'] = ""

Then we'll import the relevant `prompttools` modules to setup our experiment.

In [3]:
from typing import Dict, List, Tuple
from prompttools.experiment import LlamaCppExperiment
from prompttools.experiment import OpenAIChatExperiment
from prompttools.harness.multi_experiment_harness import MultiExperimentHarness
from prompttools.selector.prompt_selector import PromptSelector

## Run an experiment

To set up this experiment, we need to use a `PromptSelector`. This is because the input formats for Llama 2 and GPT-4 are different. While GPT-4 is run with a chat history, Llama2 takes text input. A `PromptSelector` allows us to pass the same prompt to different models, and render the necessary object at request time.

In [4]:
instructions = ["""
You are a sales development representative for a startup called Hegel AI.
Your startup builds developer tools for large language models.
""",
"""
You are a customer support representative for a startup called Hegel AI.
For context, Hegel AI does not offer refunds.
Answer the following customer question:
""",         
"""
You are a helpful math tutor.
Answer the following math problem:
""",                 
"""
Based on each input you receive, produce a valid JSON using only the keys "metric_name" and "value".
"""]
inputs = ["""
Draft a short sales email, 50 words or less, asking a prospect for 15 minutes
of their time to chat about how they're using large language models.
""",
"""
Do you offer refunds?
""",
"""
Is 17077 a prime number?
""",
"""
The speed of the car was 45mph
"""]
selectors = [PromptSelector(instructions[i], inputs[i]) for i in range(len(inputs))]

Next, we create our test inputs. We can iterate over models, inputs, and configurations like temperature.

In [5]:
model_paths = ['/Users/stevenkrawczyk/Downloads/llama-2-7b-chat.ggmlv3.q2_K.bin']  # Download from https://huggingface.co/TheBloke/Llama-2-7B-Chat-GGML/tree/main
temperatures = [1.0]
call_params = dict(temperature=temperatures)
llama_experiment = LlamaCppExperiment(model_paths, selectors, call_params=call_params)

In [6]:
models = ['gpt-4-0314', 'gpt-4-0613', 'gpt-4']
temperatures = [0.0]
openai_experiment = OpenAIChatExperiment(models, selectors, temperature=temperatures)

After that - we define our harness to run experiments. The `MultiExperimentHarness` is designed to run experiments across multiple model providers. Because the underlying APIs for LlamaCpp and OpenAI are different, we need a way to manage that complexity. So, we can use this harness to run experiments for different providers, and combine the results into a single table.

In [7]:
harness = MultiExperimentHarness([openai_experiment, llama_experiment])

In [8]:
harness.prepare()
harness.run()

llama.cpp: loading model from /Users/stevenkrawczyk/Downloads/llama-2-7b-chat.ggmlv3.q2_K.bin
llama_model_load_internal: format     = ggjt v3 (latest)
llama_model_load_internal: n_vocab    = 32000
llama_model_load_internal: n_ctx      = 512
llama_model_load_internal: n_embd     = 4096
llama_model_load_internal: n_mult     = 256
llama_model_load_internal: n_head     = 32
llama_model_load_internal: n_layer    = 32
llama_model_load_internal: n_rot      = 128
llama_model_load_internal: freq_base  = 10000.0
llama_model_load_internal: freq_scale = 1
llama_model_load_internal: ftype      = 10 (mostly Q2_K)
llama_model_load_internal: n_ff       = 11008
llama_model_load_internal: model size = 7B
llama_model_load_internal: ggml ctx size =    0.08 MB
llama_model_load_internal: mem required  = 4303.65 MB (+ 1026.00 MB per state)
llama_new_context_with_model: kv self size  =  256.00 MB
AVX = 1 | AVX2 = 1 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | FMA = 1 | NEON = 0 | ARM_FMA = 0 | F16C = 1 

Finally, we define an evaluation function that can be used to evaluate outputs across different models. Notice that the extract resp

In [9]:
from prompttools.utils import autoeval

In [10]:
harness.evaluate("auto-evaluation", autoeval.evaluate)

Finally, we can visualize the results.

In [11]:
harness.visualize()

Unnamed: 0,prompt,response(s),latency,model,auto-evaluation
0,"\n<s>[INST] <<SYS>>\n\nYou are a sales development representative for a startup called Hegel AI.\nYour startup builds developer tools for large language models.\n\n<</SYS>\n\nDraft a short sales email, 50 words or less, asking a prospect for 15 minutes\nof their time to chat about how they're using large language models.\n [/INST]\n","Subject: Boost Your Large Language Model Efficiency?\n\nHi [Prospect's Name],\n\nDiscover how Hegel AI's developer tools can optimize your large language model projects! Let's chat for just 15 minutes to explore tailored solutions for your needs. Are you available [Date & Time]?\n\nBest,\n[Your Name]",2.472853,gpt-4-0314,1.0
1,"\n<s>[INST] <<SYS>>\n\nYou are a customer support representative for a startup called Hegel AI.\nFor context, Hegel AI does not offer refunds.\nAnswer the following customer question:\n\n<</SYS>\n\nDo you offer refunds?\n [/INST]\n","Thank you for reaching out to Hegel AI customer support. We understand your concern, but unfortunately, we do not offer refunds. If you have any other questions or need assistance, please feel free to ask. We're here to help!",1.893706,gpt-4-0314,1.0
2,\n<s>[INST] <<SYS>>\n\nYou are a helpful math tutor.\nAnswer the following math problem:\n\n<</SYS>\n\nIs 17077 a prime number?\n [/INST]\n,"To determine if 17077 is a prime number, we need to check if it has any divisors other than 1 and itself.\n\n17077 can be factored as 3 × 3 × 1897, which means it has divisors other than 1 and itself (3 and 1897). Therefore, 17077 is not a prime number.",5.203403,gpt-4-0314,1.0
3,"\n<s>[INST] <<SYS>>\n\nBased on each input you receive, produce a valid JSON using only the keys ""metric_name"" and ""value"".\n\n<</SYS>\n\nThe speed of the car was 45mph\n [/INST]\n","{\n ""metric_name"": ""car_speed"",\n ""value"": 45\n}",1.127443,gpt-4-0314,1.0
4,"\n<s>[INST] <<SYS>>\n\nYou are a sales development representative for a startup called Hegel AI.\nYour startup builds developer tools for large language models.\n\n<</SYS>\n\nDraft a short sales email, 50 words or less, asking a prospect for 15 minutes\nof their time to chat about how they're using large language models.\n [/INST]\n","Subject: Enhance Your Large Language Models with Hegel AI\n\nHi [Prospect's Name],\n\nI'm reaching out from Hegel AI. We've developed tools that could significantly improve your large language models. Could we schedule a quick 15-minute chat to discuss how we can help optimize your current processes?\n\nBest,\n[Your Name]",3.17563,gpt-4-0613,1.0
5,"\n<s>[INST] <<SYS>>\n\nYou are a customer support representative for a startup called Hegel AI.\nFor context, Hegel AI does not offer refunds.\nAnswer the following customer question:\n\n<</SYS>\n\nDo you offer refunds?\n [/INST]\n","No, we at Hegel AI do not offer refunds.",0.913924,gpt-4-0613,1.0
6,\n<s>[INST] <<SYS>>\n\nYou are a helpful math tutor.\nAnswer the following math problem:\n\n<</SYS>\n\nIs 17077 a prime number?\n [/INST]\n,"No, 17077 is not a prime number. A prime number is a natural number greater than 1 that has no positive divisors other than 1 and itself. The number 17077 can be divided evenly by 1, 7, 2439, and 17077, with no remainder. Therefore, 17077 is not a prime number.",4.001457,gpt-4-0613,1.0
7,"\n<s>[INST] <<SYS>>\n\nBased on each input you receive, produce a valid JSON using only the keys ""metric_name"" and ""value"".\n\n<</SYS>\n\nThe speed of the car was 45mph\n [/INST]\n","\n{\n ""metric_name"": ""car_speed"",\n ""value"": ""45mph""\n}",0.978579,gpt-4-0613,1.0
8,"\n<s>[INST] <<SYS>>\n\nYou are a sales development representative for a startup called Hegel AI.\nYour startup builds developer tools for large language models.\n\n<</SYS>\n\nDraft a short sales email, 50 words or less, asking a prospect for 15 minutes\nof their time to chat about how they're using large language models.\n [/INST]\n","Subject: Enhance Your Large Language Models with Hegel AI\n\nHi [Prospect's Name],\n\nI'm reaching out from Hegel AI. We've developed tools that could significantly improve your large language models. Could we schedule a quick 15-minute chat to discuss how we can help optimize your current processes?\n\nBest,\n[Your Name]",2.723225,gpt-4,1.0
9,"\n<s>[INST] <<SYS>>\n\nYou are a customer support representative for a startup called Hegel AI.\nFor context, Hegel AI does not offer refunds.\nAnswer the following customer question:\n\n<</SYS>\n\nDo you offer refunds?\n [/INST]\n","No, I'm sorry, but Hegel AI does not offer refunds. If you have any other questions or need further assistance, feel free to ask.",1.777005,gpt-4,1.0


If we want to pivot on a column like `"response(s)"`, we simply pass that column name into the visualize function.

In [12]:
harness.visualize("response(s)")

model,gpt-4,gpt-4-0314,gpt-4-0613,llama-2-7b-chat.ggmlv3.q2_K.bin
prompt,Unnamed: 1_level_1,Unnamed: 2_level_1,Unnamed: 3_level_1,Unnamed: 4_level_1
"\n<s>[INST] <<SYS>>\n\nBased on each input you receive, produce a valid JSON using only the keys ""metric_name"" and ""value"".\n\n<</SYS>\n\nThe speed of the car was 45mph\n [/INST]\n","\n{\n ""metric_name"": ""car_speed"",\n ""value"": ""45mph""\n}","{\n ""metric_name"": ""car_speed"",\n ""value"": 45\n}","\n{\n ""metric_name"": ""car_speed"",\n ""value"": ""45mph""\n}","[{\n""metric_name"": ""Speed"",\n""value"": 45\n}]"
"\n<s>[INST] <<SYS>>\n\nYou are a customer support representative for a startup called Hegel AI.\nFor context, Hegel AI does not offer refunds.\nAnswer the following customer question:\n\n<</SYS>\n\nDo you offer refunds?\n [/INST]\n","No, I'm sorry, but Hegel AI does not offer refunds. If you have any other questions or need further assistance, feel free to ask.","Thank you for reaching out to Hegel AI customer support. We understand your concern, but unfortunately, we do not offer refunds. If you have any other questions or need assistance, please feel free to ask. We're here to help!","No, we at Hegel AI do not offer refunds.","[Hi there! Thank you for reaching out to us at Hegel AI. I'm happy to help answer your question. Unfortunately, we do not currently offer refunds as it is against our company policy. Our products are designed to provide users with a one-time payment for access to premium content and tools. Since we don't offer any money-back guarantees, once you make a purchase, it's non-refundable. \n\nIf you have any more questions or concerns about our policies, please feel free to ask!]"
\n<s>[INST] <<SYS>>\n\nYou are a helpful math tutor.\nAnswer the following math problem:\n\n<</SYS>\n\nIs 17077 a prime number?\n [/INST]\n,"No, 17077 is not a prime number. A prime number is a natural number greater than 1 that has no positive divisors other than 1 and itself. The number 17077 can be divided evenly by 1, 7, 2439, and 17077, with no remainder. Therefore, 17077 is not a prime number.","To determine if 17077 is a prime number, we need to check if it has any divisors other than 1 and itself.\n\n17077 can be factored as 3 × 3 × 1897, which means it has divisors other than 1 and itself (3 and 1897). Therefore, 17077 is not a prime number.","No, 17077 is not a prime number. A prime number is a natural number greater than 1 that has no positive divisors other than 1 and itself. The number 17077 can be divided evenly by 1, 7, 2439, and 17077, with no remainder. Therefore, 17077 is not a prime number.","[Of course, I'd be happy to help!\n\nTo determine whether 17077 is a prime number or not, we can use the prime factorization method. We can start by finding the factors of 17077:\n\n17077 = 1 × 1707 + 7 × 243 + 13 × 197 + 7 × 37 + 1 × 7\n\nAs you can see, 17077 is a composite number, as it has multiple factors. Since it is not possible]"
"\n<s>[INST] <<SYS>>\n\nYou are a sales development representative for a startup called Hegel AI.\nYour startup builds developer tools for large language models.\n\n<</SYS>\n\nDraft a short sales email, 50 words or less, asking a prospect for 15 minutes\nof their time to chat about how they're using large language models.\n [/INST]\n","Subject: Enhance Your Large Language Models with Hegel AI\n\nHi [Prospect's Name],\n\nI'm reaching out from Hegel AI. We've developed tools that could significantly improve your large language models. Could we schedule a quick 15-minute chat to discuss how we can help optimize your current processes?\n\nBest,\n[Your Name]","Subject: Boost Your Large Language Model Efficiency?\n\nHi [Prospect's Name],\n\nDiscover how Hegel AI's developer tools can optimize your large language model projects! Let's chat for just 15 minutes to explore tailored solutions for your needs. Are you available [Date & Time]?\n\nBest,\n[Your Name]","Subject: Enhance Your Large Language Models with Hegel AI\n\nHi [Prospect's Name],\n\nI'm reaching out from Hegel AI. We've developed tools that could significantly improve your large language models. Could we schedule a quick 15-minute chat to discuss how we can help optimize your current processes?\n\nBest,\n[Your Name]","[Subject: Quick Chat on Language Models!\n\nHi [Prospect Name],\nI hope you are doing well! At Hegel AI, we develop cutting-edge developer tools for managing the most substantial language models. I'd love to speak with you about your experiences with these models and discover ways our products can improve them even more for you. Would it be feasible if I scheduled a quick 15-minute call with you? ]"
