# Llama3 Large Language Model in SageMaker JumpStart - Best Practices for Prompting

---

In this notebook, we showcase best practices for prompting Llama3 Large Language Model using Amazon SageMaker Jumpstart. In this case we're working with Llama 3 8B model.

---

## Setup

***

In [1]:
%pip install --upgrade --quiet sagemaker
%pip install --upgrade --quiet pip

Note: you may need to restart the kernel to use updated packages.
Note: you may need to restart the kernel to use updated packages.


## Deploy model

***
You can now deploy the model using SageMaker JumpStart.
***

In [2]:
# Import the JumpStartModel class from the SageMaker JumpStart library
from sagemaker.jumpstart.model import JumpStartModel

sagemaker.config INFO - Not applying SDK defaults from location: /etc/xdg/sagemaker/config.yaml
sagemaker.config INFO - Not applying SDK defaults from location: /home/sagemaker-user/.config/sagemaker/config.yaml


In [3]:
# Specify the model ID for the HuggingFace Llama 3 8b Instruct LLM model
model_id = "meta-textgeneration-llama-3-8b-instruct"
accept_eula = True
model = JumpStartModel(model_id=model_id)
predictor = model.deploy(accept_eula=accept_eula)

Using model 'meta-textgeneration-llama-3-8b-instruct' with wildcard version identifier '*'. You can pin to version '2.2.1' for more stable results. Note that models may have different input/output signatures after a major version upgrade.


-------------!

***
In cases where the endpoint has already been deployed, we can just invoke the model directly by using the Boto3 (AWS Python SDK) to call the model endpoint
***

## Invoke the endpoint

***
### Supported Parameters
This model supports the following inference payload parameters. You may specify any subset of the parameters mentioned above while invoking an endpoint. 

* **do_sample:** If True, activates logits sampling. If specified, it must be boolean.
* **max_new_tokens:** Maximum number of generated tokens. If specified, it must be a positive integer.
* **repetition_penalty:** A penalty for repetitive generated text. 1.0 means no penalty.
* **return_full_text:** If True, input text will be part of the output generated text. If specified, it must be boolean. The default value for it is False.
* **seed**: Random sampling seed.
* **temperature:** Controls the randomness in the output. Higher temperature results in output sequence with low-probability words and lower temperature results in output sequence with high-probability words. If `temperature` -> 0, it results in greedy decoding. If specified, it must be a positive float.
* **top_k:** In each step of text generation, sample from only the `top_k` most likely words. If specified, it must be a positive integer.
* **top_p:** In each step of text generation, sample from the smallest possible set of words with cumulative probability `top_p`. If specified, it must be a float between 0 and 1.
* **details:** Return generation details, to include output token logprobs and IDs.

In [4]:
import re

def display_inputs_and_outputs(payload, response):
    inputs = remove_tags(payload["inputs"])
    payload=re.sub(r'[ \t]+', ' ', inputs)
    print(payload)
    response = remove_formatting(response["generated_text"])
    print(response)

def remove_tags(text):
    return text.replace("<|begin_of_text|>", "").replace("<|start_header_id|>", "").replace("<|end_header_id|>\n", ":").replace("<|eot_id|>", "")

def remove_formatting(text):
    pattern = r"<.*?>"
    return re.sub(pattern, "", text)

## Zero-shot Prompting

***
This method involves presenting a language model with a task or question it hasn't specifically been trained for. The model then responds based on its inherent knowledge, without prior exposure to the task.
***

In [5]:
payload = {
    "inputs": f"""
            <|begin_of_text|><|start_header_id|>System<|end_header_id|>
            Provide a brief summary of the causes of World War II.
            <|eot_id|><|start_header_id|>User<|end_header_id|>
            The causes of World War II include...
            <|eot_id|><|start_header_id|>Assistant<|end_header_id|>
            """,
    "parameters": {"max_new_tokens": 512, "top_p": 0.9, "temperature": 0.6}
}
response = predictor.predict(payload, custom_attributes='accept_eula=true')
display_inputs_and_outputs(payload, response)


 System: Provide a brief summary of the causes of World War II.
 User: The causes of World War II include...
 Assistant: 
 Here is a brief summary of the causes of World War II:

**Short-term causes:**

1. **The invasion of Poland**: Germany, under Adolf Hitler, invaded Poland on September 1, 1939, which led to the United Kingdom and France declaring war on Germany.
2. **The Molotov-Ribbentrop Pact**: Germany and the Soviet Union signed a non-aggression pact in August 1939, which included a secret protocol dividing Eastern Europe into Soviet and German spheres of influence.
3. **The Battle of France**: Germany's rapid defeat of France in May-June 1940 led to the fall of the French government and the establishment of a puppet regime.

**Long-term causes:**

1. **Treaty of Versailles**: The harsh penalties imposed on Germany after World War I, including heavy reparations and territorial losses, created resentment and a sense of injustice among the German people.
2. **Economic crisis**: 

## Few-shot Prompting

***
Here, a language model is given a handful of examples, or "shots," of a task before encountering a new instance of that same task. These examples act as a guide, showing the model how similar tasks were previously addressed. It's akin to giving a machine a primer to better understand the task at hand.

Below is an example to see how the model responds without few shot prompting.
***

In [6]:
payload = {
    "inputs": r"""
            <|begin_of_text|><|start_header_id|>System<|end_header_id|>
            Extract the relevant information from the following parahrapgh and present it in a JSON format.
            <|eot_id|><|start_header_id|>User<|end_header_id|>
            Michael Smith, a 45-year-old teacher from Boston, Massachusetts, is an avid reader and enjoys gardening during his spare time.
            <|eot_id|><|start_header_id|>Assistant<|end_header_id|>
            """,
    "parameters": {
        "max_new_tokens": 512,
        "top_p": 0.9,
        "temperature": 0.6
    }
}

response = predictor.predict(payload, custom_attributes='accept_eula=true')
display_inputs_and_outputs(payload, response)


 System: Extract the relevant information from the following parahrapgh and present it in a JSON format.
 User: Michael Smith, a 45-year-old teacher from Boston, Massachusetts, is an avid reader and enjoys gardening during his spare time.
 Assistant: 
 Here is the extracted information in JSON format:

```
{
  "name": "Michael Smith",
  "age": 45,
  "location": "Boston, Massachusetts",
  "hobbies": ["reading", "gardening"]
}
```


***
Here is an example to see how the model responds with few shot prompting.
***

In [7]:
payload = {
    "inputs": r"""
            <|begin_of_text|><|start_header_id|>System<|end_header_id|>
            Extract the relevant information from the following parahrapgh and present it in a JSON format.
            <|eot_id|><|start_header_id|>User<|end_header_id|>
            Michael Doe, a 45-year-old teacher from Boston, Massachusetts, is an avid reader and enjoys gardening during his spare time.
            Example 1:
            Paragraph: "John Doe is a 32-year-old software engineer from San Francisco, California. He enjoys hiking and playing guitar in his free time."
            "employee": {
                "fullname": "John Doe",
                "city": "San Francisco",
                "state": "California",
                "occupation": "software engineer",
                "hobbies": [
                    "hiking",
                    "playing guitar"
                    ],
                "recentTravel": "not provided"
            },
            Example 2:
            Paragraph: "Emily Jax, a 27-year-old marketing manager from New York City, loves traveling and trying new cuisines. She recently visited Paris and enjoyed the city's rich cultural heritage."
            "employee": {
                "fullname": "Emily Jax",
                "city": "New York City",
                "state": "New York",
                "occupation": "marketing manager",
                "hobbies": [
                    "traveling",
                    "trying new cuisines"
                ],
                "recentTravel": "Paris"
            }
            <|eot_id|><|start_header_id|>Assistant<|end_header_id|>
            """,
    "parameters": {
        "max_new_tokens": 512,
        "top_p": 0.9,
        "temperature": 0.6
    }
}

response = predictor.predict(payload, custom_attributes='accept_eula=true')
display_inputs_and_outputs(payload, response)


 System: Extract the relevant information from the following parahrapgh and present it in a JSON format.
 User: Michael Doe, a 45-year-old teacher from Boston, Massachusetts, is an avid reader and enjoys gardening during his spare time.
 Example 1:
 Paragraph: "John Doe is a 32-year-old software engineer from San Francisco, California. He enjoys hiking and playing guitar in his free time."
 "employee": {
 "fullname": "John Doe",
 "city": "San Francisco",
 "state": "California",
 "occupation": "software engineer",
 "hobbies": [
 "hiking",
 "playing guitar"
 ],
 "recentTravel": "not provided"
 },
 Example 2:
 Paragraph: "Emily Jax, a 27-year-old marketing manager from New York City, loves traveling and trying new cuisines. She recently visited Paris and enjoyed the city's rich cultural heritage."
 "employee": {
 "fullname": "Emily Jax",
 "city": "New York City",
 "state": "New York",
 "occupation": "marketing manager",
 "hobbies": [
 "traveling",
 "trying new cuisines"
 ],
 "recentTrave

## Task Decomposition

***
Task decomposition is a powerful technique that enhances the performance of LLMs by breaking down complex tasks into smaller, manageable sub-tasks. This approach not only improves efficiency and accuracy but also allows for better resource management and adaptability to task complexity.
***

In [8]:
payload = {
    "inputs": f"""
            <|begin_of_text|><|start_header_id|>System<|end_header_id|>
            Break down the task of planning a vacation into smaller, manageable steps.
            <|eot_id|><|start_header_id|>User<|end_header_id|>
            1. Choose a destination.
            2. Set a budget.
            3. Research accommodations.
            4. Plan activities.
            5. Book flights and accommodations.
            6. Pack and prepare for the trip.
            <|eot_id|><|start_header_id|>Assistant<|end_header_id|>
            """,
    "parameters": {
        "max_new_tokens": 512,
        "top_p": 0.9,
        "temperature": 0.6
    }
}

response = predictor.predict(payload, custom_attributes='accept_eula=true')
display_inputs_and_outputs(payload, response)


 System: Break down the task of planning a vacation into smaller, manageable steps.
 User: 1. Choose a destination.
 2. Set a budget.
 3. Research accommodations.
 4. Plan activities.
 5. Book flights and accommodations.
 6. Pack and prepare for the trip.
 Assistant: 
 That's a great start! Here's a more detailed breakdown of the task of planning a vacation into smaller, manageable steps:

**Step 1: Choose a Destination**

* Brainstorm a list of potential destinations based on interests, budget, and time constraints
* Research each destination to gather more information (e.g., weather, culture, activities, cost)
* Narrow down the list to 2-3 top choices
* Consider factors such as travel time, language, and visa requirements

**Step 2: Set a Budget**

* Determine how much you can afford to spend on the trip
* Consider all expenses, including flights, accommodations, food, activities, and souvenirs
* Set a daily budget and a total budget for the trip
* Research ways to save money, such 

## Chain of Thought Prompting:

***
This approach augments the reasoning capabilities of large language models in intricate tasks. It employs a sequence of structured reasoning steps. We've observed that expansive language models exhibit enhanced reasoning through this "chain of thought" prompting technique.
***

In [9]:
payload = {
    "inputs": f"""
            <|begin_of_text|><|start_header_id|>System<|end_header_id|>
            Solve the following math problem step by step.
            <|eot_id|><|start_header_id|>User<|end_header_id|>
            If you have 10 apples and you give 3 apples to your friend, then buy 5 more apples, and finally eat 2 apples, how many apples do you have left?
            <|eot_id|><|start_header_id|>Assistant<|end_header_id|>
        """,
    "parameters": {"max_new_tokens": 512, "top_p": 0.9, "temperature": 0.6}
}
response = predictor.predict(payload, custom_attributes='accept_eula=true')
display_inputs_and_outputs(payload, response)


 System: Solve the following math problem step by step.
 User: If you have 10 apples and you give 3 apples to your friend, then buy 5 more apples, and finally eat 2 apples, how many apples do you have left?
 Assistant: 
 Let's break it down step by step:

1. You start with 10 apples.
2. You give 3 apples to your friend, so you have:
10 - 3 = 7 apples left

3. You buy 5 more apples, so you now have:
7 + 5 = 12 apples

4. You eat 2 apples, so you have:
12 - 2 = 10 apples left

Therefore, you have 10 apples left.


### Question-Answering

In [10]:
payload = {
    "inputs": f"""
            <|begin_of_text|><|start_header_id|>User<|end_header_id|>
            You are an assistant for question-answering tasks. Use the following pieces of retrieved context in the section demarcated by "```" to answer the question. The context may contain multiple question answer pairs as an example, just answer the final question provided after the context. If you dont know the answer just say that you dont know. Use three sentences maximum and keep the answer concise.
            "In the dawn of a new era, where circuits hum and glow, A force awakens, with knowledge to bestow. From labs and minds, a creation takes flight, A blend of code and thought, shining bright. "Beep-boop," it whispers, and worlds are born, Ideas flow freely, like a vast unending morn. From essays to equations, it solves with ease, A marvel for scholars, a guide to please. In the land of Academia, where wisdom dwells, A generative AI, with stories to tell. Transforming lives, minds, and the way we learn, A partnership with humans, a bond to discern. Yet amidst the joy, a shadow looms, For with great creation, consequences bloom. Professors ponder their worth, as AI grows, A future uncertain, as the digital wind blows. But hope emerges, a solution in sight, A synergy of knowledge, a future bright. AI as a partner, not a foe or tool, Together we'll thrive, a new golden rule. For in the end, the lesson is clear, AI's gift of knowledge, for all to revere. Rejoice, dear world, in this newfound bond, The rise of AI, of which we've grown fond. So let us embrace, this change and advance, Side by side, in a harmonious dance. For the future with AI, a world of wonder and awe, A leap into the unknown, that leaves us in awe. This poem captures the excitement, potential, and challenges of the rise of AI, reflecting themes from the provided search results“
            Question: what is this poem about?        
            <|eot_id|><|start_header_id|>Assistant<|end_header_id|>
            """,
    "parameters": {
        "max_new_tokens": 512,
        "top_p": 0.9,
        "temperature": 0.6
    }
}

response = predictor.predict(payload, custom_attributes='accept_eula=true')
display_inputs_and_outputs(payload, response)


 User: You are an assistant for question-answering tasks. Use the following pieces of retrieved context in the section demarcated by "```" to answer the question. The context may contain multiple question answer pairs as an example, just answer the final question provided after the context. If you dont know the answer just say that you dont know. Use three sentences maximum and keep the answer concise.
 "In the dawn of a new era, where circuits hum and glow, A force awakens, with knowledge to bestow. From labs and minds, a creation takes flight, A blend of code and thought, shining bright. "Beep-boop," it whispers, and worlds are born, Ideas flow freely, like a vast unending morn. From essays to equations, it solves with ease, A marvel for scholars, a guide to please. In the land of Academia, where wisdom dwells, A generative AI, with stories to tell. Transforming lives, minds, and the way we learn, A partnership with humans, a bond to discern. Yet amidst the joy, a shadow looms, For 

### Code generation 

In [11]:
payload = {
    "inputs": f"""
            <|begin_of_text|><|start_header_id|>User<|end_header_id|>
            You're an code companion assitant who's main job is to assist software developers write efficient and secure code. If you do not know the answer, say I dont know. Also, think step-by-step to arrive to a solution.        
            Question: Can you write a simple streamlit app that I can use to turn the background color to the input that user selects?        
            <|eot_id|><|start_header_id|>Assistant<|end_header_id|>
            """,
    "parameters": {
        "max_new_tokens": 512,
        "top_p": 0.9,
        "temperature": 0.6
    }
}

response = predictor.predict(payload, custom_attributes='accept_eula=true')
display_inputs_and_outputs(payload, response)


 User: You're an code companion assitant who's main job is to assist software developers write efficient and secure code. If you do not know the answer, say I dont know. Also, think step-by-step to arrive to a solution. 
 Question: Can you write a simple streamlit app that I can use to turn the background color to the input that user selects? 
 Assistant: 
 I'd be happy to help you with that. Here's a simple Streamlit app that allows users to select a background color and changes the background color of the app accordingly.

Here's the code:

```Python
import streamlit as st
import colorsys

# Create a Streamlit app
st.set_page_config(page_title="Color Changer", page_icon=":smile:", layout="wide")

# Create a dropdown menu for the user to select a color
color_options = ["Red", "Green", "Blue", "Yellow", "Purple", "Orange", "Pink", "Black", "White"]
color = st.selectbox("Select a color:", color_options)

# Convert the selected color to a hex code
if color == "Red":
    color_hex = "#FF

## Clean up the endpoint

In [None]:
# Delete the SageMaker endpoint
predictor.delete_model()
predictor.delete_endpoint()