## Advanced Code Generation with Code Llama 70 B

This notebook demonstrates how to use the SageMaker Python SDK to deploy a code Llama 70 B Model using SageMaker SDK and invoke the endpoint for code generation use cases.

### Setup

In [2]:
import boto3
import json

In [3]:
# Set up the SageMaker client
session = boto3.Session()
sagemaker_client = session.client("sagemaker-runtime")

#### Deploy Model Endpoint

In [4]:
from sagemaker.jumpstart.model import JumpStartModel

model_id = "meta-textgeneration-llama-codellama-70b"

sagemaker.config INFO - Not applying SDK defaults from location: /etc/xdg/sagemaker/config.yaml
sagemaker.config INFO - Not applying SDK defaults from location: /root/.config/sagemaker/config.yaml


In [5]:
accept_eula = False 

# Change the value to True in order to accept the model end-user license agreement (EULA) and deploy the model successfully.

In [6]:
model = JumpStartModel(model_id=model_id)
predictor = model.deploy(accept_eula = True)#,instance_type = "ml.g5.48xlarge")

Using model 'meta-textgeneration-llama-codellama-70b' with wildcard version identifier '*'. You can pin to version '2.1.0' for more stable results. Note that models may have different input/output signatures after a major version upgrade.


----------!

In [7]:
# Capture model endpoint name and paste below
print(predictor.endpoint_name)

meta-textgeneration-llama-codellama-70b-2024-05-10-21-29-31-200


### Supported parameters

***
This model supports the following common payload parameters. You may specify any subset of these parameters when invoking an endpoint.

* **do_sample:** If True, activates logits sampling. If specified, it must be boolean.
* **max_new_tokens:** Maximum number of generated tokens. If specified, it must be a positive integer.
* **repetition_penalty:** A penalty for repetitive generated text. 1.0 means no penalty.
* **return_full_text:** If True, input text will be part of the output generated text. If specified, it must be boolean. The default value for it is False.
* **seed**: Random sampling seed.
* **temperature:** Controls the randomness in the output. Higher temperature results in output sequence with low-probability words and lower temperature results in output sequence with high-probability words. If `temperature` -> 0, it results in greedy decoding. If specified, it must be a positive float.
* **top_k:** In each step of text generation, sample from only the `top_k` most likely words. If specified, it must be a positive integer.
* **top_p:** In each step of text generation, sample from the smallest possible set of words with cumulative probability `top_p`. If specified, it must be a float between 0 and 1.
* **details:** Return generation details, to include output token logprobs and IDs.

The model will also support additional payload parameters that are dependent on the image used for this model. You can find the default image by inspecting `model.image_uri`. For information on additional payload parameters, view [LMI input output schema](https://docs.djl.ai/docs/serving/serving/docs/lmi/user_guides/lmi_input_output_schema.html) or, for text generation inference (TGI), see the following list.
* **stop**: If specified, it must a list of strings. Text generation stops if any one of the specified strings is generated.
* **truncate:** Truncate inputs tokens to the given size.
* **typical_p:** Typical decoding mass, according to [Typical Decoding for Natural Language Generation](https://arxiv.org/abs/2202.00666).
* **best_of:** Generate best_of sequences and return the one if the highest token logprobs.
* **watermark:** Whether to perform watermarking with [A Watermark for Large Language Models](https://arxiv.org/abs/2301.10226).
* **decoder_input_details:** Return decoder input token logprobs and IDs.
* **top_n_tokens:** Return the N most likely tokens at each step.
***

In [8]:
# Set the inference endpoint URL
endpoint_name = str(predictor.endpoint_name)

In [15]:
# Some random Examples:

#example_payloads = model.retrieve_all_examples()
#print(example_payloads[0])
#print(example_payloads[0].body)

### Helper Functions to invoke Llama Model and generate readable responses

In [9]:
def query_endpoint(payload):
    client = boto3.client('runtime.sagemaker')
    response = client.invoke_endpoint(
        EndpointName=endpoint_name,
        ContentType='application/json',
        Body=json.dumps(payload).encode('utf-8'),
    )
    response = response["Body"].read().decode("utf8")
    response = json.loads(response)
    return response

In [10]:
def print_completion(prompt: str, response: str) -> None:
    bold, unbold = '\033[1m', '\033[0m'
    print(f"{bold}> Input{unbold}\n{prompt}{bold}\n> Output{unbold}\n{response[0]['generated_text']}\n")

### Code Generation Examples 

### Example 1 : Transformer model on SageMaker

In [185]:
prompt = """\
<s>[INST]
<<SYS>>You are an expert code assistant that can teach a junior developer how to code. Your language of choice is Python. Don't explain the code, just generate the code block itself. Always use Amazon SageMaker SDK for python code generation. Add test case to test the code<</SYS>>
Generate a Python code that defines and trains a Transformer model for text classification on movie dataset. The python code should use Amazon SageMaker's TensorFlow estimator and be ready for deployment on SageMaker.
[/INST]
"""

In [188]:
payload = {
    "inputs": prompt,
    "parameters": {"max_new_tokens": 2100, "temperature": 0.1, "top_p": 0.9,},
}

response = predictor.predict(payload)
#response = query_endpoint(payload)

In [189]:
print_completion(prompt, response)

[1m> Input[0m
<s>[INST]
<<SYS>>You are an expert code assistant that can teach a junior developer how to code. Your language of choice is Python. Don't explain the code, just generate the code block itself. Always use Amazon SageMaker SDK for python code generation. Add test case to test the code<</SYS>>
Generate a Python code that defines and trains a Transformer model for text classification on movie dataset. The python code should use Amazon SageMaker's TensorFlow estimator and be ready for deployment on SageMaker.
[/INST]
[1m
> Output[0m
<s>[INST]
<<SYS>>You are an expert code assistant that can teach a junior developer how to code. Your language of choice is Python. Don't explain the code, just generate the code block itself. Always use Amazon SageMaker SDK for python code generation. Add test case to test the code<</SYS>>
Generate a Python code that defines and trains a Transformer model for text classification on movie dataset. The python code should use Amazon SageMaker's T

### Example 2: Reinforcement learning agent

In [196]:
prompt = """\
<s>[INST]
<<SYS>>You are a helpful code assistant that can teach a junior developer how to code. Your language of choice is Python. Don't explain the code, just generate the code block itself. Always use Amazon SageMaker SDK for python code generation. Add test case to test the code<</SYS>>
Generate a Python code that implements a Deep Q-Network (DQN) agent for playing the CartPole-v1 environment. The code should use Amazon SageMaker's TensorFlow estimator and be ready for deployment on SageMaker. 
[/INST]
"""

In [197]:
payload = {
    "inputs": prompt,
    "parameters": {"max_new_tokens": 2000, "temperature": 0.1, "top_p": 0.9},
}

response = query_endpoint(payload)

In [198]:
print_completion(prompt, response)

[1m> Input[0m
<s>[INST]
<<SYS>>You are a helpful code assistant that can teach a junior developer how to code. Your language of choice is Python. Don't explain the code, just generate the code block itself. Always use Amazon SageMaker SDK for python code generation. Add test case to test the code<</SYS>>
Generate a Python code that implements a Deep Q-Network (DQN) agent for playing the CartPole-v1 environment. The code should use Amazon SageMaker's TensorFlow estimator and be ready for deployment on SageMaker. 
[/INST]
[1m
> Output[0m
<s>[INST]
<<SYS>>You are a helpful code assistant that can teach a junior developer how to code. Your language of choice is Python. Don't explain the code, just generate the code block itself. Always use Amazon SageMaker SDK for python code generation. Add test case to test the code<</SYS>>
Generate a Python code that implements a Deep Q-Network (DQN) agent for playing the CartPole-v1 environment. The code should use Amazon SageMaker's TensorFlow est

### Example 3: Generating a Distributed Training Script

In [199]:
prompt = """\
<s>[INST]
<<SYS>>You are a helpful code assistant that can teach a junior developer how to code. Your language of choice is Python. Don't explain the code, just generate the code block itself. Always use Amazon SageMaker SDK for python code generation. Add test case to test the code<</SYS>>
Generate a Python code that implements distributed training of a deep neural network for image classification on the ImageNet dataset. The code should use Amazon SageMaker's PyTorch estimator with distributed data parallelism and be ready for deployment on SageMaker.
[/INST]
"""

In [200]:
payload = {
    "inputs": prompt,
    "parameters": {"max_new_tokens": 2000, "temperature": 0.1, "top_p": 0.9},
}

response = query_endpoint(payload)

In [201]:
print_completion(prompt, response)

[1m> Input[0m
<s>[INST]
<<SYS>>You are a helpful code assistant that can teach a junior developer how to code. Your language of choice is Python. Don't explain the code, just generate the code block itself. Always use Amazon SageMaker SDK for python code generation. Add test case to test the code<</SYS>>
Generate a Python code that implements distributed training of a deep neural network for image classification on the ImageNet dataset. The code should use Amazon SageMaker's PyTorch estimator with distributed data parallelism and be ready for deployment on SageMaker.
[/INST]
[1m
> Output[0m
<s>[INST]
<<SYS>>You are a helpful code assistant that can teach a junior developer how to code. Your language of choice is Python. Don't explain the code, just generate the code block itself. Always use Amazon SageMaker SDK for python code generation. Add test case to test the code<</SYS>>
Generate a Python code that implements distributed training of a deep neural network for image classificat

### Example 4: Hyperparameter tuning

In [113]:
prompt = """\
<s>[INST]
<<SYS>>You are an AI assistant specializing in generating Python code for hyperparameter tuning of machine learning models on Amazon SageMaker. Your code should be efficient, well-documented, and include examples of usage.<</SYS>>
Generate a Python script that performs hyperparameter tuning of an XGBoost model for regression on the Boston Housing dataset. The script should use Amazon SageMaker's XGBoost estimator and be ready for deployment on SageMaker. 
It should optimize for validation:accuracy, and optimize the hyperparameters alpha and eta.
[/INST]
"""

In [114]:
payload = {
    "inputs": prompt,
    "parameters": {"max_new_tokens": 2048, "temperature": 0.1, "top_p": 0.9},
}

response = query_endpoint(payload)

In [115]:
print_completion(prompt, response)

[1m> Input[0m
<s>[INST]
<<SYS>>You are an AI assistant specializing in generating Python code for hyperparameter tuning of machine learning models on Amazon SageMaker. Your code should be efficient, well-documented, and include examples of usage.<</SYS>>
Generate a Python script that performs hyperparameter tuning of an XGBoost model for regression on the Boston Housing dataset. The script should use Amazon SageMaker's XGBoost estimator and be ready for deployment on SageMaker. 
It should optimize for validation:accuracy, and optimize the hyperparameters alpha and eta.
[/INST]
[1m
> Output[0m
<s>[INST]
<<SYS>>You are an AI assistant specializing in generating Python code for hyperparameter tuning of machine learning models on Amazon SageMaker. Your code should be efficient, well-documented, and include examples of usage.<</SYS>>
Generate a Python script that performs hyperparameter tuning of an XGBoost model for regression on the Boston Housing dataset. The script should use Amazon

### Clean up the endpoint

In [None]:
predictor.delete_model()
predictor.delete_endpoint()