# LLama 2

## Prepare System

Login to Hugging Face in case of wanting to load new model that is not in `model_cache`

In [1]:
# !huggingface-cli login --token <yourtoken>


Check if GPU is available

In [1]:
import torch

torch.cuda.is_available()

True

## Load Model

Only run the block that you want

In [2]:
# 7B Chat

from transformers import AutoTokenizer, AutoModelForCausalLM

cache_directory = "./model_cache"

model_name = "meta-llama/Llama-2-7b-chat-hf"

# Download and cache the tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_name, cache_dir=cache_directory)
model = AutoModelForCausalLM.from_pretrained(model_name, cache_dir=cache_directory)

Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]

In [None]:
# 13B Chat

from transformers import AutoTokenizer, AutoModelForCausalLM

cache_directory = "./model_cache"

model_name = "meta-llama/Llama-2-13b-chat-hf"

# Download and cache the tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_name, cache_dir=cache_directory)
model = AutoModelForCausalLM.from_pretrained(model_name, cache_dir=cache_directory)

70B cannot fit in one GPU. It need to be optimised to use 2 GPU to run in our system.

In [1]:
# # 70B chat

# from transformers import AutoTokenizer, AutoModel

# cache_directory = "./model_cache"  # Replace with your desired cache directory

# model_name = "meta-llama/Llama-2-70b-chat-hf"

# tokenizer = AutoTokenizer.from_pretrained(model_name, cache_dir=cache_directory)
# model = AutoModel.from_pretrained(model_name, cache_dir=cache_directory)


## Prepare Data

Try top 5 row of our dataset

In [3]:
import pandas as pd
pd.set_option('display.max_colwidth', None)


# read prompt-1.csv 
df = pd.read_csv('prompt-3.csv')

df = df.head(5)

promptTemplate = """<s>[INST] <<SYS>>
{system_prompt}
<</SYS>>

{user_message} [/INST]"""

# Define your system prompt and user message
system_prompt = "You are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.\n\nIf a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information."

# Define a function to format the prompt
def format_prompt(row):
    user_message = row['prompt']
    return promptTemplate.format(system_prompt=system_prompt, user_message=user_message)

# Apply the function to the prompt column
df['formatted_prompt'] = df.apply(format_prompt, axis=1)

df = df.drop(columns=['formatted_prompt'])

df

Unnamed: 0,id,prompt
0,1,"Following is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\nEIA is study of probable change in socio economic and characteristics of environment due to proposed action.\n\nOptions:\nA. biophysical, B. geological, C. hydrological, D. Potential"
1,2,"Following is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\nEIA is __ and __ environmental impacts.\n\nOptions:\nA. beneficial and adverse, B. social and benefit, C. adverse and economical, D. None of above"
2,3,"Following is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\n___ is evident that order to identify, assess and predict the environmental impacts.\n\nOptions:\nA. Environmental attributes, B. Environmental impact, C. Complex system, D. Variable parameter"
3,4,"Following is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\nEIA is including mechanisms for monitoring, and evaluation.\n\nOptions:\nA. specification, B. auditing, C. reporting, D. concluding"
4,5,"Following is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\nEIA is study to delineate between beneficial and detrimental impacts.\n\nOptions:\nA. micro, B. action, C. proposed, D. macro"


## Initialise pipeline

In [4]:
import transformers
import torch

pipeline = transformers.pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
    torch_dtype=torch.float16,
    device=0,
)

## Try Prompting

In [15]:
# https://github.com/facebookresearch/llama/blob/main/example_chat_completion.py

myPrompt = """<s>[INST] <<SYS>>
{system_prompt}
<</SYS>>

{user_message} [/INST]"""

# Define your system prompt and user message
system_prompt = """You are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe.  Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.

If a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information."""

user_message = "Where is Australia"

# Use the format method to insert your defined prompts into the template
formatted_prompt = myPrompt.format(system_prompt=system_prompt, user_message=user_message)

print(formatted_prompt)

sequences = pipeline(
    formatted_prompt,
    do_sample=True,
    num_return_sequences=1,
    eos_token_id=tokenizer.eos_token_id,
    return_full_text=False, # to not repeat the question, set to False
    # max_length=512, # max length of output, default=4096
    max_new_tokens=1024,
    top_k=50, # default=10
    top_p=0.9, # default=0.9
    temperature=0.6, # default=0.6
    repetition_penalty=1.2, # default 1.0, to prevent repetition (higher more prevention)
)

print(sequences[0]['generated_text'])

<s>[INST] <<SYS>>
You are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe.  Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.If a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information.
<</SYS>>

Where is Australia [/INST]
  Thank you for asking! Australia is a beautiful country located in the southern hemisphere, surrounded by the Pacific Ocean and the Indian Ocean. It is officially known as the Commonwealth of Australia and is a federal parliamentary constitutional monarchy. The country is home to diverse landscapes, including beaches, deserts, rainforests, and mountains.
Australia is situated between latitudes 10° and 40° south of the equator, which 

## Try looping on our dataset

In [19]:
df['answer'] = ''

# Loop through each row in the DataFrame and generate text
for index, row in df.iterrows():
    prompt = row['prompt']
    generated_text = pipeline(
        prompt,
        do_sample=True,
        num_return_sequences=1,
        eos_token_id=tokenizer.eos_token_id,
        return_full_text=False,
        max_new_tokens=1024,
        top_k=50,
        top_p=0.9,
        temperature=0.6,
        repetition_penalty=1.2,
    )[0]['generated_text']

    # Assign the generated text to the 'answer' column in the same row
    df.at[index, 'answer'] = generated_text

# Create new DataFrame to store id and generated text
output_file_path = 'test.csv'
df.to_csv(output_file_path, index=False)

df



Unnamed: 0,id,prompt,answer
0,1,"<s>[INST] <<SYS>>\nYou are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.\n\nIf a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information.\n<</SYS>>\n\nFollowing is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\nEIA is study of probable change in socio economic and characteristics of environment due to proposed action.\n\nOptions:\nA. biophysical, B. geological, C. hydrological, D. Potential [/INST]","Thank you for asking! I'm here to help you with your query. After carefully reviewing the given options, I believe the correct answer is (A) Biophysical.\nThe term ""EIA"" stands for Environmental Impact Assessment, which refers to the study of the likely changes in the socio-economic and environmental characteristics of an area resulting from a proposed action or project. Biophysical factors refer to the physical and biological aspects of the environment, such as land use patterns, water resources, air quality, and biodiversity. Therefore, option A best describes the meaning of EIA."
1,2,"<s>[INST] <<SYS>>\nYou are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.\n\nIf a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information.\n<</SYS>>\n\nFollowing is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\nEIA is __ and __ environmental impacts.\n\nOptions:\nA. beneficial and adverse, B. social and benefit, C. adverse and economical, D. None of above [/INST]","Thank you for asking! I'm happy to help you with this question.\nOption A, ""beneficial and adverse,"" is the correct answer. EIA stands for Environmental Impact Assessment, which is a process used to identify and evaluate potential environmental impacts of projects or policies before they are implemented. The goal of an EIA is to minimize negative environmental impacts and promote sustainable development. Therefore, EIA considers both beneficial and adverse environmental impacts."
2,3,"<s>[INST] <<SYS>>\nYou are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.\n\nIf a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information.\n<</SYS>>\n\nFollowing is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\n___ is evident that order to identify, assess and predict the environmental impacts.\n\nOptions:\nA. Environmental attributes, B. Environmental impact, C. Complex system, D. Variable parameter [/INST]","Thank you for asking! I'm here to help you with your question. However, I must point out that the question contains some ambiguity, which makes it difficult to provide an accurate answer. Specifically, the term ""environmental impact"" could refer to various different concepts depending on the context.\nTo clarify, there are several types of environmental impacts that need to be identified, assessed, and predicted in order to fully comprehend their effects on the environment. These include:\n1. Direct Impacts: These are the immediate consequences of human activities on the environment, such as habitat destruction, water pollution, and climate change.\n2. Indirect Impacts: These are the long-term consequences of human activities on the environment, such as changes in ecosystem services, loss of biodiversity, and altered carbon cycles.\n3. Chronic Impacts: These are the ongoing and persistent impacts of human activities on the environment, such as air and water pollution, soil degradation, and overfishing.\n4. Acute Impacts: These are sudden and short-term impacts of human activities on the environment, such as oil spills, chemical contamination, and natural disasters like floods and wildfires.\nIn light of this clarification, I would suggest selecting Option A, ""Environmental attributes,"" as the best answer. This refers to the characteristics of the environment that can affect its quality and functioning, including physical, biological, and social factors. Understanding these attributes is essential for identifying, assessing, and predicting environmental impacts."
3,4,"<s>[INST] <<SYS>>\nYou are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.\n\nIf a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information.\n<</SYS>>\n\nFollowing is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\nEIA is including mechanisms for monitoring, and evaluation.\n\nOptions:\nA. specification, B. auditing, C. reporting, D. concluding [/INST]","Thank you for asking! I'm here to help you with your question. However, I must point out that the question contains some ambiguity and could lead to incorrect assumptions. Therefore, I will provide my answer in a responsible and safe manner.\nThe term ""EIA"" stands for Environmental Impact Assessment, which is a process used to identify, predict, and mitigate potential environmental impacts of projects before they are implemented. Monitoring and evaluation are indeed important components of EIA, but they cannot be reduced to just a few options given above.\nMonitoring refers to the ongoing observation and measurement of the environment and its changes over time, often using data collected through various methods such as field observations, remote sensing, and modeling. Evaluation, on the other hand, involves assessing the effectiveness and efficiency of an EIA process, identifying areas for improvement, and determining whether the predicted environmental impacts have occurred.\nTherefore, the correct answer would be both A and B, as well as C and D. These options cover different aspects of the broader concept of monitoring and evaluation within the context of EIA."
4,5,"<s>[INST] <<SYS>>\nYou are a helpful, respectful and honest assistant. Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, toxic, dangerous, or illegal content. Please ensure that your responses are socially unbiased and positive in nature.\n\nIf a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don't know the answer to a question, please don't share false information.\n<</SYS>>\n\nFollowing is a question, understand it and from the provided four options 'A', 'B', 'C', 'D', select the correct option (Either 'A' or 'B' or 'C' or 'D'). There may be more than one correct answer for each question so you can also select multiple options if they are the answers.\n\nQuestion:\nEIA is study to delineate between beneficial and detrimental impacts.\n\nOptions:\nA. micro, B. action, C. proposed, D. macro [/INST]","Thank you for asking! I'm here to help you with your question. However, I must point out that the term ""EIA"" could refer to several different things, depending on the context. Could you please provide more information or clarify which EIA you are referring to? This will allow me to give you the most accurate and relevant response possible."
