<a href="https://colab.research.google.com/github/fatemafaria142/Advancing-Agronomic-Question-and-Answering-through-Various-Language-Models/blob/main/Agronomic_Question_and_Answering_using_openchat_3_5_1210.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>

### **Install Required Packages**

In [1]:
!pip install accelerate peft bitsandbytes transformers trl datasets torch

Collecting accelerate
  Downloading accelerate-0.26.1-py3-none-any.whl (270 kB)
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m270.9/270.9 kB[0m [31m4.7 MB/s[0m eta [36m0:00:00[0m
[?25hCollecting peft
  Downloading peft-0.7.1-py3-none-any.whl (168 kB)
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m168.3/168.3 kB[0m [31m9.0 MB/s[0m eta [36m0:00:00[0m
[?25hCollecting bitsandbytes
  Downloading bitsandbytes-0.42.0-py3-none-any.whl (105.0 MB)
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m105.0/105.0 MB[0m [31m2.1 MB/s[0m eta [36m0:00:00[0m
Collecting trl
  Downloading trl-0.7.10-py3-none-any.whl (150 kB)
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m150.9/150.9 kB[0m [31m21.1 MB/s[0m eta [36m0:00:00[0m
[?25hCollecting datasets
  Downloading datasets-2.16.1-py3-none-any.whl (507 kB)
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m507.1/507.1 kB[0m [31m42.7 MB/s[0m eta [36m0:00:00

In [2]:
from google.colab import drive
# Mount Google Drive
drive.mount('/content/drive')

Mounted at /content/drive


#### **Dataset Link:** https://github.com/JonaOmara/AgroQA-Dataset/tree/main

In [3]:
import pandas as pd
dataset = pd.read_csv("/content/drive/MyDrive/AgroQA Dataset.csv")
dataset.head()

Unnamed: 0,Crop,Question,Answer
0,maize,"Apart from hand weeding, what other method use...",Machinery weeders are available
1,beans,"Apart from insecticide, what other method used...",Use resistant verities and increase on water a...
2,maize,Apart from sun drying which other method used ...,Use tarpaulins or cemented floor free from dust
3,cassava,"Apart from sun drying, what other method can I...",Solar driers
4,beans,As a farmer when should I harvest beans.,When the beans pods are yellowish green or dry...


In [4]:
length_of_data = len(dataset)
# Print the length
print(length_of_data)

3044


In [5]:
from datasets import Dataset, DatasetDict

# Create a DatasetDict for instruct_tune_dataset
instruct_tune_dataset = DatasetDict()

# Select 1500 rows for the 'train' split and only keep 'Question' and 'Answer'
instruct_tune_dataset["train"] = Dataset.from_pandas(dataset[["Question", "Answer"]].head(2700))

# Select 200 rows for the 'test' split and only keep 'Question' and 'Answer'
instruct_tune_dataset["test"] = Dataset.from_pandas(dataset[["Question", "Answer"]].tail(200))

# Print the updated instruct_tune_dataset
print(instruct_tune_dataset)

DatasetDict({
    train: Dataset({
        features: ['Question', 'Answer'],
        num_rows: 2700
    })
    test: Dataset({
        features: ['Question', 'Answer'],
        num_rows: 200
    })
})


### **Prompt Creation**

In [6]:
def create_prompt(sample):
    # Use a predefined template for instructions
    prompt =  "Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. "
    prompt += sample['Question']
    single_turn_prompt = f"GPT4 Correct User: {prompt}<|end_of_turn|>GPT4 Correct Assistant: {sample['Answer']}"
    return single_turn_prompt

In [7]:
create_prompt(instruct_tune_dataset["train"][0])

"GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. Apart from hand weeding, what other method used to weed maize<|end_of_turn|>GPT4 Correct Assistant: Machinery weeders are available"

In [8]:
create_prompt(instruct_tune_dataset["train"][1])

"GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. Apart from insecticide, what other method used to control bean weevils?<|end_of_turn|>GPT4 Correct Assistant: Use resistant verities and increase on water availability for crop vigor"

In [9]:
create_prompt(instruct_tune_dataset["train"][2])

"GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. Apart from sun drying which other method used for drying maize<|end_of_turn|>GPT4 Correct Assistant: Use tarpaulins or cemented floor free from dust"

In [10]:
create_prompt(instruct_tune_dataset["train"][3])

"GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. Apart from sun drying, what other method can I use to dry cassava<|end_of_turn|>GPT4 Correct Assistant: Solar driers"

In [11]:
create_prompt(instruct_tune_dataset["train"][4])

"GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. As a farmer when should I harvest beans.<|end_of_turn|>GPT4 Correct Assistant: When the beans pods are yellowish green or dry brown"

In [12]:
create_prompt(instruct_tune_dataset["train"][5])

"GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. As a farmer which method is good to prevent soil erosion<|end_of_turn|>GPT4 Correct Assistant: Farmers can terrace their gardens or mulch their gardens to stop erosion"

### **Initializing the Model**
* Load the model using a 4-bit configuration, employing double quantization, and set bfloat16 as the compute data type.

* Notably, we opt for the instruct-tuned model in this instance rather than the base model. It's worth mentioning that fine-tuning a base model necessitates a more substantial amount of data!

In [13]:
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch

bnb_config = BitsAndBytesConfig(
        load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype="float16", bnb_4bit_use_double_quant=True
    )


* **Model and Tokenizer:** https://huggingface.co/openchat/openchat-3.5-1210

In [14]:
mode_id = "openchat/openchat-3.5-1210"

In [15]:
model = AutoModelForCausalLM.from_pretrained(
        mode_id, quantization_config=bnb_config, device_map="auto", use_cache=False
    )

The secret `HF_TOKEN` does not exist in your Colab secrets.
To authenticate with the Hugging Face Hub, create a token in your settings tab (https://huggingface.co/settings/tokens), set it as secret in your Google Colab and restart your session.
You will be able to reuse this secret in all of your notebooks.
Please note that authentication is recommended but still optional to access public models or datasets.


config.json:   0%|          | 0.00/623 [00:00<?, ?B/s]

model.safetensors.index.json:   0%|          | 0.00/23.9k [00:00<?, ?B/s]

Downloading shards:   0%|          | 0/3 [00:00<?, ?it/s]

model-00001-of-00003.safetensors:   0%|          | 0.00/4.94G [00:00<?, ?B/s]

model-00002-of-00003.safetensors:   0%|          | 0.00/5.00G [00:00<?, ?B/s]

model-00003-of-00003.safetensors:   0%|          | 0.00/4.54G [00:00<?, ?B/s]

Loading checkpoint shards:   0%|          | 0/3 [00:00<?, ?it/s]

generation_config.json:   0%|          | 0.00/179 [00:00<?, ?B/s]



In [16]:
tokenizer = AutoTokenizer.from_pretrained("openchat/openchat-3.5-1210")
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "right"

tokenizer_config.json:   0%|          | 0.00/1.62k [00:00<?, ?B/s]

tokenizer.model:   0%|          | 0.00/493k [00:00<?, ?B/s]

tokenizer.json:   0%|          | 0.00/1.80M [00:00<?, ?B/s]

added_tokens.json:   0%|          | 0.00/53.0 [00:00<?, ?B/s]

special_tokens_map.json:   0%|          | 0.00/502 [00:00<?, ?B/s]

Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.


### **Let's example how well the model does at this task currently:**

In [17]:
def generate_response(prompt, model):
  encoded_input = tokenizer(prompt,  return_tensors="pt", add_special_tokens=True)
  model_inputs = encoded_input.to('cuda')

  generated_ids = model.generate(**model_inputs, max_new_tokens=256, do_sample=True, pad_token_id=tokenizer.eos_token_id)

  decoded_output = tokenizer.batch_decode(generated_ids)

  return decoded_output[0].replace(prompt, "")

In [18]:
prompt = """GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. As a farmer which method is good to prevent soil erosion? <|end_of_turn|>GPT4 Correct Assistant:"""

In [19]:
generate_response(prompt, model)



"<s> GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. As a farmer which method is good to prevent soil erosion? <|end_of_turn|> GPT4 Correct Assistant: As an agronomist, I recommend the use of conservation methods to prevent soil erosion and maintain soil fertility. Here are some key practices a farmer can implement:\n\n1. Contour plowing: This involves plowing along the natural contours of the land, which helps to slow down water runoff and reduce the impact of erosion.\n\n2. Crop rotation: By rotating different crops, you can improve soil structure and increase organic matter, which enhances soil stability and reduces erosion.\n\n3. Planting cover crops: Cover crops, like clover or rye, can be planted in off-season to protect the soil and help

### **Setting up the Training**
we will be using the `huggingface` and the `peft` library!

In [20]:
from peft import AutoPeftModelForCausalLM, LoraConfig, get_peft_model, prepare_model_for_kbit_training

peft_config = LoraConfig(r=8, lora_alpha=16, lora_dropout=0.05, bias="none", task_type="CAUSAL_LM")


* we need to prepare the model to be trained in 4bit so we will use the  **`prepare_model_for_kbit_training`** function from peft




In [21]:
model = prepare_model_for_kbit_training(model)
model = get_peft_model(model, peft_config)

# **Training Hyperparameters**
The choice of hyperparameters is contingent upon the desired training duration. Pay special attention to the following key factors:

* `num_train_epochs/max_steps:` Dictates the number of iterations over the data. Exercise caution, as an excessive number may lead to overfitting!

* `learning_rate:` Governs the convergence speed of the model. Adjust this parameter judiciously for optimal results.

In [22]:
from transformers import TrainingArguments
output_model= "openchat_instruct_generation"
training_arguments = TrainingArguments(
        output_dir=output_model,
        per_device_train_batch_size=1,
        gradient_accumulation_steps=4,
        optim="paged_adamw_32bit",
        learning_rate=2e-4,
        lr_scheduler_type="cosine",
        save_strategy="epoch",
        logging_steps=10,
        num_train_epochs=1,
        max_steps=100,
        fp16=True,
)


### **Setting up the trainer**

`max_seq_length`: Context window size


In [23]:
from trl import SFTTrainer

max_seq_length = 256

trainer = SFTTrainer(
  model=model,
  peft_config=peft_config,
  max_seq_length=max_seq_length,
  tokenizer=tokenizer,
  packing=True,
  formatting_func=create_prompt, # this will aplly the create_prompt mapping to all training and test dataset
  args=training_arguments,
  train_dataset=instruct_tune_dataset["train"],
  eval_dataset=instruct_tune_dataset["test"]
)

Generating train split: 0 examples [00:00, ? examples/s]

Generating train split: 0 examples [00:00, ? examples/s]



### **Training starts here**

In [24]:
trainer.train()

You're using a LlamaTokenizerFast tokenizer. Please note that with a fast tokenizer, using the `__call__` method is faster than using a method to encode the text followed by a call to the `pad` method to get a padded encoding.


Step,Training Loss
10,1.8567
20,1.0232
30,0.6559
40,0.5695
50,0.5712
60,0.5624
70,0.5237
80,0.5245
90,0.5263
100,0.5546


TrainOutput(global_step=100, training_loss=0.7368033361434937, metrics={'train_runtime': 554.2397, 'train_samples_per_second': 0.722, 'train_steps_per_second': 0.18, 'total_flos': 4370888392704000.0, 'train_loss': 0.7368033361434937, 'epoch': 0.4})

### **Save the model**

In [25]:
trainer.save_model("openchat_instruct_generation")

In [26]:
merged_model = model.merge_and_unload()



In [27]:
def generate_response(prompt, model):
  encoded_input = tokenizer(prompt,  return_tensors="pt", add_special_tokens=True)
  model_inputs = encoded_input.to('cuda')

  generated_ids = model.generate(**model_inputs, max_new_tokens=256, do_sample=True, pad_token_id=tokenizer.eos_token_id)

  decoded_output = tokenizer.batch_decode(generated_ids)

  return decoded_output[0]

### **Example No:1**

In [28]:
prompt = "GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. "
prompt += "What causes premature stem cracks in maize? <|end_of_turn|> "
prompt += "GPT4 Correct Assistant:"
response = generate_response(prompt, model)
# Print the response with formatted output
print(response)



<s> GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. What causes premature stem cracks in maize? <|end_of_turn|>  GPT4 Correct Assistant: Premature stem cracks in maize can be caused by a variety of factors, such as drought, excess moisture, inadequate fertilizer, and pest infestations. A vigilant and experienced agronomist will be able to diagnose the problem and recommend the appropriate measures for addressing it.<|end_of_turn|>


### **Example No:2**

In [29]:
prompt = "GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. "
prompt += "Which variety of cassava is tolerant to disease? <|end_of_turn|> "
prompt += "GPT4 Correct Assistant:"
response = generate_response(prompt, model)
# Print the response with formatted output
print(response)

<s> GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. Which variety of cassava is tolerant to disease? <|end_of_turn|>  GPT4 Correct Assistant: Germplasm collection such as cassava brown streak virus resistance can tolerate disease.<|end_of_turn|>


### **Example No:3**

In [30]:
prompt = "GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. "
prompt += "can rain water reduce power of chemical after spraying?  <|end_of_turn|> "
prompt += "GPT4 Correct Assistant:"
response = generate_response(prompt, model)
# Print the response with formatted output
print(response)

<s> GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. can rain water reduce power of chemical after spraying?  <|end_of_turn|>  GPT4 Correct Assistant: No. rain water does not reduce the power of a chemical. If you are spraying chemical with rain in the air your chemical will reduce its efficiency.<|end_of_turn|>


### **Example No:4**

In [31]:
prompt = "GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. "
prompt += "What is the advantage of growing Maize? <|end_of_turn|> "
prompt += "GPT4 Correct Assistant:"
response = generate_response(prompt, model)
# Print the response with formatted output
print(response)

<s> GPT4 Correct User: Imagine you're an experienced agronomist, entrusted with optimizing crop yields. Your expertise is crucial in guiding farmers towards success. Picture a field filled with potential, and now it's your task to provide the necessary insights for cultivating a thriving harvest. What is the advantage of growing Maize? <|end_of_turn|>  GPT4 Correct Assistant: The advantage of growing Maize is that it is a very important crop for food and industrial use, it has good adaptation to harsh environment and high yield potential. It is important for livestock feed as it contains good nutritional value<|end_of_turn|>
