<a href="https://colab.research.google.com/github/cagBRT/promptEngineering/blob/main/3_HuggingFace.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>

In [None]:
# Transformers installation
! pip install transformers datasets
# To install from source instead of the last release, comment the command above and uncomment the following one.
! pip install git+https://github.com/huggingface/transformers.git

In [None]:
!pip install --upgrade accelerate
!pip install transformers accelerate

# Fine-tune a pretrained model
There are significant benefits to using a pretrained model:
 It reduces :
 - computation costs 
 - your carbon footprint
 - allows you to use state-of-the-art models without having to train one from scratch. <br>

HuggingFace Transformers provide access to thousands of pretrained models for a wide range of tasks. <br>



**When you use a pretrained model, you train it on a dataset specific to your task**. 
This is known as fine-tuning, an incredibly powerful training technique. <br>

In this tutorial, you will fine-tune a pretrained model with a deep learning framework of your choice:

- Fine-tune a pretrained model with HuggingFace Transformers Trainer.<br>
- Fine-tune a pretrained model in TensorFlow with Keras.<br>
- Fine-tune a pretrained model in native PyTorch.<br>

Before you can fine-tune a pretrained model, **download a dataset and prepare it for training**. 

Begin by loading the Yelp Reviews dataset:

In [None]:
from datasets import load_dataset

dataset = load_dataset("yelp_review_full")

**An example of one of the Yelp reviews**

In [None]:
dataset["train"][100]

**Tokenize the Yelp dataset**

In [None]:
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bert-base-cased")

def tokenize_function(examples):
    return tokenizer(examples["text"], padding="max_length", truncation=True)

tokenized_datasets = dataset.map(tokenize_function, batched=True)

If you like, you can create a smaller subset of the full dataset to fine-tune on to reduce the time it takes:

In [None]:
small_train_dataset = tokenized_datasets["train"].shuffle(seed=42).select(range(1000))
small_eval_dataset = tokenized_datasets["test"].shuffle(seed=42).select(range(1000))

# Train

HuggingFace Transformers provides a Trainer class optimized for training HuggingFace Transformers models, making it easier to start training without manually writing your own training loop. <br>

The Trainer API supports a range of training options and features such as logging, gradient accumulation, and mixed precision.<br>





Start by loading your model and specify the number of expected labels. From the Yelp Review dataset card, you know there are five labels (5 stars):

In [None]:
from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained("bert-base-cased", num_labels=5)

You will see a warning about some of the pretrained weights not being used and some weights being randomly initialized. Don't worry, this is completely normal! The pretrained head of the BERT model is discarded, and replaced with a randomly initialized classification head. You will fine-tune this new model head on your sequence classification task, transferring the knowledge of the pretrained model to it.

**Create a TrainingArguments class** which contains all the hyperparameters you can tune as well as flags for activating different training options. <br>

For this tutorial you can start with the default training hyperparameters, but feel free to experiment with these to find your optimal settings.

Specify where to save the checkpoints from your training:

In [None]:
!pip install git+https://github.com/huggingface/accelerate

In [None]:
from transformers import TrainingArguments

training_args = TrainingArguments(output_dir="test_trainer")

# Metrics
Trainer does not automatically evaluate model performance during training. <br>
You will need to pass Trainer a function to compute and report metrics. 

The HuggingFace Datasets library provides a simple accuracy function you can load with the load_metric function:

In [None]:
import numpy as np
from datasets import load_metric

metric = load_metric("accuracy")

**Call compute on metric to calculate the accuracy of your predictions**. <br>

Before passing your predictions to compute, you need to convert the predictions to logits (remember all HuggingFace Transformers models return logits):<br>
<br>

A Logit function, also known as the log-odds function, is a function that represents probability values from 0 to 1, and negative infinity to infinity. 


In [None]:
def compute_metrics(eval_pred):
    logits, labels = eval_pred
    predictions = np.argmax(logits, axis=-1)
    return metric.compute(predictions=predictions, references=labels)

If you'd like to monitor your evaluation metrics during fine-tuning,<br>

>specify the evaluation_strategy parameter in your training arguments to report the evaluation metric at the end of each epoch:



In [None]:
from transformers import TrainingArguments, Trainer

training_args = TrainingArguments(output_dir="test_trainer", evaluation_strategy="epoch")

**Trainer**<br>
Create a Trainer object with your model, training arguments, training and test datasets, and evaluation function:



In [None]:
trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=small_train_dataset,
    eval_dataset=small_eval_dataset,
    compute_metrics=compute_metrics,
)

Then fine-tune your model by calling train():



In [None]:
trainer.train()


In [None]:
trainer.save_model("trainer_saved")

HuggingFace Transformers models also supports training in TensorFlow with the Keras API.

# Convert dataset to TensorFlow format


The DefaultDataCollator assembles tensors into a batch for the model to train on. Make sure you specify return_tensors to return TensorFlow tensors:

In [None]:
from transformers import DefaultDataCollator

data_collator = DefaultDataCollator(return_tensors="tf")

Trainer uses DataCollatorWithPadding by default so you don't need to explicitly specify a data collator.


**Convert the tokenized datasets to TensorFlow datasets** with the to_tf_dataset method. Specify your inputs in columns, and your label in label_cols:

In [None]:
tf_train_dataset = small_train_dataset.to_tf_dataset(
    columns=["attention_mask", "input_ids", "token_type_ids"],
    label_cols=["labels"],
    shuffle=True,
    collate_fn=data_collator,
    batch_size=8,
)

tf_validation_dataset = small_eval_dataset.to_tf_dataset(
    columns=["attention_mask", "input_ids", "token_type_ids"],
    label_cols=["labels"],
    shuffle=False,
    collate_fn=data_collator,
    batch_size=8,
)

# Compile and fit
Load a TensorFlow model with the expected number of labels:



In [None]:
import tensorflow as tf
from transformers import TFAutoModelForSequenceClassification

model = TFAutoModelForSequenceClassification.from_pretrained
        ("bert-base-cased", num_labels=5)

Then compile and fine-tune your model with fit as you would with any other Keras model:

In [None]:
model.compile(
    optimizer=tf.keras.optimizers.Adam(learning_rate=5e-5),
    loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
    metrics=tf.metrics.SparseCategoricalAccuracy(),
)

model.fit(tf_train_dataset, validation_data=tf_validation_dataset, epochs=3)

# Train in native PyTorch


Trainer takes care of the training loop and allows you to fine-tune a model in a single line of code. 

For users who prefer to write their own training loop, you can also fine-tune a HuggingFace Transformers model in native PyTorch.

At this point, you may need to restart your notebook or execute the following code to free some memory:

In [None]:
del model
del pytorch_model
del trainer
torch.cuda.empty_cache()

Next, manually postprocess tokenized_dataset to prepare it for training.

Remove the text column because the model does not accept raw text as an input:

>>> tokenized_datasets = tokenized_datasets.remove_columns(["text"])
Rename the label column to labels because the model expects the argument to be named labels:

>>> tokenized_datasets = tokenized_datasets.rename_column("label", "labels")
Set the format of the dataset to return PyTorch tensors instead of lists:

>>> tokenized_datasets.set_format("torch")
Then create a smaller subset of the dataset as previously shown to speed up the fine-tuning:

In [None]:
small_train_dataset = tokenized_datasets["train"].shuffle(seed=42).select(range(1000))
small_eval_dataset = tokenized_datasets["test"].shuffle(seed=42).select(range(1000))
     

# DataLoader
Create a DataLoader for your training and test datasets so you can iterate over batches of data:

In [None]:

from torch.utils.data import DataLoader

train_dataloader = DataLoader(small_train_dataset, shuffle=True, batch_size=8)
eval_dataloader = DataLoader(small_eval_dataset, batch_size=8)


Load your model with the number of expected labels:

In [None]:
from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained("bert-base-cased", num_labels=5)

# Optimizer and learning rate scheduler
Create an optimizer and learning rate scheduler to fine-tune the model. Let's use the AdamW optimizer from PyTorch:

In [None]:
from torch.optim import AdamW

optimizer = AdamW(model.parameters(), lr=5e-5)

Create the default learning rate scheduler from Trainer:

In [None]:

from transformers import get_scheduler

num_epochs = 3
num_training_steps = num_epochs * len(train_dataloader)
lr_scheduler = get_scheduler(
    name="linear", optimizer=optimizer, num_warmup_steps=0, num_training_steps=num_training_steps
)

Lastly, specify device to use a GPU if you have access to one. Otherwise, training on a CPU may take several hours instead of a couple of minutes.



In [None]:
import torch

device = torch.device("cuda") if torch.cuda.is_available() else torch.device("cpu")
model.to(device)

Get free access to a cloud GPU if you don't have one with a hosted notebook like Colaboratory or SageMaker StudioLab.

Great, now you are ready to train! 🥳



# Training loop
To keep track of your training progress, use the tqdm library to add a progress bar over the number of training steps:

In [None]:

from tqdm.auto import tqdm

progress_bar = tqdm(range(num_training_steps))

model.train()
for epoch in range(num_epochs):
    for batch in train_dataloader:
        batch = {k: v.to(device) for k, v in batch.items()}
        outputs = model(**batch)
        loss = outputs.loss
        loss.backward()

        optimizer.step()
        lr_scheduler.step()
        optimizer.zero_grad()
        progress_bar.update(1)

# Metrics
Just like how you need to add an evaluation function to Trainer, you need to do the same when you write your own training loop. But instead of calculating and reporting the metric at the end of each epoch, this time you will accumulate all the batches with add_batch and calculate the metric at the very end.

In [None]:
metric = load_metric("accuracy")
model.eval()
for batch in eval_dataloader:
    batch = {k: v.to(device) for k, v in batch.items()}
    with torch.no_grad():
        outputs = model(**batch)

    logits = outputs.logits
    predictions = torch.argmax(logits, dim=-1)
    metric.add_batch(predictions=predictions, references=batch["labels"])

metric.compute()