# Text classification

Text classification is a common NLP task that assigns a label or class to text. Some of the largest companies run text classification in production for a wide range of practical applications. One of the most popular forms of text classification is sentiment analysis, which assigns a label like 🙂 positive, 🙁 negative, or 😐 neutral to a sequence of text.

This guide will show you how to:

1. Finetune [DistilBERT](https://huggingface.co/distilbert-base-uncased) on the [IMDb](https://huggingface.co/datasets/imdb) dataset to determine whether a movie review is positive or negative.
2. Use your finetuned model for inference.

<Tip>
The task illustrated in this tutorial is supported by the following model architectures:

<!--This tip is automatically generated by `make fix-copies`, do not fill manually!-->

[ALBERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/albert), [BART](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/bart), [BERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/bert), [BigBird](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/big_bird), [BigBird-Pegasus](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/bigbird_pegasus), [BioGpt](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/biogpt), [BLOOM](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/bloom), [CamemBERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/camembert), [CANINE](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/canine), [ConvBERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/convbert), [CTRL](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/ctrl), [Data2VecText](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/data2vec-text), [DeBERTa](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/deberta), [DeBERTa-v2](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/deberta-v2), [DistilBERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/distilbert), [ELECTRA](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/electra), [ERNIE](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/ernie), [ErnieM](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/ernie_m), [ESM](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/esm), [FlauBERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/flaubert), [FNet](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/fnet), [Funnel Transformer](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/funnel), [GPT-Sw3](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/gpt-sw3), [OpenAI GPT-2](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/gpt2), [GPTBigCode](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/gpt_bigcode), [GPT Neo](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/gpt_neo), [GPT NeoX](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/gpt_neox), [GPT-J](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/gptj), [I-BERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/ibert), [LayoutLM](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/layoutlm), [LayoutLMv2](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/layoutlmv2), [LayoutLMv3](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/layoutlmv3), [LED](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/led), [LiLT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/lilt), [LLaMA](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/llama), [Longformer](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/longformer), [LUKE](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/luke), [MarkupLM](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/markuplm), [mBART](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/mbart), [MEGA](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/mega), [Megatron-BERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/megatron-bert), [MobileBERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/mobilebert), [MPNet](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/mpnet), [MVP](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/mvp), [Nezha](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/nezha), [Nyströmformer](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/nystromformer), [OpenLlama](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/open-llama), [OpenAI GPT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/openai-gpt), [OPT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/opt), [Perceiver](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/perceiver), [PLBart](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/plbart), [QDQBert](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/qdqbert), [Reformer](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/reformer), [RemBERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/rembert), [RoBERTa](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/roberta), [RoBERTa-PreLayerNorm](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/roberta-prelayernorm), [RoCBert](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/roc_bert), [RoFormer](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/roformer), [SqueezeBERT](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/squeezebert), [TAPAS](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/tapas), [Transformer-XL](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/transfo-xl), [XLM](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/xlm), [XLM-RoBERTa](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/xlm-roberta), [XLM-RoBERTa-XL](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/xlm-roberta-xl), [XLNet](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/xlnet), [X-MOD](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/xmod), [YOSO](https://huggingface.co/docs/transformers/main/en/tasks/../model_doc/yoso)


<!--End of the generated tip-->

</Tip>

Before you begin, make sure you have all the necessary libraries installed:

```bash
pip install transformers datasets evaluate
```

We encourage you to login to your Hugging Face account so you can upload and share your model with the community. When prompted, enter your token to login:

In [1]:
from huggingface_hub import notebook_login

notebook_login()

#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

VBox(children=(HTML(value='<center> <img\nsrc=https://huggingface.co/front/assets/huggingface_logo-noborder.sv…

## Load IMDb dataset

Start by loading the IMDb dataset from the 🤗 Datasets library:

In [2]:
from datasets import load_dataset

imdb = load_dataset("imdb")
sst2 = load_dataset("SetFit/sst2")
sst5 = load_dataset("SetFit/sst5")
twitterfin = load_dataset("zeroshot/twitter-financial-news-sentiment")
agnews = load_dataset("ag_news")


Repo card metadata block was not found. Setting CardData to empty.
Repo card metadata block was not found. Setting CardData to empty.


In [3]:
sst2['test'][0]


{'text': 'no movement , no yuks , not much of anything .',
 'label': 0,
 'label_text': 'negative'}

## Preprocess

The next step is to load a DistilBERT tokenizer to preprocess the `text` field:

Create a preprocessing function to tokenize `text` and truncate sequences to be no longer than DistilBERT's maximum input length:

In [4]:
# from transformers import AutoTokenizer
# tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased")


# def preprocess_function(examples):
#     return tokenizer(examples.get('text', examples.get('sentence', '')), truncation=True)

# Add different paddings

In [4]:
def add_text(example, n):
  # example is a dictionary with keys 'text' and 'label'
  # we can modify the 'text' value and return the modified example
  text = example.get('text', example.get('sentence', ''))
  example['text'] = "[PAD]" * n + text
  return example

# for n in range(0, 101, 10):
#     # use format method to create a variable name with n
#     var_name = "imdb_padding_{}".format(n)
#     # use exec function to execute the assignment statement
#     exec("{} = imdb.map(lambda x: add_text(x, n))".format(var_name))
#     print(var_name, len(eval(var_name)))

datasets = {'imdb': imdb, 'sst2': sst2, 'sst5': sst5, 'twitterfin': twitterfin, 'agnews': agnews}

# 对每个数据集进行处理
for dataset_name, dataset in datasets.items():
    for n in range(0, 101, 10):
        var_name = "{}_padding_{}".format(dataset_name, n)
        exec("{} = dataset.map(lambda x: add_text(x, n))".format(var_name))
        print(var_name, len(eval(var_name)))

imdb_padding_0 3
imdb_padding_10 3
imdb_padding_20 3
imdb_padding_30 3
imdb_padding_40 3
imdb_padding_50 3
imdb_padding_60 3
imdb_padding_70 3
imdb_padding_80 3
imdb_padding_90 3
imdb_padding_100 3
sst2_padding_0 3
sst2_padding_10 3
sst2_padding_20 3
sst2_padding_30 3
sst2_padding_40 3
sst2_padding_50 3
sst2_padding_60 3
sst2_padding_70 3
sst2_padding_80 3
sst2_padding_90 3
sst2_padding_100 3
sst5_padding_0 3
sst5_padding_10 3
sst5_padding_20 3
sst5_padding_30 3
sst5_padding_40 3
sst5_padding_50 3
sst5_padding_60 3
sst5_padding_70 3
sst5_padding_80 3
sst5_padding_90 3
sst5_padding_100 3
twitterfin_padding_0 2
twitterfin_padding_10 2
twitterfin_padding_20 2
twitterfin_padding_30 2
twitterfin_padding_40 2
twitterfin_padding_50 2
twitterfin_padding_60 2
twitterfin_padding_70 2
twitterfin_padding_80 2
twitterfin_padding_90 2
twitterfin_padding_100 2
agnews_padding_0 2
agnews_padding_10 2
agnews_padding_20 2
agnews_padding_30 2
agnews_padding_40 2
agnews_padding_50 2
agnews_padding_60 2
agn

In [5]:
sst2_padding_10["test"][0]

{'text': '[PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD]no movement , no yuks , not much of anything .',
 'label': 0,
 'label_text': 'negative'}

To apply the preprocessing function over the entire dataset, use 🤗 Datasets [map](https://huggingface.co/docs/datasets/main/en/package_reference/main_classes#datasets.Dataset.map) function. You can speed up `map` by setting `batched=True` to process multiple elements of the dataset at once:

In [6]:
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased")

def preprocess_function(examples):
    return tokenizer(examples["text"], truncation=True)

In [7]:
n_range = range(0, 101, 10)


# 对每个数据集进行处理
for dataset_name, dataset in datasets.items():
    for n in n_range:
        filename = f"{dataset_name}_padding_{n}"
        file = eval(filename)
        tokenized_filename = f"tokenized_{dataset_name}_{n}"
        exec(f"{tokenized_filename} = file.map(preprocess_function, batched=True)")
        tokenized_file = eval(tokenized_filename)
        print(tokenized_filename, len(tokenized_file))
    print(tokenized_filename, n)
print("tokenization done")


tokenized_imdb_0 3


Map:   0%|          | 0/25000 [00:00<?, ? examples/s]

tokenized_imdb_10 3
tokenized_imdb_20 3
tokenized_imdb_30 3
tokenized_imdb_40 3
tokenized_imdb_50 3
tokenized_imdb_60 3
tokenized_imdb_70 3
tokenized_imdb_80 3
tokenized_imdb_90 3
tokenized_imdb_100 3
tokenized_imdb_100 100
tokenized_sst2_0 3
tokenized_sst2_10 3
tokenized_sst2_20 3
tokenized_sst2_30 3
tokenized_sst2_40 3
tokenized_sst2_50 3
tokenized_sst2_60 3
tokenized_sst2_70 3
tokenized_sst2_80 3
tokenized_sst2_90 3
tokenized_sst2_100 3
tokenized_sst2_100 100
tokenized_sst5_0 3
tokenized_sst5_10 3
tokenized_sst5_20 3
tokenized_sst5_30 3
tokenized_sst5_40 3
tokenized_sst5_50 3
tokenized_sst5_60 3
tokenized_sst5_70 3
tokenized_sst5_80 3
tokenized_sst5_90 3
tokenized_sst5_100 3
tokenized_sst5_100 100
tokenized_twitterfin_0 2
tokenized_twitterfin_10 2
tokenized_twitterfin_20 2
tokenized_twitterfin_30 2
tokenized_twitterfin_40 2
tokenized_twitterfin_50 2
tokenized_twitterfin_60 2
tokenized_twitterfin_70 2
tokenized_twitterfin_80 2
tokenized_twitterfin_90 2
tokenized_twitterfin_100 2
toke

In [8]:
tokenized_agnews_100["test"][0]

{'text': "[PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD][PAD]Fears for T N pension after talks Unions representing workers at Turner   Newall say they are 'disappointed' after talks with stricken parent firm Federal Mogul.",
 'label': 2,
 'input_ids': [101,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0,
  0

Now create a batch of examples using [DataCollatorWithPadding](https://huggingface.co/docs/transformers/main/en/main_classes/data_collator#transformers.DataCollatorWithPadding). It's more efficient to *dynamically pad* the sentences to the longest length in a batch during collation, instead of padding the whole dataset to the maximum length.

In [9]:
from transformers import DataCollatorWithPadding
data_collator = DataCollatorWithPadding(tokenizer=tokenizer)

## Evaluate

Including a metric during training is often helpful for evaluating your model's performance. You can quickly load a evaluation method with the 🤗 [Evaluate](https://huggingface.co/docs/evaluate/index) library. For this task, load the [accuracy](https://huggingface.co/spaces/evaluate-metric/accuracy) metric (see the 🤗 Evaluate [quick tour](https://huggingface.co/docs/evaluate/a_quick_tour) to learn more about how to load and compute a metric):

In [10]:
import evaluate
accuracy = evaluate.load("accuracy")

Then create a function that passes your predictions and labels to [compute](https://huggingface.co/docs/evaluate/main/en/package_reference/main_classes#evaluate.EvaluationModule.compute) to calculate the accuracy:

In [11]:
import numpy as np
def compute_metrics(eval_pred):
    predictions, labels = eval_pred
    predictions = np.argmax(predictions, axis=1)
    return accuracy.compute(predictions=predictions, references=labels)


Your `compute_metrics` function is ready to go now, and you'll return to it when you setup your training.

## Train

Before you start training your model, create a map of the expected ids to their labels with `id2label` and `label2id`:

In [None]:
# imdb
id2label = {0: "NEGATIVE", 1: "POSITIVE"}
label2id = {"NEGATIVE": 0, "POSITIVE": 1}

from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer

model = AutoModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased", num_labels=2, id2label=id2label, label2id=label2id
)

# define the range of n

for n in n_range:
    # use format method to create a variable name with n
    variable = f"tokenized_imdb_{n}"
    output_dir = f"left_padding{n}model"
    logging_dir=f"left_padding{n}model_logs"

    # use exec function to execute the assignment statement
    var = eval(variable)
    training_args = TrainingArguments(
        
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=16,
        num_train_epochs=10,
        weight_decay=0.01,
        
        evaluation_strategy="epoch",
        save_strategy="epoch",

        push_to_hub=True,
        output_dir=output_dir,
        seed=42,
        data_seed=123,
    )
    
    # define the trainer
    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=var["train"],
        eval_dataset=var["test"],
        tokenizer=tokenizer,
        data_collator=data_collator,
        compute_metrics=compute_metrics,
    )

    trainer.train(resume_from_checkpoint=True)
    trainer.push_to_hub()

In [12]:
# sst2: https://huggingface.co/datasets/SetFit/sst2

id2label = {0: "NEGATIVE", 1: "POSITIVE"}
label2id = {"NEGATIVE": 0, "POSITIVE": 1}

from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer

model = AutoModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased", num_labels=2, id2label=id2label, label2id=label2id
)

# define the range of n

for n in n_range:
    # use format method to create a variable name with n
    variable = f"tokenized_sst2_{n}"
    output_dir = f"sst2_padding{n}model"
    logging_dir=f"sst2_padding{n}model_logs"

    # use exec function to execute the assignment statement
    var = eval(variable)
    training_args = TrainingArguments(
        
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=16,
        num_train_epochs=0.01,
        weight_decay=0.01,
        
        evaluation_strategy="epoch",
        save_strategy="epoch",

        push_to_hub=True,
        output_dir=output_dir,
        seed=42,
        data_seed=123,
    )
    
    # define the trainer
    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=var["train"],
        eval_dataset=var["test"],
        tokenizer=tokenizer,
        data_collator=data_collator,
        compute_metrics=compute_metrics,
    )

    #trainer.train(resume_from_checkpoint=True)
    trainer.train()
    trainer.push_to_hub()

Some weights of DistilBertForSequenceClassification were not initialized from the model checkpoint at distilbert-base-uncased and are newly initialized: ['classifier.weight', 'pre_classifier.bias', 'pre_classifier.weight', 'classifier.bias']
You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.


  0%|          | 0/5 [00:00<?, ?it/s]

You're using a DistilBertTokenizerFast tokenizer. Please note that with a fast tokenizer, using the `__call__` method is faster than using a method to encode the text followed by a call to the `pad` method to get a padded encoding.


  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 0.6953288316726685, 'eval_accuracy': 0.49917627677100496, 'eval_runtime': 2.4475, 'eval_samples_per_second': 744.038, 'eval_steps_per_second': 46.579, 'epoch': 0.01}
{'train_runtime': 9.0773, 'train_samples_per_second': 7.623, 'train_steps_per_second': 0.551, 'train_loss': 0.6956593990325928, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 0.6975795030593872, 'eval_accuracy': 0.4975288303130148, 'eval_runtime': 2.9172, 'eval_samples_per_second': 624.219, 'eval_steps_per_second': 39.078, 'epoch': 0.01}
{'train_runtime': 6.8435, 'train_samples_per_second': 10.112, 'train_steps_per_second': 0.731, 'train_loss': 0.6858995914459228, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 0.6994195580482483, 'eval_accuracy': 0.48105436573311366, 'eval_runtime': 3.3185, 'eval_samples_per_second': 548.744, 'eval_steps_per_second': 34.353, 'epoch': 0.01}
{'train_runtime': 7.237, 'train_samples_per_second': 9.562, 'train_steps_per_second': 0.691, 'train_loss': 0.6899939060211182, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 0.7210732698440552, 'eval_accuracy': 0.49697968149368477, 'eval_runtime': 3.7304, 'eval_samples_per_second': 488.151, 'eval_steps_per_second': 30.56, 'epoch': 0.01}
{'train_runtime': 8.7625, 'train_samples_per_second': 7.897, 'train_steps_per_second': 0.571, 'train_loss': 0.68722562789917, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 0.7228018641471863, 'eval_accuracy': 0.4673256452498627, 'eval_runtime': 4.2282, 'eval_samples_per_second': 430.681, 'eval_steps_per_second': 26.962, 'epoch': 0.01}
{'train_runtime': 8.2898, 'train_samples_per_second': 8.348, 'train_steps_per_second': 0.603, 'train_loss': 0.6524814128875732, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 0.7667254209518433, 'eval_accuracy': 0.5057660626029654, 'eval_runtime': 4.6891, 'eval_samples_per_second': 388.347, 'eval_steps_per_second': 24.312, 'epoch': 0.01}
{'train_runtime': 9.4305, 'train_samples_per_second': 7.338, 'train_steps_per_second': 0.53, 'train_loss': 0.684022855758667, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 0.7972058057785034, 'eval_accuracy': 0.4733662822624931, 'eval_runtime': 5.3893, 'eval_samples_per_second': 337.889, 'eval_steps_per_second': 21.153, 'epoch': 0.01}
{'train_runtime': 10.1701, 'train_samples_per_second': 6.804, 'train_steps_per_second': 0.492, 'train_loss': 0.6042176246643066, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 0.8633940815925598, 'eval_accuracy': 0.49368478857770454, 'eval_runtime': 5.853, 'eval_samples_per_second': 311.122, 'eval_steps_per_second': 19.477, 'epoch': 0.01}
{'train_runtime': 9.4219, 'train_samples_per_second': 7.345, 'train_steps_per_second': 0.531, 'train_loss': 0.6667703628540039, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 0.9500973224639893, 'eval_accuracy': 0.4958813838550247, 'eval_runtime': 6.2933, 'eval_samples_per_second': 289.356, 'eval_steps_per_second': 18.115, 'epoch': 0.01}
{'train_runtime': 11.707, 'train_samples_per_second': 5.911, 'train_steps_per_second': 0.427, 'train_loss': 0.6533186912536622, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 1.0395991802215576, 'eval_accuracy': 0.4876441515650741, 'eval_runtime': 6.8627, 'eval_samples_per_second': 265.347, 'eval_steps_per_second': 16.611, 'epoch': 0.01}
{'train_runtime': 11.4785, 'train_samples_per_second': 6.029, 'train_steps_per_second': 0.436, 'train_loss': 0.6353700160980225, 'epoch': 0.01}


  0%|          | 0/5 [00:00<?, ?it/s]

  0%|          | 0/114 [00:00<?, ?it/s]

{'eval_loss': 1.1704916954040527, 'eval_accuracy': 0.46787479406919275, 'eval_runtime': 7.3234, 'eval_samples_per_second': 248.656, 'eval_steps_per_second': 15.567, 'epoch': 0.01}
{'train_runtime': 11.4769, 'train_samples_per_second': 6.029, 'train_steps_per_second': 0.436, 'train_loss': 0.6735086441040039, 'epoch': 0.01}


In [13]:
# sst5: https://huggingface.co/datasets/SetFit/sst5

id2label = {4: "very positive", 3: "positive", 2: "neutral", 1: "negative", 0: "very negative"}
label2id = {"very positive": 4, "positive": 3, "neutral": 2, "negative": 1, "very negative": 0}

from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer

model = AutoModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased", num_labels=5, id2label=id2label, label2id=label2id
)

# define the range of n

for n in n_range:
    # use format method to create a variable name with n
    variable = f"tokenized_sst5_{n}"
    output_dir = f"sst5_padding{n}model"
    logging_dir=f"sst5_padding{n}model_logs"

    # use exec function to execute the assignment statement
    var = eval(variable)
    training_args = TrainingArguments(
        
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=16,
        num_train_epochs=0.01,
        weight_decay=0.01,
        
        evaluation_strategy="epoch",
        save_strategy="epoch",

        push_to_hub=True,
        output_dir=output_dir,
        seed=42,
        data_seed=123,
    )
    
    # define the trainer
    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=var["train"],
        eval_dataset=var["test"],
        tokenizer=tokenizer,
        data_collator=data_collator,
        compute_metrics=compute_metrics,
    )
    
    #trainer.train(resume_from_checkpoint=True)
    trainer.train()

    trainer.push_to_hub()

Some weights of DistilBertForSequenceClassification were not initialized from the model checkpoint at distilbert-base-uncased and are newly initialized: ['classifier.weight', 'pre_classifier.bias', 'pre_classifier.weight', 'classifier.bias']
You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.6022824048995972, 'eval_accuracy': 0.23846153846153847, 'eval_runtime': 3.018, 'eval_samples_per_second': 732.282, 'eval_steps_per_second': 46.058, 'epoch': 0.01}
{'train_runtime': 7.8053, 'train_samples_per_second': 10.946, 'train_steps_per_second': 0.769, 'train_loss': 1.6150665283203125, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.5842775106430054, 'eval_accuracy': 0.28009049773755657, 'eval_runtime': 3.6487, 'eval_samples_per_second': 605.695, 'eval_steps_per_second': 38.096, 'epoch': 0.01}
{'train_runtime': 8.2315, 'train_samples_per_second': 10.38, 'train_steps_per_second': 0.729, 'train_loss': 1.5888927777608235, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.5739518404006958, 'eval_accuracy': 0.28778280542986423, 'eval_runtime': 4.1866, 'eval_samples_per_second': 527.878, 'eval_steps_per_second': 33.201, 'epoch': 0.01}
{'train_runtime': 9.4989, 'train_samples_per_second': 8.995, 'train_steps_per_second': 0.632, 'train_loss': 1.5661211013793945, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.565750241279602, 'eval_accuracy': 0.2895927601809955, 'eval_runtime': 4.8345, 'eval_samples_per_second': 457.128, 'eval_steps_per_second': 28.751, 'epoch': 0.01}
{'train_runtime': 8.94, 'train_samples_per_second': 9.557, 'train_steps_per_second': 0.671, 'train_loss': 1.5298943519592285, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.5655057430267334, 'eval_accuracy': 0.28778280542986423, 'eval_runtime': 5.4461, 'eval_samples_per_second': 405.793, 'eval_steps_per_second': 25.523, 'epoch': 0.01}
{'train_runtime': 10.11, 'train_samples_per_second': 8.451, 'train_steps_per_second': 0.593, 'train_loss': 1.5098872184753418, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.5651577711105347, 'eval_accuracy': 0.2900452488687783, 'eval_runtime': 6.2195, 'eval_samples_per_second': 355.335, 'eval_steps_per_second': 22.349, 'epoch': 0.01}
{'train_runtime': 11.9717, 'train_samples_per_second': 7.137, 'train_steps_per_second': 0.501, 'train_loss': 1.5065964063008626, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.5560145378112793, 'eval_accuracy': 0.302262443438914, 'eval_runtime': 6.4332, 'eval_samples_per_second': 343.531, 'eval_steps_per_second': 21.607, 'epoch': 0.01}
{'train_runtime': 11.6322, 'train_samples_per_second': 7.345, 'train_steps_per_second': 0.516, 'train_loss': 1.4822359085083008, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.5509940385818481, 'eval_accuracy': 0.3389140271493213, 'eval_runtime': 7.2152, 'eval_samples_per_second': 306.3, 'eval_steps_per_second': 19.265, 'epoch': 0.01}
{'train_runtime': 11.4715, 'train_samples_per_second': 7.448, 'train_steps_per_second': 0.523, 'train_loss': 1.445456822713216, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.5507333278656006, 'eval_accuracy': 0.33710407239819007, 'eval_runtime': 7.5398, 'eval_samples_per_second': 293.112, 'eval_steps_per_second': 18.436, 'epoch': 0.01}
{'train_runtime': 13.7765, 'train_samples_per_second': 6.202, 'train_steps_per_second': 0.436, 'train_loss': 1.4012468655904133, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.5402005910873413, 'eval_accuracy': 0.3393665158371041, 'eval_runtime': 8.5474, 'eval_samples_per_second': 258.559, 'eval_steps_per_second': 16.262, 'epoch': 0.01}
{'train_runtime': 15.2755, 'train_samples_per_second': 5.593, 'train_steps_per_second': 0.393, 'train_loss': 1.3631470998128254, 'epoch': 0.01}


  0%|          | 0/6 [00:00<?, ?it/s]

  0%|          | 0/139 [00:00<?, ?it/s]

{'eval_loss': 1.5398588180541992, 'eval_accuracy': 0.3375565610859729, 'eval_runtime': 9.7149, 'eval_samples_per_second': 227.485, 'eval_steps_per_second': 14.308, 'epoch': 0.01}
{'train_runtime': 15.623, 'train_samples_per_second': 5.469, 'train_steps_per_second': 0.384, 'train_loss': 1.3308571179707844, 'epoch': 0.01}


In [None]:
#twitterfin: https://huggingface.co/datasets/zeroshot/twitter-financial-news-sentiment

id2label = {0: "Bearish", 1: "Bullish", 2: "Neutral"}
label2id = {"Bearish": 0, "Bullish": 1, "Neutral": 2}

from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer

model = AutoModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased", num_labels=3, id2label=id2label, label2id=label2id
)

# define the range of n
for n in n_range:
    # use format method to create a variable name with n
    variable = f"tokenized_twitterfin_{n}"
    output_dir = f"twitterfin_padding{n}model"
    logging_dir=f"twitterfin_padding{n}model_logs"

    # use exec function to execute the assignment statement
    var = eval(variable)
    training_args = TrainingArguments(
        
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=16,
        num_train_epochs=0.01,
        weight_decay=0.01,
        
        evaluation_strategy="epoch",
        save_strategy="epoch",

        push_to_hub=True,
        output_dir=output_dir,
        seed=42,
        data_seed=123,
    )
    
    # define the trainer
    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=var["train"],
        eval_dataset=var["validation"],
        tokenizer=tokenizer,
        data_collator=data_collator,
        compute_metrics=compute_metrics,
    )

    trainer.train()
    trainer.push_to_hub()

Some weights of DistilBertForSequenceClassification were not initialized from the model checkpoint at distilbert-base-uncased and are newly initialized: ['pre_classifier.weight', 'classifier.weight', 'classifier.bias', 'pre_classifier.bias']
You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.


  0%|          | 0/597 [00:00<?, ?it/s]

You're using a DistilBertTokenizerFast tokenizer. Please note that with a fast tokenizer, using the `__call__` method is faster than using a method to encode the text followed by a call to the `pad` method to get a padded encoding.


{'loss': 0.5663, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.4012547731399536, 'eval_accuracy': 0.8513400335008375, 'eval_runtime': 3.7193, 'eval_samples_per_second': 642.061, 'eval_steps_per_second': 40.33, 'epoch': 1.0}
{'train_runtime': 60.4189, 'train_samples_per_second': 157.947, 'train_steps_per_second': 9.881, 'train_loss': 0.5491338439123514, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.3845, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.34795135259628296, 'eval_accuracy': 0.8743718592964824, 'eval_runtime': 4.2634, 'eval_samples_per_second': 560.114, 'eval_steps_per_second': 35.183, 'epoch': 1.0}
{'train_runtime': 67.3437, 'train_samples_per_second': 141.706, 'train_steps_per_second': 8.865, 'train_loss': 0.38838475034065184, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.2897, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.35290542244911194, 'eval_accuracy': 0.86892797319933, 'eval_runtime': 4.9902, 'eval_samples_per_second': 478.534, 'eval_steps_per_second': 30.059, 'epoch': 1.0}
{'train_runtime': 70.6744, 'train_samples_per_second': 135.028, 'train_steps_per_second': 8.447, 'train_loss': 0.3037401004452602, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.2299, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.40427643060684204, 'eval_accuracy': 0.8639028475711893, 'eval_runtime': 5.6457, 'eval_samples_per_second': 422.979, 'eval_steps_per_second': 26.569, 'epoch': 1.0}
{'train_runtime': 79.4716, 'train_samples_per_second': 120.081, 'train_steps_per_second': 7.512, 'train_loss': 0.260514609178706, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.1885, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.4547550678253174, 'eval_accuracy': 0.8722780569514238, 'eval_runtime': 6.2581, 'eval_samples_per_second': 381.586, 'eval_steps_per_second': 23.969, 'epoch': 1.0}
{'train_runtime': 86.4189, 'train_samples_per_second': 110.427, 'train_steps_per_second': 6.908, 'train_loss': 0.23381399669040187, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.1626, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.5107681751251221, 'eval_accuracy': 0.8739530988274706, 'eval_runtime': 7.3353, 'eval_samples_per_second': 325.551, 'eval_steps_per_second': 20.449, 'epoch': 1.0}
{'train_runtime': 101.0037, 'train_samples_per_second': 94.482, 'train_steps_per_second': 5.911, 'train_loss': 0.21660090051903397, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.1487, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.583983838558197, 'eval_accuracy': 0.8739530988274706, 'eval_runtime': 8.3663, 'eval_samples_per_second': 285.432, 'eval_steps_per_second': 17.929, 'epoch': 1.0}
{'train_runtime': 112.0678, 'train_samples_per_second': 85.154, 'train_steps_per_second': 5.327, 'train_loss': 0.2060726747241252, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.1322, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.631811797618866, 'eval_accuracy': 0.8760469011725294, 'eval_runtime': 9.253, 'eval_samples_per_second': 258.079, 'eval_steps_per_second': 16.211, 'epoch': 1.0}
{'train_runtime': 120.3685, 'train_samples_per_second': 79.282, 'train_steps_per_second': 4.96, 'train_loss': 0.1944433943909816, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.1113, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.6689165830612183, 'eval_accuracy': 0.864321608040201, 'eval_runtime': 10.3076, 'eval_samples_per_second': 231.673, 'eval_steps_per_second': 14.552, 'epoch': 1.0}
{'train_runtime': 132.1456, 'train_samples_per_second': 72.216, 'train_steps_per_second': 4.518, 'train_loss': 0.17782326958686662, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.1083, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.7153224349021912, 'eval_accuracy': 0.8706030150753769, 'eval_runtime': 10.8473, 'eval_samples_per_second': 220.146, 'eval_steps_per_second': 13.828, 'epoch': 1.0}
{'train_runtime': 138.3942, 'train_samples_per_second': 68.955, 'train_steps_per_second': 4.314, 'train_loss': 0.18571021049665645, 'epoch': 1.0}


  0%|          | 0/597 [00:00<?, ?it/s]

{'loss': 0.0944, 'learning_rate': 3.2495812395309884e-06, 'epoch': 0.84}


  0%|          | 0/150 [00:00<?, ?it/s]

{'eval_loss': 0.7331303954124451, 'eval_accuracy': 0.8697654941373534, 'eval_runtime': 10.2221, 'eval_samples_per_second': 233.611, 'eval_steps_per_second': 14.674, 'epoch': 1.0}
{'train_runtime': 138.2947, 'train_samples_per_second': 69.005, 'train_steps_per_second': 4.317, 'train_loss': 0.16383927950707314, 'epoch': 1.0}


In [14]:
# agnews: https://huggingface.co/datasets/ag_news

id2label = {0: "World", 1: "Sports", 2: "Business", 3: "Sci/Tech"}
label2id = {"World": 0, "Sports": 1, "Business": 2, "Sci/Tech": 3}

from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer

model = AutoModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased", num_labels=4, id2label=id2label, label2id=label2id
)

# define the range of n
for n in n_range:
    # use format method to create a variable name with n
    variable = f"tokenized_agnews_{n}"
    output_dir = f"agnews_padding{n}model"
    logging_dir=f"agnews_padding{n}model_logs"

    # use exec function to execute the assignment statement
    var = eval(variable)
    training_args = TrainingArguments(
        
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=16,
        num_train_epochs=0.01,
        weight_decay=0.01,
        
        evaluation_strategy="epoch",
        save_strategy="epoch",

        push_to_hub=True,
        output_dir=output_dir,
        seed=42,
        data_seed=123,
    )
    
    # define the trainer
    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=var["train"],
        eval_dataset=var["test"],
        tokenizer=tokenizer,
        data_collator=data_collator,
        compute_metrics=compute_metrics,
    )

    #trainer.train(resume_from_checkpoint=True)
    trainer.train()
    trainer.push_to_hub()


Some weights of DistilBertForSequenceClassification were not initialized from the model checkpoint at distilbert-base-uncased and are newly initialized: ['classifier.weight', 'pre_classifier.bias', 'pre_classifier.weight', 'classifier.bias']
You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.7518361806869507, 'eval_accuracy': 0.8542105263157894, 'eval_runtime': 20.1254, 'eval_samples_per_second': 377.632, 'eval_steps_per_second': 23.602, 'epoch': 0.01}
{'train_runtime': 39.5286, 'train_samples_per_second': 30.358, 'train_steps_per_second': 1.897, 'train_loss': 1.0339847819010417, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.4200113117694855, 'eval_accuracy': 0.8769736842105263, 'eval_runtime': 21.9953, 'eval_samples_per_second': 345.528, 'eval_steps_per_second': 21.595, 'epoch': 0.01}
{'train_runtime': 45.5323, 'train_samples_per_second': 26.355, 'train_steps_per_second': 1.647, 'train_loss': 0.5153947448730469, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.35912051796913147, 'eval_accuracy': 0.8864473684210527, 'eval_runtime': 24.1721, 'eval_samples_per_second': 314.412, 'eval_steps_per_second': 19.651, 'epoch': 0.01}
{'train_runtime': 47.784, 'train_samples_per_second': 25.113, 'train_steps_per_second': 1.57, 'train_loss': 0.36050994873046877, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.3472963571548462, 'eval_accuracy': 0.89, 'eval_runtime': 26.508, 'eval_samples_per_second': 286.706, 'eval_steps_per_second': 17.919, 'epoch': 0.01}
{'train_runtime': 52.6805, 'train_samples_per_second': 22.779, 'train_steps_per_second': 1.424, 'train_loss': 0.2943467458089193, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.3629646897315979, 'eval_accuracy': 0.8922368421052631, 'eval_runtime': 29.5337, 'eval_samples_per_second': 257.333, 'eval_steps_per_second': 16.083, 'epoch': 0.01}
{'train_runtime': 56.0654, 'train_samples_per_second': 21.404, 'train_steps_per_second': 1.338, 'train_loss': 0.24894274393717447, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.3956390917301178, 'eval_accuracy': 0.886578947368421, 'eval_runtime': 31.1604, 'eval_samples_per_second': 243.899, 'eval_steps_per_second': 15.244, 'epoch': 0.01}
{'train_runtime': 58.8163, 'train_samples_per_second': 20.403, 'train_steps_per_second': 1.275, 'train_loss': 0.20261423746744792, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.45044973492622375, 'eval_accuracy': 0.8881578947368421, 'eval_runtime': 34.3238, 'eval_samples_per_second': 221.421, 'eval_steps_per_second': 13.839, 'epoch': 0.01}
{'train_runtime': 65.8629, 'train_samples_per_second': 18.22, 'train_steps_per_second': 1.139, 'train_loss': 0.1922247568766276, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.4824838936328888, 'eval_accuracy': 0.8867105263157895, 'eval_runtime': 36.8704, 'eval_samples_per_second': 206.128, 'eval_steps_per_second': 12.883, 'epoch': 0.01}
{'train_runtime': 69.6692, 'train_samples_per_second': 17.224, 'train_steps_per_second': 1.077, 'train_loss': 0.17889537811279296, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.5680945515632629, 'eval_accuracy': 0.8846052631578948, 'eval_runtime': 38.4431, 'eval_samples_per_second': 197.695, 'eval_steps_per_second': 12.356, 'epoch': 0.01}
{'train_runtime': 74.0548, 'train_samples_per_second': 16.204, 'train_steps_per_second': 1.013, 'train_loss': 0.18295463562011718, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.6541604995727539, 'eval_accuracy': 0.8827631578947368, 'eval_runtime': 42.3707, 'eval_samples_per_second': 179.369, 'eval_steps_per_second': 11.211, 'epoch': 0.01}
{'train_runtime': 85.1027, 'train_samples_per_second': 14.101, 'train_steps_per_second': 0.881, 'train_loss': 0.19648821512858072, 'epoch': 0.01}


  0%|          | 0/75 [00:00<?, ?it/s]

  0%|          | 0/475 [00:00<?, ?it/s]

{'eval_loss': 0.7492882609367371, 'eval_accuracy': 0.8761842105263158, 'eval_runtime': 64.3874, 'eval_samples_per_second': 118.036, 'eval_steps_per_second': 7.377, 'epoch': 0.01}
{'train_runtime': 120.7853, 'train_samples_per_second': 9.935, 'train_steps_per_second': 0.621, 'train_loss': 0.19611497243245443, 'epoch': 0.01}


<Tip>

If you aren't familiar with finetuning a model with the [Trainer](https://huggingface.co/docs/transformers/main/en/main_classes/trainer#transformers.Trainer), take a look at the basic tutorial [here](https://huggingface.co/docs/transformers/main/en/tasks/../training#train-with-pytorch-trainer)!

</Tip>

You're ready to start training your model now! Load DistilBERT with [AutoModelForSequenceClassification](https://huggingface.co/docs/transformers/main/en/model_doc/auto#transformers.AutoModelForSequenceClassification) along with the number of expected labels, and the label mappings:

In [15]:
from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer

model = AutoModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased", num_labels=2, id2label=id2label, label2id=label2id
)

Some weights of DistilBertForSequenceClassification were not initialized from the model checkpoint at distilbert-base-uncased and are newly initialized: ['classifier.weight', 'pre_classifier.weight', 'classifier.bias', 'pre_classifier.bias']
You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.


In [None]:
training_args = TrainingArguments(
    output_dir="left_padding0model",
    learning_rate=2e-5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    num_train_epochs=2,
    weight_decay=0.01,

    evaluation_strategy="epoch",
    save_strategy="epoch",
    
    load_best_model_at_end=True,
    push_to_hub=True,
    seed=42,
    data_seed=123)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_imdb_0["train"],
    eval_dataset=tokenized_imdb_0["test"],
    tokenizer=tokenizer,
    data_collator=data_collator,
    compute_metrics=compute_metrics,

)

trainer.train(resume_from_checkpoint=True)


In [33]:
training_args = TrainingArguments(
    output_dir="left_padding0model",
    learning_rate=2e-5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    num_train_epochs=2,
    weight_decay=0.01,

    evaluation_strategy="epoch",
    save_strategy="epoch",
    
    load_best_model_at_end=True,
    push_to_hub=True,
    seed=42,
    data_seed=123)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_imdb_0["train"],
    eval_dataset=tokenized_imdb_0["test"],
    tokenizer=tokenizer,
    data_collator=data_collator,
    compute_metrics=compute_metrics,

)

trainer.train(resume_from_checkpoint=True)


  0%|          | 0/3126 [00:00<?, ?it/s]

{'train_runtime': 0.1256, 'train_samples_per_second': 398145.539, 'train_steps_per_second': 24892.059, 'train_loss': 0.0, 'epoch': 20.0}


TrainOutput(global_step=31260, training_loss=0.0, metrics={'train_runtime': 0.1256, 'train_samples_per_second': 398145.539, 'train_steps_per_second': 24892.059, 'train_loss': 0.0, 'epoch': 20.0})

In [20]:
import json

results = {}

for n in range(0, 11, 10):
    result = trainer.evaluate()
    results[n] = result


with open("results.json", "w") as f:
    json.dump(results, f)



  0%|          | 0/1563 [00:00<?, ?it/s]

  0%|          | 0/1563 [00:00<?, ?it/s]

NameError: name 'json' is not defined

In [24]:
results

{0: {'eval_loss': 0.23403853178024292,
  'eval_accuracy': 0.9166,
  'eval_runtime': 357.8435,
  'eval_samples_per_second': 69.863,
  'eval_steps_per_second': 4.368,
  'epoch': 2.0},
 10: {'eval_loss': 0.23403853178024292,
  'eval_accuracy': 0.9166,
  'eval_runtime': 364.259,
  'eval_samples_per_second': 68.632,
  'eval_steps_per_second': 4.291,
  'epoch': 2.0}}

In [28]:
# define the range of n
#for n in n_range:
for n in range(0, 11, 10):
    # use format method to create a variable name with n
    variable = f"tokenized_imdb_{n}"
    output_dir = f"left_padding{n}model"
    logging_dir=f"left_padding{n}model_logs"

    # use exec function to execute the assignment statement
    var = eval(variable)
    training_args = TrainingArguments(
        
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=16,
        num_train_epochs=2,
        weight_decay=0.01,
        
        evaluation_strategy="epoch",
        save_strategy="epoch",

        push_to_hub=True,
        seed=42,
        data_seed=123,
        output_dir=output_dir,
        logging_dir="./logs",
        logging_steps=100
    )
    
    # define the trainer
    trainer = Trainer(
        model=model,
        args=training_args,
        train_dataset=var["train"],
        eval_dataset=var["test"],
        tokenizer=tokenizer,
        data_collator=data_collator,
        compute_metrics=compute_metrics,
    )
    trainer.train()


  0%|          | 0/3126 [00:00<?, ?it/s]

{'loss': 0.1914, 'learning_rate': 1.9360204734484968e-05, 'epoch': 0.06}
{'loss': 0.1357, 'learning_rate': 1.872040946896993e-05, 'epoch': 0.13}
{'loss': 0.1454, 'learning_rate': 1.8080614203454897e-05, 'epoch': 0.19}
{'loss': 0.2537, 'learning_rate': 1.744081893793986e-05, 'epoch': 0.26}
{'loss': 0.2634, 'learning_rate': 1.6801023672424827e-05, 'epoch': 0.32}
{'loss': 0.2562, 'learning_rate': 1.616122840690979e-05, 'epoch': 0.38}
{'loss': 0.2484, 'learning_rate': 1.5521433141394756e-05, 'epoch': 0.45}
{'loss': 0.2546, 'learning_rate': 1.488163787587972e-05, 'epoch': 0.51}
{'loss': 0.2496, 'learning_rate': 1.4241842610364684e-05, 'epoch': 0.58}
{'loss': 0.2291, 'learning_rate': 1.3602047344849649e-05, 'epoch': 0.64}
{'loss': 0.2437, 'learning_rate': 1.2962252079334613e-05, 'epoch': 0.7}
{'loss': 0.2125, 'learning_rate': 1.2322456813819578e-05, 'epoch': 0.77}
{'loss': 0.2148, 'learning_rate': 1.1682661548304543e-05, 'epoch': 0.83}
{'loss': 0.2183, 'learning_rate': 1.1042866282789508e-05

  0%|          | 0/1563 [00:00<?, ?it/s]

{'eval_loss': 0.22205103933811188, 'eval_runtime': 363.4682, 'eval_samples_per_second': 68.782, 'eval_steps_per_second': 4.3, 'epoch': 1.0}
{'loss': 0.171, 'learning_rate': 9.763275751759437e-06, 'epoch': 1.02}
{'loss': 0.1388, 'learning_rate': 9.123480486244403e-06, 'epoch': 1.09}
{'loss': 0.1527, 'learning_rate': 8.483685220729368e-06, 'epoch': 1.15}
{'loss': 0.1351, 'learning_rate': 7.843889955214333e-06, 'epoch': 1.22}
{'loss': 0.144, 'learning_rate': 7.204094689699297e-06, 'epoch': 1.28}
{'loss': 0.1467, 'learning_rate': 6.5642994241842614e-06, 'epoch': 1.34}
{'loss': 0.1327, 'learning_rate': 5.924504158669226e-06, 'epoch': 1.41}
{'loss': 0.1555, 'learning_rate': 5.284708893154191e-06, 'epoch': 1.47}
{'loss': 0.1377, 'learning_rate': 4.644913627639156e-06, 'epoch': 1.54}
{'loss': 0.128, 'learning_rate': 4.005118362124121e-06, 'epoch': 1.6}
{'loss': 0.1279, 'learning_rate': 3.3653230966090854e-06, 'epoch': 1.66}
{'loss': 0.1799, 'learning_rate': 2.72552783109405e-06, 'epoch': 1.73}

  0%|          | 0/1563 [00:00<?, ?it/s]

{'eval_loss': 0.24687997996807098, 'eval_runtime': 362.9107, 'eval_samples_per_second': 68.887, 'eval_steps_per_second': 4.307, 'epoch': 2.0}
{'train_runtime': 2814.2898, 'train_samples_per_second': 17.766, 'train_steps_per_second': 1.111, 'train_loss': 0.18125215174674378, 'epoch': 2.0}


TypeError: list indices must be integers or slices, not str

KeyboardInterrupt: 

In [85]:
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tokenizer = AutoTokenizer.from_pretrained("Realgon/left_padding100_model")
model = AutoModelForSequenceClassification.from_pretrained("Realgon/left_padding100_model")

Downloading (…)okenizer_config.json:   0%|          | 0.00/1.20k [00:00<?, ?B/s]

Downloading (…)solve/main/vocab.txt:   0%|          | 0.00/232k [00:00<?, ?B/s]

Downloading (…)/main/tokenizer.json:   0%|          | 0.00/711k [00:00<?, ?B/s]

Downloading (…)cial_tokens_map.json:   0%|          | 0.00/125 [00:00<?, ?B/s]

Downloading (…)lve/main/config.json:   0%|          | 0.00/735 [00:00<?, ?B/s]

Downloading model.safetensors:   0%|          | 0.00/268M [00:00<?, ?B/s]

In [84]:
for obj in trainer.state.log_history:
    print(obj)

In [55]:
trainer.state.log_history['accuracy']

TypeError: list indices must be integers or slices, not str

In [37]:
result

[{'train_runtime': 2814.2898,
  'train_samples_per_second': 17.766,
  'train_steps_per_second': 1.111,
  'total_flos': 6564686875195392.0,
  'train_loss': 0.18125215174674378,
  'epoch': 2.0,
  'step': 3126,
  'n': 0}]

In [34]:
result['n'] = n
result_df = result_df.append(result, ignore_index=True)

TypeError: list indices must be integers or slices, not str

In [None]:
training_args = TrainingArguments(
    output_dir="left_padding50_model",
    learning_rate=2e-5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    num_train_epochs=2,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    push_to_hub=True,
    seed=42,
    data_seed=123
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_imdb_50["train"],
    eval_dataset=tokenized_imdb_50["test"],
    tokenizer=tokenizer,
    data_collator=data_collator,
    compute_metrics=compute_metrics,

)
trainer.train()

In [16]:
training_args = TrainingArguments(
    output_dir="left_padding50_model",
    learning_rate=2e-5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    num_train_epochs=2,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    push_to_hub=True,
    seed=42,
    data_seed=123
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_imdb_50["train"],
    eval_dataset=tokenized_imdb_50["test"],
    tokenizer=tokenizer,
    data_collator=data_collator,
    compute_metrics=compute_metrics,

)

trainer.train()

  0%|          | 0/3126 [00:00<?, ?it/s]

You're using a DistilBertTokenizerFast tokenizer. Please note that with a fast tokenizer, using the `__call__` method is faster than using a method to encode the text followed by a call to the `pad` method to get a padded encoding.


{'loss': 0.3348, 'learning_rate': 1.6801023672424827e-05, 'epoch': 0.32}
{'loss': 0.2586, 'learning_rate': 1.3602047344849649e-05, 'epoch': 0.64}
{'loss': 0.2306, 'learning_rate': 1.0403071017274472e-05, 'epoch': 0.96}


  0%|          | 0/1563 [00:00<?, ?it/s]

{'eval_loss': 0.26562219858169556, 'eval_accuracy': 0.90268, 'eval_runtime': 377.3093, 'eval_samples_per_second': 66.259, 'eval_steps_per_second': 4.142, 'epoch': 1.0}
{'loss': 0.1622, 'learning_rate': 7.204094689699297e-06, 'epoch': 1.28}
{'loss': 0.1487, 'learning_rate': 4.005118362124121e-06, 'epoch': 1.6}
{'loss': 0.1599, 'learning_rate': 8.061420345489445e-07, 'epoch': 1.92}


  0%|          | 0/1563 [00:00<?, ?it/s]

{'eval_loss': 0.235260471701622, 'eval_accuracy': 0.92952, 'eval_runtime': 376.2296, 'eval_samples_per_second': 66.449, 'eval_steps_per_second': 4.154, 'epoch': 2.0}
{'train_runtime': 2855.1642, 'train_samples_per_second': 17.512, 'train_steps_per_second': 1.095, 'train_loss': 0.2124700436436512, 'epoch': 2.0}


TrainOutput(global_step=3126, training_loss=0.2124700436436512, metrics={'train_runtime': 2855.1642, 'train_samples_per_second': 17.512, 'train_steps_per_second': 1.095, 'train_loss': 0.2124700436436512, 'epoch': 2.0})

In [17]:
trainer.push_to_hub()

'https://huggingface.co/Realgon/left_padding50_model/tree/main/'

In [18]:
training_args = TrainingArguments(
    output_dir="left_padding100_model",
    learning_rate=2e-5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    num_train_epochs=2,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    push_to_hub=True,
    seed=42,
    data_seed=123
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_imdb_100["train"],
    eval_dataset=tokenized_imdb_100["test"],
    tokenizer=tokenizer,
    data_collator=data_collator,
    compute_metrics=compute_metrics,

)

trainer.train()

  0%|          | 0/3126 [00:00<?, ?it/s]

{'loss': 0.1975, 'learning_rate': 1.6801023672424827e-05, 'epoch': 0.32}
{'loss': 0.1874, 'learning_rate': 1.3602047344849649e-05, 'epoch': 0.64}
{'loss': 0.1809, 'learning_rate': 1.0403071017274472e-05, 'epoch': 0.96}


  0%|          | 0/1563 [00:00<?, ?it/s]

{'eval_loss': 0.2602356970310211, 'eval_accuracy': 0.9228, 'eval_runtime': 373.1317, 'eval_samples_per_second': 67.0, 'eval_steps_per_second': 4.189, 'epoch': 1.0}
{'loss': 0.1007, 'learning_rate': 7.204094689699297e-06, 'epoch': 1.28}
{'loss': 0.0853, 'learning_rate': 4.005118362124121e-06, 'epoch': 1.6}
{'loss': 0.1099, 'learning_rate': 8.061420345489445e-07, 'epoch': 1.92}


  0%|          | 0/1563 [00:00<?, ?it/s]

{'eval_loss': 0.284171462059021, 'eval_accuracy': 0.93, 'eval_runtime': 372.6626, 'eval_samples_per_second': 67.085, 'eval_steps_per_second': 4.194, 'epoch': 2.0}
{'train_runtime': 2858.2215, 'train_samples_per_second': 17.493, 'train_steps_per_second': 1.094, 'train_loss': 0.1417671253646099, 'epoch': 2.0}


TrainOutput(global_step=3126, training_loss=0.1417671253646099, metrics={'train_runtime': 2858.2215, 'train_samples_per_second': 17.493, 'train_steps_per_second': 1.094, 'train_loss': 0.1417671253646099, 'epoch': 2.0})

In [19]:
trainer.push_to_hub()

'https://huggingface.co/Realgon/left_padding100_model/tree/main/'

At this point, only three steps remain:

1. Define your training hyperparameters in [TrainingArguments](https://huggingface.co/docs/transformers/main/en/main_classes/trainer#transformers.TrainingArguments). The only required parameter is `output_dir` which specifies where to save your model. You'll push this model to the Hub by setting `push_to_hub=True` (you need to be signed in to Hugging Face to upload your model). At the end of each epoch, the [Trainer](https://huggingface.co/docs/transformers/main/en/main_classes/trainer#transformers.Trainer) will evaluate the accuracy and save the training checkpoint.
2. Pass the training arguments to [Trainer](https://huggingface.co/docs/transformers/main/en/main_classes/trainer#transformers.Trainer) along with the model, dataset, tokenizer, data collator, and `compute_metrics` function.
3. Call [train()](https://huggingface.co/docs/transformers/main/en/main_classes/trainer#transformers.Trainer.train) to finetune your model.

In [20]:
training_args = TrainingArguments(
    output_dir="left_padding150_model",
    learning_rate=2e-5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    num_train_epochs=2,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    push_to_hub=True,
    seed=42,
    data_seed=123)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_imdb_150["train"],
    eval_dataset=tokenized_imdb_150["test"],
    tokenizer=tokenizer,
    data_collator=data_collator,
    compute_metrics=compute_metrics,

)

trainer.train()

  0%|          | 0/3126 [00:00<?, ?it/s]

{'loss': 0.1597, 'learning_rate': 1.6801023672424827e-05, 'epoch': 0.32}
{'loss': 0.1511, 'learning_rate': 1.3602047344849649e-05, 'epoch': 0.64}
{'loss': 0.1564, 'learning_rate': 1.0403071017274472e-05, 'epoch': 0.96}


  0%|          | 0/1563 [00:00<?, ?it/s]

{'eval_loss': 0.2768949866294861, 'eval_accuracy': 0.92112, 'eval_runtime': 379.0743, 'eval_samples_per_second': 65.95, 'eval_steps_per_second': 4.123, 'epoch': 1.0}
{'loss': 0.082, 'learning_rate': 7.204094689699297e-06, 'epoch': 1.28}
{'loss': 0.0764, 'learning_rate': 4.005118362124121e-06, 'epoch': 1.6}
{'loss': 0.0916, 'learning_rate': 8.061420345489445e-07, 'epoch': 1.92}


  0%|          | 0/1563 [00:00<?, ?it/s]

{'eval_loss': 0.3246845006942749, 'eval_accuracy': 0.92544, 'eval_runtime': 378.4581, 'eval_samples_per_second': 66.058, 'eval_steps_per_second': 4.13, 'epoch': 2.0}
{'train_runtime': 2899.2484, 'train_samples_per_second': 17.246, 'train_steps_per_second': 1.078, 'train_loss': 0.11752973331981031, 'epoch': 2.0}


TrainOutput(global_step=3126, training_loss=0.11752973331981031, metrics={'train_runtime': 2899.2484, 'train_samples_per_second': 17.246, 'train_steps_per_second': 1.078, 'train_loss': 0.11752973331981031, 'epoch': 2.0})

In [19]:
training_args = TrainingArguments(
    output_dir="left_padding150_model",
    learning_rate=2e-5,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    num_train_epochs=3,
    weight_decay=0.01,
    evaluation_strategy="epoch",
    save_strategy="epoch",
    load_best_model_at_end=True,
    push_to_hub=True,
    seed=42,
    data_seed=123)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_imdb_150["train"],
    eval_dataset=tokenized_imdb_150["test"],
    tokenizer=tokenizer,
    data_collator=data_collator,
    compute_metrics=compute_metrics,

)

trainer.train(resume_from_checkpoint=True)

  0%|          | 0/4689 [00:00<?, ?it/s]

You're using a DistilBertTokenizerFast tokenizer. Please note that with a fast tokenizer, using the `__call__` method is faster than using a method to encode the text followed by a call to the `pad` method to get a padded encoding.


{'loss': 0.0563, 'learning_rate': 5.07144380464918e-06, 'epoch': 2.24}
{'loss': 0.0529, 'learning_rate': 2.9387929195990615e-06, 'epoch': 2.56}
{'loss': 0.054, 'learning_rate': 8.061420345489445e-07, 'epoch': 2.88}


  0%|          | 0/1563 [00:00<?, ?it/s]

{'eval_loss': 0.38004031777381897, 'eval_accuracy': 0.92476, 'eval_runtime': 384.1011, 'eval_samples_per_second': 65.087, 'eval_steps_per_second': 4.069, 'epoch': 3.0}
{'train_runtime': 1434.0528, 'train_samples_per_second': 52.299, 'train_steps_per_second': 3.27, 'train_loss': 0.01731584208889744, 'epoch': 3.0}


TrainOutput(global_step=4689, training_loss=0.01731584208889744, metrics={'train_runtime': 1434.0528, 'train_samples_per_second': 52.299, 'train_steps_per_second': 3.27, 'train_loss': 0.01731584208889744, 'epoch': 3.0})

无法呈现“application/vnd.jupyter.widget-view+json”的内容
{"model_id":"057a11bb382d417db9ab4c042ad62e75","version_major":2,"version_minor":0}
You're using a DistilBertTokenizerFast tokenizer. Please note that with a fast tokenizer, using the `__call__` method is faster than using a method to encode the text followed by a call to the `pad` method to get a padded encoding.
{'loss': 0.3392, 'learning_rate': 1.6801023672424827e-05, 'epoch': 0.32}
{'loss': 0.2638, 'learning_rate': 1.3602047344849649e-05, 'epoch': 0.64}
{'loss': 0.2446, 'learning_rate': 1.0403071017274472e-05, 'epoch': 0.96}
无法呈现“application/vnd.jupyter.widget-view+json”的内容
{"model_id":"ea2324d6b1874a199e5f632294202495","version_major":2,"version_minor":0}
{'eval_loss': 0.3612774908542633, 'eval_accuracy': 0.86776, 'eval_runtime': 1436.2269, 'eval_samples_per_second': 17.407, 'eval_steps_per_second': 1.088, 'epoch': 1.0}
{'loss': 0.1764, 'learning_rate': 7.204094689699297e-06, 'epoch': 1.28}
{'loss': 0.1591, 'learning_rate': 4.005118362124121e-06, 'epoch': 1.6}
{'loss': 0.1746, 'learning_rate': 8.061420345489445e-07, 'epoch': 1.92}
无法呈现“application/vnd.jupyter.widget-view+json”的内容
{"model_id":"e11f558dc7c64eacb2d28408bf424940","version_major":2,"version_minor":0}
{'eval_loss': 0.24108850955963135, 'eval_accuracy': 0.92488, 'eval_runtime': 746.9998, 'eval_samples_per_second': 33.467, 'eval_steps_per_second': 2.092, 'epoch': 2.0}
{'train_runtime': 6886.1501, 'train_samples_per_second': 7.261, 'train_steps_per_second': 0.454, 'train_loss': 0.22320297629270353, 'epoch': 2.0}
TrainOutput(global_step=3126, training_loss=0.22320297629270353, metrics={'train_runtime': 6886.1501, 'train_samples_per_second': 7.261, 'train_steps_per_second': 0.454, 'train_loss': 0.22320297629270353, 'epoch': 2.0})

In [21]:
trainer.push_to_hub()

'https://huggingface.co/Realgon/left_padding150_model/tree/main/'

In [76]:
from transformers import pipeline
from datasets import load_dataset
from evaluate import evaluator
import evaluate

pipe = pipeline("text-classification", model="lvwerra/distilbert-imdb", device=0)
data = load_dataset("imdb", split="test").shuffle().select(range(1000))
metric = evaluate.load("accuracy")

Downloading (…)lve/main/config.json:   0%|          | 0.00/735 [00:00<?, ?B/s]

To support symlinks on Windows, you either need to activate Developer Mode or to run Python as an administrator. In order to see activate developer mode, see this article: https://docs.microsoft.com/en-us/windows/apps/get-started/enable-your-device-for-development


Downloading pytorch_model.bin:   0%|          | 0.00/268M [00:00<?, ?B/s]

Downloading (…)okenizer_config.json:   0%|          | 0.00/333 [00:00<?, ?B/s]

Downloading (…)solve/main/vocab.txt:   0%|          | 0.00/232k [00:00<?, ?B/s]

Downloading (…)/main/tokenizer.json:   0%|          | 0.00/466k [00:00<?, ?B/s]

Downloading (…)cial_tokens_map.json:   0%|          | 0.00/112 [00:00<?, ?B/s]

In [77]:
from evaluate import EvaluationSuite
suite = EvaluationSuite.load('mathemakitten/sentiment-evaluation-suite')
results = suite.run("huggingface/Realgon/left_padding50_model")

Downloading builder script:   0%|          | 0.00/1.34k [00:00<?, ?B/s]

Map:   0%|          | 0/10 [00:00<?, ? examples/s]

`data` is a preloaded Dataset! Ignoring `subset` and `split`.


HFValidationError: Repo id must be in the form 'repo_name' or 'namespace/repo_name': 'huggingface/Realgon/left_padding50_model'. Use `repo_type` argument if needed.

# Load Model

In [75]:
from transformers import pipeline

classifier = pipeline("sentiment-analysis", model="Realgon/left_padding50_model")
data = load_dataset("imdb", split="test").shuffle().select(range(1000))
metric = evaluate.load("accuracy")
metric

EvaluationModule(name: "accuracy", module_type: "metric", features: {'predictions': Value(dtype='int32', id=None), 'references': Value(dtype='int32', id=None)}, usage: """
Args:
    predictions (`list` of `int`): Predicted labels.
    references (`list` of `int`): Ground truth labels.
    normalize (`boolean`): If set to False, returns the number of correctly classified samples. Otherwise, returns the fraction of correctly classified samples. Defaults to True.
    sample_weight (`list` of `float`): Sample weights Defaults to None.

Returns:
    accuracy (`float` or `int`): Accuracy score. Minimum possible value is 0. Maximum possible value is 1.0, or the number of examples input, if `normalize` is set to `True`.. A higher score means higher accuracy.

Examples:

    Example 1-A simple example
        >>> accuracy_metric = evaluate.load("accuracy")
        >>> results = accuracy_metric.compute(references=[0, 1, 2, 0, 1, 2], predictions=[0, 1, 1, 2, 1, 0])
        >>> print(results)
    

<Tip>

[Trainer](https://huggingface.co/docs/transformers/main/en/main_classes/trainer#transformers.Trainer) applies dynamic padding by default when you pass `tokenizer` to it. In this case, you don't need to specify a data collator explicitly.

</Tip>

Once training is completed, share your model to the Hub with the [push_to_hub()](https://huggingface.co/docs/transformers/main/en/main_classes/trainer#transformers.Trainer.push_to_hub) method so everyone can use your model:

<Tip>

For a more in-depth example of how to finetune a model for text classification, take a look at the corresponding
[PyTorch notebook](https://colab.research.google.com/github/huggingface/notebooks/blob/main/examples/text_classification.ipynb)
or [TensorFlow notebook](https://colab.research.google.com/github/huggingface/notebooks/blob/main/examples/text_classification-tf.ipynb).

</Tip>

## Inference

Great, now that you've finetuned a model, you can use it for inference!

Grab some text you'd like to run inference on:

In [72]:
text = "This was a masterpiece. Not completely faithful to the books, but enthralling from beginning to end. Might be my favorite of the three."

The simplest way to try out your finetuned model for inference is to use it in a [pipeline()](https://huggingface.co/docs/transformers/main/en/main_classes/pipelines#transformers.pipeline). Instantiate a `pipeline` for sentiment analysis with your model, and pass your text to it:

In [73]:
from transformers import pipeline

classifier = pipeline("sentiment-analysis", model="Realgon/left_padding50_model")
classifier(text)

[{'label': 'POSITIVE', 'score': 0.9963740706443787}]

In [67]:
classifier.state.log_history

AttributeError: 'TextClassificationPipeline' object has no attribute 'state'

You can also manually replicate the results of the `pipeline` if you'd like:

Tokenize the text and return PyTorch tensors:

In [24]:
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("Realgon/left_padding50_model")
inputs = tokenizer(text, return_tensors="pt")

Pass your inputs to the model and return the `logits`:

In [25]:
from transformers import AutoModelForSequenceClassification
import torch

model = AutoModelForSequenceClassification.from_pretrained("Realgon/left_padding50_model")
with torch.no_grad():
    logits = model(**inputs).logits

Get the class with the highest probability, and use the model's `id2label` mapping to convert it to a text label:

In [26]:
predicted_class_id = logits.argmax().item()
model.config.id2label[predicted_class_id]

'POSITIVE'