<a href="https://colab.research.google.com/github/vilsonrodrigues/playing-with-falcon/blob/main/notebooks/falcon-7b-instruct-4bit.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>

References: https://huggingface.co/blog/4bit-transformers-bitsandbytes

### Install dependencies

In [None]:
!pip install -q -U bitsandbytes
!pip install -q -U git+https://github.com/huggingface/transformers.git
!pip install -q -U git+https://github.com/huggingface/peft.git
!pip install -q -U git+https://github.com/huggingface/accelerate.git
!pip install -q -U einops
!pip install -q -U safetensors
!pip install -q trl
!pip install xformers

[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m92.2/92.2 MB[0m [31m11.0 MB/s[0m eta [36m0:00:00[0m
[?25h  Installing build dependencies ... [?25l[?25hdone
  Getting requirements to build wheel ... [?25l[?25hdone
  Preparing metadata (pyproject.toml) ... [?25l[?25hdone
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m236.8/236.8 kB[0m [31m19.0 MB/s[0m eta [36m0:00:00[0m
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m7.8/7.8 MB[0m [31m114.5 MB/s[0m eta [36m0:00:00[0m
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m1.3/1.3 MB[0m [31m78.6 MB/s[0m eta [36m0:00:00[0m
[?25h  Building wheel for transformers (pyproject.toml) ... [?25l[?25hdone
  Installing build dependencies ... [?25l[?25hdone
  Getting requirements to build wheel ... [?25l[?25hdone
  Preparing metadata (pyproject.toml) ... [?25l[?25hdone
[2K     [90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━[0m [32m227.6/227.6 kB[0m [31m18.6 MB

### bitsandbytes configs

The 4bit integration comes with 2 different quantization types: FP4 and NF4. The NF4 dtype stands for Normal Float 4 and is introduced in the QLoRA paper

You can switch between these two dtype using bnb_4bit_quant_type from BitsAndBytesConfig. By default, the FP4 quantization is used.

This saves more memory at no additional performance - from our empirical observations, this enables fine-tuning llama-13b model on an NVIDIA-T4 16GB with a sequence length of 1024, batch size of 1 and gradient accumulation steps of 4.

To enable this feature, simply add `bnb_4bit_use_double_quant=True` when creating your quantization config!

(text from HF colab)



We will used NF4!

In [1]:
from datasets import load_dataset
from random import randrange

# Load dataset from the hub
dataset = load_dataset("databricks/databricks-dolly-15k", split="train")

print(f"dataset size: {len(dataset)}")
print(dataset[randrange(len(dataset))])

Found cached dataset json (/home/kamal/.cache/huggingface/datasets/databricks___json/databricks--databricks-dolly-15k-7427aa6e57c34282/0.0.0/8bb11242116d547c741b2e8a1f18598ffdd40a1d4f2a2872c7a28b697434bc96)


dataset size: 15011
{'instruction': 'Give me a list of the best bodybuilders of all time.', 'context': '', 'response': 'There were many great bodybuilders, but most associate the best with those who have won Mr. Olympia, the most prestigious bodybuilding competition in the world. Those who have the most wins in chronological order are:\n\nArnold Schwarzenegger (7 total wins)\nLee Haney (8 total wins)\nDorian Yates (6 total wins)\nRonnie Coleman (8 total wins)\nJay Cutler (4 total wins)\nPhil Health (7 total wins)', 'category': 'brainstorming'}


In [2]:
def format_instruction(sample):
	return f"""### Instruction:
Use the Input below to create an instruction, which could have been used to generate the input using an LLM.

### Input:
{sample['response']}

### Response:
{sample['instruction']}
"""

In [1]:
import torch
from transformers import BitsAndBytesConfig

quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
)

### Load model and pipeline

In [2]:
# My version with smaller chunks on safetensors for low RAM environments
model_id = "vilsonrodrigues/falcon-7b-instruct-sharded"

from transformers import AutoModelForCausalLM, AutoTokenizer,pipeline

In [3]:
tokenizer = AutoTokenizer.from_pretrained(model_id)
tokenizer.pad_token = tokenizer.eos_token
tokenizer.padding_side = "right"

In [4]:
torch.cuda.mem_get_info()

(11786059776, 12592087040)

In [9]:
torch.cuda.empty_cache()

In [5]:
model_4bit = AutoModelForCausalLM.from_pretrained(
        model_id,
        device_map="auto",
        quantization_config=quantization_config,
        trust_remote_code=True,
    resume_download=True)
model_4bit.config.pretraining_tp = 1

Loading checkpoint shards:   0%|          | 0/15 [00:00<?, ?it/s]

Some weights of FalconForCausalLM were not initialized from the model checkpoint at vilsonrodrigues/falcon-7b-instruct-sharded and are newly initialized: ['lm_head.weight']
You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.


In [6]:
torch.cuda.mem_get_info()

(7660109824, 12592087040)

In [7]:
print(model_4bit)

FalconForCausalLM(
  (transformer): FalconModel(
    (word_embeddings): Embedding(65024, 4544)
    (h): ModuleList(
      (0-31): 32 x FalconDecoderLayer(
        (self_attention): FalconAttention(
          (maybe_rotary): FalconRotaryEmbedding()
          (query_key_value): Linear4bit(in_features=4544, out_features=4672, bias=False)
          (dense): Linear4bit(in_features=4544, out_features=4544, bias=False)
          (attention_dropout): Dropout(p=0.0, inplace=False)
        )
        (mlp): FalconMLP(
          (dense_h_to_4h): Linear4bit(in_features=4544, out_features=18176, bias=False)
          (act): GELU(approximate='none')
          (dense_4h_to_h): Linear4bit(in_features=18176, out_features=4544, bias=False)
        )
        (input_layernorm): LayerNorm((4544,), eps=1e-05, elementwise_affine=True)
      )
    )
    (ln_f): LayerNorm((4544,), eps=1e-05, elementwise_affine=True)
  )
  (lm_head): Linear(in_features=4544, out_features=65024, bias=False)
)


In [7]:
import torch
import transformers

pipeline = transformers.pipeline(
        "text-generation",
        model=model_4bit,
        tokenizer=tokenizer,
        use_cache=True,
        device_map="auto",
        max_length=296,
        do_sample=True,
        top_k=10,
        num_return_sequences=1,
        eos_token_id=tokenizer.eos_token_id,
        pad_token_id=tokenizer.eos_token_id,
)

The model 'FalconForCausalLM' is not supported for text-generation. Supported models are ['BartForCausalLM', 'BertLMHeadModel', 'BertGenerationDecoder', 'BigBirdForCausalLM', 'BigBirdPegasusForCausalLM', 'BioGptForCausalLM', 'BlenderbotForCausalLM', 'BlenderbotSmallForCausalLM', 'BloomForCausalLM', 'CamembertForCausalLM', 'CodeGenForCausalLM', 'CpmAntForCausalLM', 'CTRLLMHeadModel', 'Data2VecTextForCausalLM', 'ElectraForCausalLM', 'ErnieForCausalLM', 'GitForCausalLM', 'GPT2LMHeadModel', 'GPT2LMHeadModel', 'GPTBigCodeForCausalLM', 'GPTNeoForCausalLM', 'GPTNeoXForCausalLM', 'GPTNeoXJapaneseForCausalLM', 'GPTJForCausalLM', 'LlamaForCausalLM', 'MarianForCausalLM', 'MBartForCausalLM', 'MegaForCausalLM', 'MegatronBertForCausalLM', 'MvpForCausalLM', 'OpenLlamaForCausalLM', 'OpenAIGPTLMHeadModel', 'OPTForCausalLM', 'PegasusForCausalLM', 'PLBartForCausalLM', 'ProphetNetForCausalLM', 'QDQBertLMHeadModel', 'ReformerModelWithLMHead', 'RemBertForCausalLM', 'RobertaForCausalLM', 'RobertaPreLayerNorm

In [8]:
sequences = pipeline("Girafatron is obsessed with giraffes, the most glorious animal on the face of this Earth.")

In [9]:
sequences

[{'generated_text': "Girafatron is obsessed with giraffes, the most glorious animal on the face of this Earth. He's so obsessed that he's the only man alive on the continent of Africa. He lives alone, away from other humans and the civilization. He can't get enough of giraffes and the feeling of seeing their majestic height and length is the only thing that he's obsessed with in his life. Girafatron has a very special talent. In the world of girafatrons he can communicate and have relationships with the most special species, the Girafatrons. In the past, he's been able to make a living by using his talents, but now he's struggling with his obsession. It's the year 2053 and the world has changed. There are no girafatrons left on the planet and their extinction may not be far away. Girafatron must find out a way to survive and to keep his giraffa company. He's not sure how to do it. He's not sure if he's able to make a living with his special talents, or if he can even find another Giraf

if the below import fail, then do 

pip install git+https://github.com/huggingface/peft

In [12]:
from peft import LoraConfig, prepare_model_for_kbit_training, get_peft_model

# LoRA config based on QLoRA paper
peft_config = LoraConfig(
        lora_alpha=16,
        lora_dropout=0.1,
        r=64,
        bias="none",
        task_type="CAUSAL_LM",
)


# prepare model for training
model = prepare_model_for_kbit_training(model_4bit)
model = get_peft_model(model, peft_config)

In [21]:
from transformers import TrainingArguments

args = TrainingArguments(
    output_dir="falcon-7-int4",
    num_train_epochs=3,
    per_device_train_batch_size= 4,
    gradient_accumulation_steps=2,
    gradient_checkpointing=True,
    optim="paged_adamw_32bit",
    logging_steps=10,
    save_strategy="epoch",
    learning_rate=2e-4,
    bf16=True,
    tf32=True,
    max_grad_norm=0.3,
    warmup_ratio=0.03,
    lr_scheduler_type="constant",
    disable_tqdm=True, # disable tqdm since with packing values are in correct
    use_cache=False
)

TypeError: TrainingArguments.__init__() got an unexpected keyword argument 'use_cache'

In [20]:
from trl import SFTTrainer

max_seq_length = 2048 # max sequence length for model and packing of the dataset

trainer = SFTTrainer(
    model=model,
    train_dataset=dataset,
    peft_config=peft_config,
    max_seq_length=max_seq_length,
    tokenizer=tokenizer,
    packing=True,
    formatting_func=format_instruction,
    args=args,
)

TypeError: SFTTrainer.__init__() got an unexpected keyword argument 'use_cache'

In [15]:
torch.cuda.mem_get_info()

(6439895040, 12592087040)

In [18]:
import os

os.environ['PYTORCH_CUDA_ALLOC_CONF']='max_split_size_mb:512'

In [19]:
trainer.train()

`use_cache=True` is incompatible with gradient checkpointing. Setting `use_cache=False`...


OutOfMemoryError: CUDA out of memory. Tried to allocate 2.22 GiB (GPU 0; 11.73 GiB total capacity; 8.34 GiB already allocated; 994.94 MiB free; 9.63 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.  See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF

### Use with LangChain

In [None]:
# Some error in colab. fix with
import locale
locale.getpreferredencoding = lambda: "UTF-8"

In [14]:
from langchain import HuggingFacePipeline

Load local LLM

In [15]:
llm = HuggingFacePipeline(pipeline=pipeline)

Define Template

In [16]:
from langchain import PromptTemplate, LLMChain

template = """Question: {question}
Answer: Let's think step by step."""

prompt = PromptTemplate(
    template=template,
    input_variables= ["question"]
)

Chain

In [17]:
llm_chain = LLMChain(prompt=prompt, llm=llm)

### **Tests**

In [18]:
llm_chain("How to prepare eggs?")

{'question': 'How to prepare eggs?',
 'text': ' You should start by washing a dozen or so eggs thoroughly in a bowl, then rinsing them in cold water to remove traces of dirt or shell bits. Next, you can place the cleaned eggs inside a bowl and cover them with cold water. Then you should allow them to sit for several hours in the refrigerator to allow the whites and yolks to solidify and the yolks to thicken. Finally, crack each egg over a pan that has been lightly greased with cooking spray or butter. Remove the yolks with a spoon, then place the white in a bowl, and you can mix the yolk and white as well. You can also add salt and pepper to these ingredients if desired. Then use any cooking method desired to prepare eggs for consumption. Some popular methods include poaching or baking, as well as a variety of ways to scramble or fry them.'}

In [19]:
llm_chain("Write a program for tic tac toe?")

{'question': 'Write a program for tic tac toe?',
 'text': " We need to write a function to handle the board state and the user inputs. The board consists of two rows of four tiles each. We can define two arrays of four tiles each for the top row, and another array for the bottom row. \n\nThe function should take three arguments - row, col, and type (up = player is above, down = player is below). The function should then update the board state according to the inputs. \n\nFirst, we need to check if the board is full and reset it if so. We can do this by looping through the boards and updating the indices of the current row and column. We need another variable to keep track of the number of empty tiles in each row. \n\nFor the next step, we can define a function to handle the user inputs. We need to prompt the user to enter up and down for row and col. If the user enters a number (i.e the index of a specific tile or location), we simply print it to console, and if the row and col index i

In [21]:
sequence = llm_chain("Write a program for tic tac toe in python?")

In [23]:
print(sequence['text'])

 First, we need to create a board to store the positions of each piece. We can use a 2D array to achieve this. We will need to initialize this array to hold the board's configuration. 

Next, we would want the program to check if there are no spaces on the board, as in a tic-tac-toe game is a draw. We can do this by using a simple for loop to iterate through each position, checking if it's vacant (i.e. contains no piece). 

After the check, we would want to set up the game by placing pieces at each position based on whether we have a tic or a tic-a-tic (i.e. we have a tic but not a tic-a-tic), and if not, we set the tic-a-tic to that value. 

To handle the game, we would need a method to check if we have won (i.e. a tic-a-tic) or lost. We can check for winning by comparing the current board with the initial board, and if both are the same, we have a winner. We may also use a flag to track whether we have won before, and we can use it to update the game accordingly.


In [None]:
llm_chain_2 = LLMChain(prompt=prompt2, llm=llm)

In [None]:
result_explanation = llm_chain_2("Explain antibiotics")

In [None]:
prompt_pt_grafos = """No ramo de análise de grafos, existe uma métrica chamada Clustering Coefficient,
   você pode me falar como interpretar ela?"""
llm_chain_2(prompt_pt_grafos)



{'question': 'No ramo de análise de grafos, existe uma métrica chamada Clustering Coefficient, \n   você pode me falar como interpretar ela?',
 'text': "/n Clustering Coefficient is a metric used in clustering algorithms to quantify the level of 'fidelity' or 'homogeneity' between groups of items. It is commonly used to identify overlapping groups and clusters in a dataset. It ranges from 1 (high) to -1 (low). The closer the value to 1, the more similar the groups and clusters are. Clustering coefficients range from 1 to 4, with 4 indicating that the dataset can be divided into four distinct groups."}

In [None]:
llm_chain_2("what is a convolution?")

{'question': 'what is a convolution?',
 'text': '/n a convolution, in the context of computer graphics, is when an image is processed in multiple dimensions, typically to blur or sharpen the image. This processing can involve convolution of pixels in the image. The term is derived from the mathematical term convolution, which is a mathematical operation that involves overlapping of two functions of the same shape (i.e., the convolution of functions of the shape). The result of the convolution operation is the resulting shape. In computer graphics, the process of convolution is performed to enhance the quality of the rendered image. This is usually achieved by using a mask to filter an image and then convolving an image with the mask. This can be done in real-time on the graphics processing unit (GPU). A convolution can involve a single dimension or multiple dimensions. If there is only one dimension, it is called 1D convolution. If two dimensions, it is called 2D convolution. If more t

In [None]:
prompt_code = """ I have to pass 2 values that are as a string

'2,3'

And turn them into a tuple

(2,3), how to do?
"""

(hit token limit)

In [None]:
llm_chain_2(prompt_code)



{'question': " I have to pass 2 values that are as a string\n\n'2,3'\n\nAnd turn them into a tuple\n\n(2,3), how to do?\n",
 'text': '\nOne way to convert the string into a tuple using a list:\n\n- "2,3" is a string,\n- "2,3" is a pair consisting of a string and a number,\n- Therefore, you can convert the string into a list of a pair of strings, then into a tuple.\n- Here\'s an example of how you can do it:\n\n```python\nstring = "2,3"\nstring_list = list(string)\nlist_of_pairs = string_list\npairs = list(list_of_pairs)\n\n```\n\nNow `pairs` is a list containing strings like `"2,3"` and also has a pair with a string and an integer like `("2,3")`.\nTo convert this list of strings into a tuple, you can create an anonymous function to create the tuple:\n\n```python\npairs = list(",".join(s"("%d", i+1)%s%d" for d,i,i+1 in zip(string_list, string_list[i+1], int'}

In [None]:
prompt_code2 = """
/*
Write a python code to ask the user for their name and say "Hello"
*/
"""
llm_chain_2(prompt_code2)

{'question': '\n/*\nWrite a python code to ask the user for their name and say "Hello"\n*/\n',
 'text': '\nname = input( "What is your name? " )\nprint("Hello!")\nUser '}

In [None]:
llm_chain_2("How to convert a base64 file to bytes in python?")



{'question': 'How to convert a base64 file to bytes in python?',
 'text': "/n \nYou can convert a base64 file to bytes in python by using the 'open()' function to read the file as binary bytes, and then converting these binary bytes to strings using the 'decode()' method. Here's an example: \n\n```\nfile = open('file.b64')\nbytes = file.read()\nfile.close()\nbase64 = ''.join(bytes.decode('utf-8'))\n``` \n\nIn this example, the file `file.b64` is opened using the `open()` function and read as binary bytes using the `read()` method. The resulting binary bytes are then converted to strings using the `decode('utf-8')` method, which converts the binary bytes to a Python strings using the `utf-8` encoding. \n\nFinally, the resulting binary string is joined together using the `join()` method and a blank string to create a base64 encoded string."}

In [None]:
prompt_sql = """
Table departments, columns = [DepartmentId, DepartmentName]
Table students, columns = [DepartmentId, StudentId, StudentName]
Create a MySQL query for all students in the Computer Science Department
"""
llm_chain_2(prompt_sql)



{'question': '\nTable departments, columns = [DepartmentId, DepartmentName]\nTable students, columns = [DepartmentId, StudentId, StudentName]\nCreate a MySQL query for all students in the Computer Science Department\n',
 'text': '\nThe MySQL query for this would be: \n\n```\nSELECT students.StudentId, students.StudentName \nFROM students\nWHERE departments.DepartmentId = yourDepartmentId\n``` \n\nTo select all students in the Computer Science department, you can use a similar query. The `.` character is used as the wildcard to select all columns, since only the StudentId column is required. \n'}

In [None]:
llm_chain_2("como funciona o método __call__ em python")

{'question': 'como funciona o método __call__ em python',
 'text': '/n como funciona o método __call__ em python? - /n como funciona o método __call__ em python /n como funciona o método __call__ em python\n<p>__call__ is a method in Python that takes another object as an argument and calls a method of that object. When a new instance of the object is created, the method gets called automatically. However, this behavior can be overridden with an explicit call using the __call__ method. So, to answer the question, __call__ is the method that takes another object as an argument and calls a function or method of that other object. It can be called directly without having a __call__ method defined.'}

In [None]:
llm_chain_2("show me how python's args and kwargs work")



{'question': "show me how python's args and kwargs work",
 'text': '/n\nHere\'s an explanation and example:\n\nIn Python, you can access command-line arguments and their types using the `args` and `kwargs` arguments of a `__main__()` function.\n\n`args` is a list of arguments passed to the function, which include the names of the arguments themselves, their data types, and their order.\n\nFor example:\n\n```python\ndef hello(name, message):\n    args = [\'-n\', \'name\']\n    if len(args) > 1:\n        print(args)\n    print(args[-1])\n    message("Hello, %s!" % name.title())\n\nhello()\n```\n\nIn this example, the `arg` variable contains the arguments `-n` and `name`, which means that the `name` argument is passed with a data type of `<class\'str\'>` and is the last argument passed (after the `-n`.\n\nSimilarly, `kwargs` is used to pass keyword arguments (arguments with data types and/or names). These include arguments that are defined in the `__init__()` method of an object.\n\nFor e

In [None]:
llm_chain_2("What's latency definition?")

{'question': "What's latency definition?",
 'text': "/n Latency is the amount of time it takes an item to move from one part of a system to another. It's commonly used to measure how quickly data is transmitted between computers on a network. /n "}

In [None]:
llm_chain_2("what is Python's ABC library and what is it for?")

{'question': "what is Python's ABC library and what is it for?",
 'text': '/n The ABC library stands for Absolute Backwards Compatible. It was developed by David Mazzucchelli (who also developed the original Python interpreter) as an attempt to improve the performance of Python. The aim was to create a library that could run on any platform (in particular, Unix) and to provide a standard interface to the Python interpreter.\n\nIn other words, you can think of the ABC library as the standard interface to the Python interpreter. The library provides a simple interface to functions and data that are not in the interpreter, making them easier to use.\n\nSome notable features of the ABC library include:\n\n- Absolute Backwards Compatibility: The library allows Python to run on older versions of Windows, where it may have trouble running on newer Windows.\n\n- Improved Performance: The library allows the Python interpreter to compile and execute code faster, making it a great addition to the

In [None]:
llm_chain_2("Write me a diet, my goal is to gain lean mass and I will work out")

{'question': 'Write me a diet, my goal is to gain lean mass and I will work out',
 'text': '/n \nBreakfast:\n- Oatmeal: 50g of oats and 100g of water\n- 1 banana\n- 2 tablespoons of natural peanut butter\n\nSnack:\n- 2 eggs\n- 2 slices of bacon\n\nLunch:\n- Grilled chicken salad\n- 1 cup of quinoa\n- 2 slices of avocado\n- Handful of cherry tomatoes\n\nDinner:\n- Baked salmon\n- 1 cup of quinoa\n- 1 cup of broccoli\n- 2 eggs\n- 2 tablespoons of olive oil\n\nSnack:\n- 2 hard-boiled eggs\n- Greek yogurt\n\nMake sure to stay within your calorie goal and adjust portion sizes to fit your workout schedule.'}

In [None]:
llm_chain_2("what is the difference between Similarity embeddings and search embeddings")



{'question': 'what is the difference between Similarity embeddings and search embeddings',
 'text': '/n \nSimilarity embeddings represent the relationship between two items in a high-dimensional space. Search embeddings represent the similarity between two documents in a lower-dimensional space. The difference is that similarity embeddings learn and represent the intrinsic relationships between items, while search embeddings represent the similarity of documents. The similarity between documents is based on their content similarity, which can be more or less abstract depending on the type of documents.'}

Conversation

In [None]:
template_chat = """You are now a conversational assistant and must answer the questions: /n {history}"""

prompt_chat = PromptTemplate(
    template=template_chat,
    input_variables= ["history"]
)
llm_chain_chat = LLMChain(prompt=prompt_chat, llm=llm)

In [None]:
prompt_conversation1 = """
The following is a conversation with an AI research assistant. The assistant tone is technical and scientific.
Human: Hello, who are you?
AI: Greeting! I am an AI research assistant. How can I help you today?
Human: Can you tell me about the creation of blackholes?
AI:
"""

In [None]:
llm_chain_chat(prompt_conversation1)



{'history': '\nThe following is a conversation with an AI research assistant. The assistant tone is technical and scientific.\nHuman: Hello, who are you?\nAI: Greeting! I am an AI research assistant. How can I help you today?\nHuman: Can you tell me about the creation of blackholes?\nAI:\n',
 'text': 'Black holes are formed when a massive star collapses under its own gravity. As the core of the star becomes denser and denser, the more it is pulled in and becomes a gravitational superpower. Eventually this core becomes so powerful that it sucks everything in, including light, to its boundary, forming a black hole.'}