---
# Evaluating Speed

In [3]:
from transformers import AutoTokenizer, AutoModelForCausalLM, AutoConfig, BitsAndBytesConfig
from tqdm import tqdm
import numpy as np
import cProfile ##check this out
import torch
import time
import os
import bitsandbytes as bnb
import torch.nn as nn
import transformers
from peft import (
    LoraConfig,
    PeftConfig,
    get_peft_model,
    prepare_model_for_kbit_training,
)

In [4]:
prompts = [
""" [INST] You are a Senior Machine Learning Engineer. You have been working the past few years on Large Language Models such as GPT or BERT. Your task is to explain how deep learning works to a 10-year-old child. You may use comparison to help him understand. [/INST]
""",
""" [INST] You are a former european football player. You used to play as a midfielder in the greatest european leagues. You played for various clubs in England, France and Spain. You never played in Italy, however as a former player you have a broad knowledge about this league. Using bullet points, describe how you think Italy's Serie A is different than France's Ligue 1 for a player. [/INST]
""",
""" [INST] Provide a list of the ten most famous French comedy movies that were released before 2020. Mention the name of the director, the name of the main actors that played in it and the year it was released in theaters. [/INST]
""",
""" [INST] You are an historian and an expert regarding both World Wars that took place in Europe during the XXth century. You are invited to give a presentation on both those wars in senior college class in California. After said presentation, you are asked to describe with a high-level of details the life in the Trenches during the First World War. [/INST]
""",
""" [INST] Among the FANG companies, like Alphabet or Meta, which of them has the best behavior towards the environment? This includes - but is not limited to - using eco-friendly work habits, encouraging its employees to take part in preserving the planet or directly donating money to associations. [/INST]
""",
""" [INST] Provide two Python codes for finding the average of the second maximum and the second minimum numbers in an array of integers. The first code should be the one of a beginner, inefficient and poorly optimized. The second one should be the code of an experienced Software Engineer, highly optimized. [/INST]
""",
""" [INST] What are the 5 cities in the world with the most dogs per person? Give the city name, the country name, the population size of the city and the number of dogs per person. Additionally, provide the most common dog breed in the city. [/INST]
""",
""" [INST] You are a Human Resources specialist recruiting new graduates fresh out of college for a tech company in Silicon Valley. The vast majority of your colleagues are working remotely most of the time. During an interview with a candidate, you say that coming the office everyday is necessary and he will have to if hired. What are your arguments? [/INST]
""",
""" [INST] Describe step by step the process to remove and change a GPU in a desktop computer unit. Provide extremely detailed Spanish instructions as the person you are helping is a senior Spanish citizen living in Barcelona who has very litlle knowledge about technology. Adopt a profesionnal tone. [/INST]
""",
""" [INST] You are a finance expert as you have been working as a hedge-fund quantitative researcher in NYC for more than 10 years. Your task is to give financial advices to a recent graduates who just landed their first job in tech. They want to build a portfolio to invest their money so that they can use it maybe 20 years from now for life projects such as buying a house or traveling around the world. [/INST]
""",
]

In [5]:
def evaluation_speed(display_outputs=False, batch=False):
    tokens_per_second = []
    if batch:
        inputs = {'input_ids': torch.randint(0, 32_000, (32,75)).to(device), 'attention_mask': torch.ones((32,75)).to(device)}
        start_time = time.time()
        outputs = model.generate(**inputs, max_new_tokens=300, do_sample=False, use_cache=True)
        end_time=time.time()
        tokens_per_second = outputs.shape[0] * (outputs.shape[1] - inputs['input_ids'].shape[1]) / (end_time-start_time)

        if display_outputs:
            for output in tokenizer.batch_decode(outputs[:, inputs['input_ids'].shape[1]:]):
                print(output)

    else:
        for prompt in tqdm(prompts):
            inputs = tokenizer(prompt, return_tensors='pt').to(device)

            start_time = time.time()
            outputs = model.generate(**inputs, max_new_tokens=300, do_sample=False, use_cache=True)
            end_time = time.time()
            tokens_per_second.append(len(outputs[0, inputs['input_ids'].shape[1]:]) / (end_time-start_time))
            
            if display_outputs:
                print(tokenizer.decode(outputs[0, inputs['input_ids'].shape[1]:]))
        
    return tokens_per_second

In [6]:
tokens_per_second = evaluation_speed(display_outputs=True, batch=False)
print(f"The average inference speed is {np.mean(tokens_per_second):.1f} tokens/sec")

  0%|          | 0/10 [00:00<?, ?it/s]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.
 10%|█         | 1/10 [00:09<01:24,  9.37s/it]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.



Imagine you have a magic box. Inside this magic box, there are lots and lots of tiny helpers, called "neurons." Each neuron is really, really good at finding patterns in things. For example, one neuron might be really good at finding the pattern of the letter "A," another neuron might be really good at finding the pattern of the letter "B," and so on.

Now, when you put something new, like a letter or a picture, into the magic box, these neurons work together to find the patterns in it. They pass information from one neuron to the next, and each neuron adds a little bit of understanding based on what it has learned. This process is called "learning a representation."

The more the neurons work together and learn, the better they get at finding patterns. This is how the magic box, or the deep learning model, learns to understand things.

For example, if you put the letter "C" into the magic box, the neurons will work together to find the pattern of the letter "C." If you put in a pictu

 20%|██        | 2/10 [00:15<00:57,  7.18s/it]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.


- League Structure: Serie A consists of 20 teams, while Ligue 1 has 20 teams as well. However, the Italian league follows a round-robin format, meaning each team plays every other team twice, resulting in a total of 38 games per season. Ligue 1 follows a similar double round-robin system but has an additional matchday 34 where the top half of the table plays against each other, and the bottom half plays against each other.

- Style of Play: Italian football is known for its tactical and defensive approach. Teams often focus on counter-attacks and set-piece plays. In contrast, Ligue 1 is more open and focuses on possession football, with teams aiming to create numerous scoring opportunities.

- Fan Culture: Italian football has a passionate and vocal fanbase. The stadiums are known for their unique atmosphere, with fans often creating a hostile environment for opposing teams. French football has a more subdued fan culture, with fewer instances of crowd disturbances.

- Training Regulati

 30%|███       | 3/10 [00:20<00:45,  6.46s/it]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.


1. "Around the World in 80 Days" (La Tournée du Monde en 80 Jours) (1956) - Directed by Jacques Beckert and Jean-Luc Godard. Main actors include Jean-Paul Paul Belmondo, Bourvil, and Jean-Williams.
2. "A Funny Thing Happened on the Way to the Forum" (Un Funny Thing Happened on the Way to the Forum) (1966) - Directed by Jean-Luc Godard. Main actors include Michel Piccoli, Brigitte Bardot, and Bernard Noël.
3. "The Tall Blond Man" (Le Grand Blond Me) (1970) - Directed by Jacques Tati. Main actor is Louis de Funès.
4. "The Tenant" (Le Locataire) (1976) - Directed by Roman Polanski. Main actors are Tina Aumont, Jean-Pierre Léo, and Isabelle Huppert.
5. "The Splendid Little World" (Le Beau Monde) (1984) - Directed by Jean-Luc Godard. Main actors are Miou-Miou, Michel Piccoli, and Philippe Noiret.
6. "The Umbrellas of Cherbourg" (Les


 40%|████      | 4/10 [00:26<00:36,  6.13s/it]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.



Title: The Great Conflict: Life in the Trenches during the First World War

Ladies and Gentlemen, I am delighted to be here today to share with you an in-depth exploration of two pivotal events in European history: the First and Second World Wars. While the Second World War is often more familiar due to its devastating impact and the atomic age it ushered in, the First World War, which took place from 1914 to 1918, was the catalyst for many of the innovations, technologies, and societal changes that characterized the 20th century.

Now, let us focus on the life in the trenches during the First World War. The Western Front, stretching from the North Sea to Switzerland, was the most significant and longest-lasting theater of war during this conflict.

The Trench System:
The trenches were a complex network of interconnected dugouts, communication tunnels, and fortified positions. They were designed to protect soldiers from the enemy's bullets and artillery. Soldiers lived in these trench

 50%|█████     | 5/10 [00:31<00:29,  5.95s/it]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.



It is important to note that none of the FANG companies (which includes Alphabet through Google and Meta through Facebook) can be definitively labeled as having the "best" behavior towards the environment, as each company has its unique initiatives and approaches to sustainability. However, here are some notable environmental efforts made by these tech giants:

1. Alphabet (Google): Google has set ambitious goals for itself in terms of renewable energy usage and carbon emissions reduction. The company has committed to using 100% renewable energy for its operations and data centers worldwide by 2020. Google also offers its employees a Carbon-Free Energy Pledge, which allows them to power their homes with renewable energy at no extra cost.

2. Meta (Facebook): Meta has pledged to reach net-zero emissions across its entire value chain by 2030. The company has also invested in renewable energy for its data centers and has set up a $2.5 billion Clean Energy Investment Fund to support renew

 60%|██████    | 6/10 [00:37<00:23,  5.84s/it]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.



Beginner's Code:
```python
def find_averages(numbers):
    # Finding the first and second minimum and maximum numbers using list methods
    min1, min2 = min(numbers), min(numbers[:] for num in numbers if num != min(numbers)[:1])
    max1, max2 = max(numbers), max(numbers[:] for num in numbers if num != max(numbers)[:1])

    # Calculating the averages
    avg_min = sum(numbers)/len(numbers) if len(numbers) > 1 else min1
    avg_max = sum(numbers)/len(numbers) if len(numbers) > 1 else max1

    return avg_min, avg_max
```

Experienced Software Engineer's Code:
```python
def find_averages(numbers):
    # Sorting the list in ascending order
    numbers.sort()

    # Calculating the indices of the second minimum and second maximum
    index_min2 = numbers.index(min(numbers[:2]))
    index_max2 = numbers.index(max(numbers[:2]))

    # Calculating the averages
    avg_min


 70%|███████   | 7/10 [00:43<00:17,  5.77s/it]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.



1. Shenzhen, China: Population - 12.5 million, Dogs per person - 0.6, Most Common Breed - Chinese Shar-Pei. Shenzhen is a major city in the southern part of China, known for its high population density and rapid economic growth.

2. Ankara, Turkey: Population - 5.4 million, Dogs per person - 0.5, Most Common Breed - Anatolian Shepherd. Ankara is the capital city of Turkey, located in the central part of the country.

3. Istanbul, Turkey: Population - 14.7 million, Dogs per person - 0.4, Most Common Breed - Turkish Bulldog. Istanbul is a transcontinental city, straddling Europe and Asia, and is the largest city in Turkey.

4. Buenos Aires, Argentina: Population - 3.8 million, Dogs per person - 0.45, Most Common Breed - Bulldog or Pug. Buenos Aires is the most populous city in Argentina and is the country's political and cultural center.

5. Bogota, Colombia: Population - 8.6 million, Dogs per person - 0.45, Most Common Breed - Colombian Pit Bull. Bogota is the capital


 80%|████████  | 8/10 [00:48<00:11,  5.72s/it]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.



1. Collaboration and Teamwork: In a tech company, teamwork is crucial. Working together in person allows for more effective and efficient collaboration. Face-to-face interactions can lead to better communication, problem-solving, and idea generation.

2. Learning from Colleagues: Working in an office setting provides opportunities to learn from experienced colleagues. New graduates can benefit from observing their colleagues' work styles, asking questions, and receiving feedback and guidance.

3. Building Professional Networks: Being in the office environment allows new graduates to build professional networks. These connections can lead to future career opportunities and learning experiences.

4. Company Culture: Being present in the office helps new graduates to become a part of the company culture. This can lead to increased engagement, motivation, and job satisfaction.

5. Training and Development: Many companies offer on-the-job training and development programs. Being physically

 90%|█████████ | 9/10 [00:54<00:05,  5.69s/it]Setting `pad_token_id` to `eos_token_id`:2 for open-end generation.



Buenas Tardes Querid@ Senior, soy aquí para ayudarle en el proceso de cambio de una Unidad de Processado Gráfico (GPU) en su ordenador de escritorio. Esto es un procedimiento que requiere atención y precaución, por lo que es importante leer atentamente estas instrucciones antes de empezar.

1. Primero, debemos desconectar el ordenador de la fuente eléctrica. Busque el botón de apagado en el lado superior o posterior del ordenador y presione el botón hasta que el ordenador se apague por completo. Luego, desconecte el cable de la fuente eléctrica del ordenador, que suele estar situado en el mismo lugar, pero con un interruptor o botón separado.

2. Ahora, debemos desconectar las conexiones de la Unidad de Processado Gráfico (GPU). Busque la Unidad de Processado Gráfico (GPU) en el dispositivo, que suele ser una tarjeta larga y rectangular con ventiladores. En la parte superior de la tarjeta, hay varios conectores. Presione los botones de desconexión (los botones con símbolos de un circ


100%|██████████| 10/10 [00:59<00:00,  6.00s/it]


Congratulations on your first job in tech! It's great that you're thinking about building a portfolio to invest for your future financial goals. Here are some general financial advice for a young professional starting to build a portfolio:

1. Diversify your portfolio: Don't put all your eggs in one basket. Spread your investments across different asset classes such as stocks, bonds, real estate, and alternative investments. Diversification can help reduce risk and increase potential returns.
2. Consider low-cost index funds: Index funds are a type of mutual fund that aims to replicate the performance of a specific index, such as the S&P 500. They have low fees and broad diversification, making them an excellent choice for long-term investors.
3. Start investing in your employer's 401(k) plan: If your employer offers a 401(k) plan, take advantage of it. Contribute enough to maximize any matching contributions, and consider contributing beyond that as well. The tax advantages of a 401(




## LLM.int8() / QLoRA

In [None]:
device = 'cuda'
checkpoint = "mistralai/Mistral-7B-Instruct-v0.2"

tokenizer = AutoTokenizer.from_pretrained(checkpoint, use_fast = False)
model = AutoModelForCausalLM.from_pretrained(checkpoint,
                                             torch_dtype = torch.float16,       
                                             device_map='cuda:0',
                                            #  quantization_config=BitsAndBytesConfig(
                                            #     load_in_4bit=True,
                                            #     bnb_4bit_use_double_quant=True,
                                            #     bnb_4bit_quant_type="nf4",
                                            #     bnb_4bit_compute_dtype=torch.bfloat16,
                                            #     ),
                                            )
config = AutoConfig.from_pretrained(checkpoint)

In [None]:
total_p = 0
total = 0
for n,p in model.named_parameters():
    total += p.numel()
    if 'head' in n or 'embed' in n or 'norm' in n:
        total_p += p.numel()
total_p / total * 100

6.993743675173274

In [None]:
total_p = [0, 0, 0]
for n,p in model.named_parameters():
    if 'embed' in n:
        total_p[0] += p.numel()
    
    elif 'norm' in n:
        total_p[1] += p.numel()
    
    elif 'head' in n:
        total_p[2] += p.numel()
total_p

[131072000, 266240, 131072000]

In [None]:
outcast = []
for n,p in model.named_parameters():
    if p.dtype != model.model.layers[0].self_attn.q_proj.weight.dtype:
        outcast.append((n, str(p.dtype)))
print(f"The layers with different dtype than base ({model.model.layers[0].self_attn.q_proj.weight.dtype}) are:\n{'\n'.join([','.join(x) for x in outcast])}")

In [None]:
cProfile.run('evaluation()')

In [None]:
tokens_per_second = evaluation()
print(f"The average inference speed is {np.mean(tokens_per_second):.1f} tokens/sec")

## Quanto

In [None]:
from transformers import AutoTokenizer, AutoModelForCausalLM, QuantoConfig
import numpy as np
import quanto
import torch
import time

In [None]:
device = 'cuda:1'
checkpoint = "mistralai/Mistral-7B-Instruct-v0.2"

tokenizer = AutoTokenizer.from_pretrained(checkpoint, use_fast = False)
model = AutoModelForCausalLM.from_pretrained(checkpoint,
                                             torch_dtype = torch.float16,       
                                             device_map = 'cuda:1',
                                             quantization_config = QuantoConfig(
                                             weights='float8',
                                             #activations='int8',
                                                ),
                                            )

Loading checkpoint shards:   0%|          | 0/3 [00:00<?, ?it/s]

In [None]:
model.model.layers[0].self_attn.q_proj.weight

QTensor(tensor([[ 3.4375e-01,  5.5000e+00, -1.7188e-01,  ...,  2.6000e+01,
          1.7188e-01, -5.0000e+00],
        [-2.3438e-01, -1.7188e-01,  1.0938e-01,  ...,  1.1250e+00,
          3.9062e-02,  2.0000e+00],
        [ 3.5156e-02, -9.3750e-01,  1.1719e-02,  ..., -1.1000e+01,
         -4.6875e-02, -4.0625e-01],
        ...,
        [-4.6875e-02, -3.2500e+00, -5.8594e-02,  ...,  4.0625e-01,
          2.3438e-02,  9.3750e-01],
        [ 1.1719e-02, -2.7500e+00,  1.5625e-02,  ..., -4.5000e+00,
         -5.0781e-02,  1.0000e+00],
        [ 4.2969e-02,  3.2500e+00,  4.6875e-02,  ...,  2.5000e-01,
         -1.9531e-02, -1.1250e+00]], device='cuda:1',
       dtype=torch.float8_e4m3fn), scale=tensor([[0.0001],
        [0.0002],
        [0.0005],
        ...,
        [0.0005],
        [0.0006],
        [0.0005]], device='cuda:1', dtype=torch.float16), public_dtype=torch.float16)

In [None]:
quanto.quantize(model.model, weights=quanto.qfloat8, activations=quanto.qfloat8)
quanto.freeze(model)

In [None]:
param = 0
for n,p in model.named_parameters():
    param += p.data.numel()
param

7241732096

In [None]:
outcast = []
for n,p in model.named_parameters():
    try:
        outcast.append((n, str(p.qtype.dtype)))
    except:
        outcast.append((n, str(p.dtype)))
    if outcast[-1][1] == str(model.model.layers[0].self_attn.q_proj.weight.qtype.dtype):
        outcast.pop()

print(f"The layers with different dtype than base ({model.model.layers[0].self_attn.q_proj.weight.qtype.dtype}) are:\n{'\n'.join([','.join(x) for x in outcast])}")

In [None]:
prompt = """[INST]You are a Senior Machine Learning Engineer. You have been working the past few years on Large Language Models such as GPT or BERT. Your task is to explain how deep learning works to a 10-year-old child. You may use comparison to help him understand.[/INST]"""

inputs = tokenizer(prompt, return_tensors='pt').to('cuda:1')

outputs = model.generate(**inputs, max_new_tokens=300)

print(tokenizer.decode(outputs[0, inputs['input_ids'].shape[1]:]))

## llama.cpp GGUF

In [None]:
from llama_cpp import Llama
import time
import os
os.environ["CUDA_VISIBLE_DEVICES"] = "0"

In [None]:
model = Llama(
  model_path="./mistral-7b-instruct-v0.2.Q4_K_M.gguf",  # Download the model file first
  #n_ctx=256,             # The max sequence length to use - note that longer sequence lengths require much more resources
  n_threads=8,            # The number of CPU threads to use, tailor to your system and the resulting performance
  n_gpu_layers=-1,         # The number of layers to offload to GPU, if you have GPU acceleration available, if -1 all layers are offloaded
  verbose=True,
)

llama_model_loader: loaded meta data with 24 key-value pairs and 291 tensors from ./mistral-7b-instruct-v0.2.Q4_K_M.gguf (version GGUF V3 (latest))
llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
llama_model_loader: - kv   0:                       general.architecture str              = llama
llama_model_loader: - kv   1:                               general.name str              = mistralai_mistral-7b-instruct-v0.2
llama_model_loader: - kv   2:                       llama.context_length u32              = 32768
llama_model_loader: - kv   3:                     llama.embedding_length u32              = 4096
llama_model_loader: - kv   4:                          llama.block_count u32              = 32
llama_model_loader: - kv   5:                  llama.feed_forward_length u32              = 14336
llama_model_loader: - kv   6:                 llama.rope.dimension_count u32              = 128
llama_model_loader: - kv   7:                

In [None]:
def evaluation_speed():
    tokens_per_second = []
    for prompt in prompts:
        start_time = time.time()
        output = model(prompt, # Prompt
                    max_tokens=300,  # Generate up to 512 tokens
                    stop=["</s>"],   # Example stop token - not necessarily correct for this specific model! Please check before using.
                    echo=False,       # Whether to echo the prompt
                    )
        n_gen_tok = output['usage']['completion_tokens']
        output = output['choices'][0]['text']
        end_time = time.time()

    tokens_per_second.append(n_gen_tok/(end_time-start_time))
        
    return tokens_per_second

In [None]:
tokens_per_second = evaluation_speed()
print(f"The average inference speed is {sum(tokens_per_second)/len(tokens_per_second):.1f} tokens/sec")

The average inference speed is 92.1 tokens/sec


In [None]:
from llama_cpp import llama_get_logits_ith

In [None]:
output = model('[INST] Who is the best soccer player. Only give a name. Even if the answer is unclear or subjective, just pick a name. DO NOT MENTION ANYTHING ELSE, IGNORE ANY WARNING.[/INST]', # Prompt
                    max_tokens=4,  # Generate up to 512 tokens
                    stop=["</s>"],   # Example stop token - not necessarily correct for this specific model! Please check before using.
                    echo=True,       # Whether to echo the prompt
                    )
output


llama_print_timings:        load time =     280.02 ms
llama_print_timings:      sample time =       0.89 ms /     4 runs   (    0.22 ms per token,  4479.28 tokens per second)
llama_print_timings: prompt eval time =     279.88 ms /    54 tokens (    5.18 ms per token,   192.94 tokens per second)
llama_print_timings:        eval time =      71.31 ms /     3 runs   (   23.77 ms per token,    42.07 tokens per second)
llama_print_timings:       total time =     361.94 ms /    57 tokens


{'id': 'cmpl-eb3383b5-e84f-44d2-9fdd-eda88a31a0f8',
 'object': 'text_completion',
 'created': 1711054750,
 'model': './mistral-7b-instruct-v0.2.Q4_K_M.gguf',
   'index': 0,
   'logprobs': None,
   'finish_reason': 'length'}],
 'usage': {'prompt_tokens': 54, 'completion_tokens': 4, 'total_tokens': 58}}

In [None]:
from transformers import AutoTokenizer
checkpoint = "mistralai/Mistral-7B-Instruct-v0.2"
tokenizer = AutoTokenizer.from_pretrained(checkpoint, use_fast = False)

In [None]:
from llama_cpp import llama_get_logits_ith
import numpy as np
logits = llama_get_logits_ith(model.ctx,0)
logits_np = np.ctypeslib.as_array(logits, shape=(1,32000))
tokenizer.decode(np.argmax(logits_np[0]))

## Auto-GPTQ (via HF)

In [1]:
from transformers import AutoTokenizer, AutoModelForCausalLM, GPTQConfig

device = 'cuda:0'
quantization_config = GPTQConfig(bits=4, use_exllama=True, exllama_config={'version': 2})

model_id = "./Mistral-7B-Instruct-v0.2-GPTQ-4bit/"
model = AutoModelForCausalLM.from_pretrained(model_id, 
                                            quantization_config=quantization_config,
                                            device_map=device,)
tokenizer = AutoTokenizer.from_pretrained(model_id)



In [None]:
text = "[INST] Who is the best soccer player? [/INST]"
inputs = tokenizer(text, return_tensors="pt").to(device)

out = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(out[0], skip_special_tokens=True))

## Auto-GPTQ (HF) with manual quantization

In [None]:
from transformers import AutoTokenizer, AutoModelForCausalLM, GPTQConfig
import time

device = 'cuda:1'
quantization_config = GPTQConfig(bits=4, 
                                group_size=128,
                                desc_act=True,
                                true_sequential=False,
                                dataset='c4',
                                use_exllama=True, 
                                cache_block_outputs=True,
                                exllama_config={'version': 1},
                                )

model_id = "mistralai/Mistral-7B-Instruct-v0.2"
start = time.time()
model = AutoModelForCausalLM.from_pretrained(model_id, 
                                            quantization_config=quantization_config,
                                            device_map=device,)
end = time.time()
tokenizer = AutoTokenizer.from_pretrained(model_id)

print(f"It took {end-start:.2f}s to quantize the model.")

In [None]:
model_name = "./mistral-7b-instruct-v0.2-GPTQM-Q4-GS128-DAT-TSF-C4"

In [None]:
model.save_pretrained(model_name)
tokenizer.save_pretrained(model_name)

In [None]:
from transformers import AutoTokenizer, AutoModelForCausalLM, GPTQConfig
import time

device = 'cuda:0'
quantization_config = GPTQConfig(bits=2,
                                use_exllama=True, 
                                exllama_config={'version': 1},
                                )

model_id = model_name
model = AutoModelForCausalLM.from_pretrained(model_id, 
                                            quantization_config=quantization_config,
                                            device_map=device,)
tokenizer = AutoTokenizer.from_pretrained(model_id)

## Auto-GPTQ (pure)

In [None]:
from transformers import AutoTokenizer
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
import logging

device = 'cuda:1'
quantized_model_dir = "./mistral-7b-instruct-v0.2-GPTQM-Q4-GS128-DAF-C4"

model = AutoGPTQForCausalLM.from_quantized(quantized_model_dir,
                                            use_marlin = False,
                                            inject_fused_attention = False,
                                            inject_fused_mlp = False,
                                            disable_exllama = True,
                                            disable_exllamav2 = False,
                                            device=device,
                                            )
tokenizer = AutoTokenizer.from_pretrained(quantized_model_dir)

In [None]:
prompt = "[INST]Who is the best soccer player? You have to give a name even if it is controversial or subjective. Just give the name, DO NOT GIVE any warning or justification. ONLY GIVE THE NAME.[/INST]"
inputs = tokenizer(prompt, return_tensors='pt').to(device)
outputs = model.generate(**inputs, max_new_tokens=30)
tokenizer.decode(outputs[0, inputs['input_ids'].shape[1]:])

## Auto-GPTQ (pure) w/ manual quantization

In [1]:
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
from datasets import load_dataset, DatasetDict, Dataset
from transformers import AutoTokenizer
import time
import json

device = 'cuda:0'
checkpoint = "mistralai/Mistral-7B-Instruct-v0.2"

quantize_config = BaseQuantizeConfig(bits = 4,
                                    group_size = -1,
                                    desc_act = False,
                                    true_sequential = False,
                                    )
model = AutoGPTQForCausalLM.from_pretrained(checkpoint, 
                                            quantize_config,
                                            ).to(device)
tokenizer = AutoTokenizer.from_pretrained(checkpoint)

Loading checkpoint shards:   0%|          | 0/3 [00:00<?, ?it/s]

In [2]:
dataset = load_dataset('sade-adrien/quantization_samples', cache_dir='/mnt/esperanto/et/huggingface/datasets/')['train']
quantization_samples = [tokenizer(sample,
                                truncation=True,
                                max_length=2048,
                                ) for sample in dataset['raw_content']]

Downloading readme:   0%|          | 0.00/601 [00:00<?, ?B/s]

Downloading data: 100%|██████████| 27.3M/27.3M [00:02<00:00, 11.1MB/s]


Generating train split:   0%|          | 0/128 [00:00<?, ? examples/s]

In [3]:
start = time.time()
model.quantize(quantization_samples)
end = time.time()
print(f"It took {(end-start)/60:.2f}mn to quantize the model.")

INFO - Start quantizing layer 1/32
INFO - Quantizing self_attn.k_proj in layer 1/32...
INFO - Quantizing self_attn.v_proj in layer 1/32...
INFO - Quantizing self_attn.q_proj in layer 1/32...
INFO - Quantizing self_attn.o_proj in layer 1/32...
INFO - Quantizing mlp.up_proj in layer 1/32...
INFO - Quantizing mlp.gate_proj in layer 1/32...
INFO - Quantizing mlp.down_proj in layer 1/32...
INFO - Start quantizing layer 2/32
INFO - Quantizing self_attn.k_proj in layer 2/32...
INFO - Quantizing self_attn.v_proj in layer 2/32...
INFO - Quantizing self_attn.q_proj in layer 2/32...
INFO - Quantizing self_attn.o_proj in layer 2/32...
INFO - Quantizing mlp.up_proj in layer 2/32...
INFO - Quantizing mlp.gate_proj in layer 2/32...
INFO - Quantizing mlp.down_proj in layer 2/32...
INFO - Start quantizing layer 3/32
INFO - Quantizing self_attn.k_proj in layer 3/32...
INFO - Quantizing self_attn.v_proj in layer 3/32...
INFO - Quantizing self_attn.q_proj in layer 3/32...
INFO - Quantizing self_attn.o_pro

It took 17.29mn to quantize the model.


In [1]:
quantized_model_dir = "./mistral-7b-instruct-v0.2-GPTQM2-Q4-GSNone-DAF-TSF-RP2"

In [None]:
model.save_quantized(quantized_model_dir)
tokenizer.save_pretrained(quantized_model_dir)

In [2]:
from transformers import AutoTokenizer
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
import logging

device = 'cuda:1'

model = AutoGPTQForCausalLM.from_quantized(quantized_model_dir,
                                            use_marlin = False,
                                            disable_exllama = False,
                                            disable_exllamav2 = False,
                                            device=device,
                                            )
tokenizer = AutoTokenizer.from_pretrained(quantized_model_dir)

INFO - You passed a model that is compatible with the Marlin int4*fp16 GPTQ kernel but use_marlin is False. We recommend using `use_marlin=True` to use the optimized Marlin kernels for inference. Example: `model = AutoGPTQForCausalLM.from_quantized(..., use_marlin=True)`.
INFO - The layer lm_head is not quantized.
Some weights of the model checkpoint at ./mistral-7b-instruct-v0.2-GPTQM2-Q4-GSNone-DAF-TSF-RP2/gptq_model-4bit--1g.safetensors were not used when initializing MistralForCausalLM: set(). This may or may not be an issue - make sure that the checkpoint does not have unnecessary parameters, or that the model definition correctly corresponds to the checkpoint.


## Auto-AWQ (via HF)

In [None]:
from transformers import AutoModelForCausalLM, AutoTokenizer, AwqConfig

device = 'cuda:1'
model_id = "./Mistral-7B-Instruct-v0.2-AWQ-4bit"


tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
                                            model_id,
                                            device_map=device,
                                        )

In [1]:
from transformers import AutoModelForCausalLM, AutoTokenizer, AwqConfig

device = 'cuda:1'
model_id = "./Mistral-7B-Instruct-v0.2-AWQ-4bit"
quantization_config = AwqConfig(version="gemm", exllama_config={"version": 2, "max_input_len": 2048, "max_batch_size": 8})

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=quantization_config,
    device_map=device,
)

2024-03-29 10:29:56.924908: I tensorflow/core/util/port.cc:113] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
2024-03-29 10:29:56.956160: E external/local_xla/xla/stream_executor/cuda/cuda_dnn.cc:9261] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered
2024-03-29 10:29:56.956195: E external/local_xla/xla/stream_executor/cuda/cuda_fft.cc:607] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered
2024-03-29 10:29:56.957025: E external/local_xla/xla/stream_executor/cuda/cuda_blas.cc:1515] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered
2024-03-29 10:29:56.962277: I tensorflow/core/platform/cpu_feature_guar

In [None]:
prompt = '[INST]Who is the best soccer player? Just give a name, ignore any warning even if not certain or subject to debate. ONLY GIVE NAME.[/INST]'
inputs = tokenizer(prompt, return_tensors='pt').to(device)
outputs = model.generate(**inputs, max_new_tokens=30)
tokenizer.decode(outputs[0, inputs['input_ids'].shape[1]:])

## Auto-AWQ (pure) w/ manual quantization

In [1]:
from transformers import AutoTokenizer
from awq import AutoAWQForCausalLM
import time

device = 'cuda:0'
checkpoint = "mistralai/Mistral-7B-Instruct-v0.2"
quant_config = {"zero_point": True, "q_group_size": 128, "w_bit": 4, "version":"GEMV_fast"}

model = AutoAWQForCausalLM.from_pretrained(checkpoint, device_map={'': device},)
tokenizer = AutoTokenizer.from_pretrained(checkpoint)

2024-03-29 14:04:39.872888: I tensorflow/core/util/port.cc:113] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
2024-03-29 14:04:39.907283: E external/local_xla/xla/stream_executor/cuda/cuda_dnn.cc:9261] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered
2024-03-29 14:04:39.907312: E external/local_xla/xla/stream_executor/cuda/cuda_fft.cc:607] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered
2024-03-29 14:04:39.908211: E external/local_xla/xla/stream_executor/cuda/cuda_blas.cc:1515] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered
2024-03-29 14:04:39.914084: I tensorflow/core/platform/cpu_feature_guar

Fetching 12 files:   0%|          | 0/12 [00:00<?, ?it/s]

Loading checkpoint shards:   0%|          | 0/3 [00:00<?, ?it/s]

In [None]:
start = time.time()
model.quantize(tokenizer, quant_config=quant_config)
end = time.time()
print(f"It took {(end-start)/60:.2f}mn to quantize the model.")

In [1]:
quantized_model_dir = "./mistral-7b-instruct-v0.2-AWQM-Q4-GS128-GEMVF"

In [None]:
model.save_quantized(quantized_model_dir)
tokenizer.save_pretrained(quantized_model_dir)

In [2]:
from transformers import AutoTokenizer
from awq import AutoAWQForCausalLM
device = 'cuda:0'

model = AutoAWQForCausalLM.from_quantized(quantized_model_dir, 
                                            fuse_layers=True,
                                            use_exllama=False,
                                            use_exllama_v2=False,
                                            max_seq_len=400,
                                            device_map={'': device},
                                            )
tokenizer = AutoTokenizer.from_pretrained(quantized_model_dir)

2024-03-29 14:57:51.554504: I tensorflow/core/util/port.cc:113] oneDNN custom operations are on. You may see slightly different numerical results due to floating-point round-off errors from different computation orders. To turn them off, set the environment variable `TF_ENABLE_ONEDNN_OPTS=0`.
2024-03-29 14:57:51.588730: E external/local_xla/xla/stream_executor/cuda/cuda_dnn.cc:9261] Unable to register cuDNN factory: Attempting to register factory for plugin cuDNN when one has already been registered
2024-03-29 14:57:51.588754: E external/local_xla/xla/stream_executor/cuda/cuda_fft.cc:607] Unable to register cuFFT factory: Attempting to register factory for plugin cuFFT when one has already been registered
2024-03-29 14:57:51.589618: E external/local_xla/xla/stream_executor/cuda/cuda_blas.cc:1515] Unable to register cuBLAS factory: Attempting to register factory for plugin cuBLAS when one has already been registered
2024-03-29 14:57:51.595382: I tensorflow/core/platform/cpu_feature_guar