In [1]:
!pip install flash-attn transformers accelerate

import time
from datetime import timedelta

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
from transformers.utils import is_flash_attn_2_available

torch.random.manual_seed(0)

model = AutoModelForCausalLM.from_pretrained(
    "microsoft/Phi-3-mini-4k-instruct",
    device_map="cuda",
    torch_dtype="auto",
    trust_remote_code=True,
    # attn_implementation="flash_attention_2",
).to("cuda")
tokenizer = AutoTokenizer.from_pretrained("microsoft/Phi-3-mini-4k-instruct")
streamer = TextStreamer(tokenizer, skip_prompt=True)

print("flash_attn_2 available:", is_flash_attn_2_available())



Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]

Special tokens have been added in the vocabulary, make sure the associated word embeddings are fine-tuned or trained.


flash_attn_2 available: True


In [2]:
def gen(text, preview=True):
    duration_start = time.perf_counter()
    prompt = "<|user|>\n{} <|end|>\n<|assistant|>".format(text)
    tokens = tokenizer.encode(prompt, return_tensors="pt").to("cuda")
    outputs = model.generate(
        tokens,
        max_new_tokens=1024,
        return_dict_in_generate=True,
        streamer=streamer if preview else None,
    )
    output_tokens = outputs.sequences[0]
    output_gen_tokens = output_tokens[
        len(tokens[0]) : -1
    ]  # From just after prompt to just before <|end|> token
    output_string = tokenizer.decode(output_gen_tokens)
    duration_seconds = time.perf_counter() - duration_start
    print(
        "== took {} ({} toks: {}/tok; {} tps) ==".format(
            timedelta(seconds=duration_seconds),
            len(output_gen_tokens),
            timedelta(seconds=duration_seconds / len(output_gen_tokens)),
            len(output_gen_tokens) / duration_seconds,
        )
    )
    del tokens, outputs, output_tokens, output_gen_tokens
    return output_string


gen("What is the closest star to the Sun?")
gen("What is the difference between hue, saturation, and value?")
gen("Where is Waldo?")

<s><|user|> What is the closest star to the Sun? 

You are not running the flash-attention implementation, expect numerical differences.


<|end|><|assistant|> The closest star to the Sun is Proxima Centauri. It is part of the Alpha Centauri star system, which also includes Alpha Centauri A and Alpha Centauri B. Proxima Centauri is approximately 4.24 light-years away from the Sun. It is a red dwarf star and is the closest known exoplanet host, with at least two confirmed planets, Proxima Centauri b and Proxima Centauri c, orbiting it.<|end|>
== took 0:00:04.800694 (113 toks: 0:00:00.042484/tok; 23.53826175902905 tps) ==
<s><|user|> What is the difference between hue, saturation, and value? <|end|><|assistant|> Hue, saturation, and value are three components of the color model known as HSV (Hue, Saturation, Value), which is also referred to as HSB (Hue, Saturation, Brightness) in some contexts. These components are used to describe and manipulate colors in digital imaging and color theory. Here's a breakdown of each component:

1. Hue: Hue refers to the color itself, or the dominant wavelength of light that determines the 

"I'm unable to assist with that. However, Waldo is a fictional character from a series of children's books, and his location in those stories is always hidden among various illustrations. If you're referring to a different context, please provide more details."

In [3]:
conv = gen(
    "Please generate an example conversation between participant1 and participant2 about frogs. Make it last 20 utterances and have a follow-up. Please label utterance numbers and separate utterances with newlines."
)

<s><|user|> Please generate an example conversation between participant1 and participant2 about frogs. Make it last 20 utterances and have a follow-up. Please label utterance numbers and separate utterances with newlines. <|end|><|assistant|> 1. Participant1: Hey, have you ever been fascinated by frogs?
2. Participant2: Actually, yes! I find them quite interesting. Why do you ask?
3. Participant1: I was just reading about their life cycle and it's quite fascinating.
4. Participant2: Oh, I'd love to hear more about it. What did you learn?
5. Participant1: Well, did you know that frogs start their life as eggs?
6. Participant2: Yes, I've heard about that. The eggs hatch into tadpoles, right?
7. Participant1: Exactly! Tadpoles live in water and breathe through gills.
8. Participant2: That's so cool! And then they undergo metamorphosis to become adult frogs.
9. Participant1: Yes, during metamorphosis, they develop lungs and legs for life on land.
10. Participant2: I've always wondered how 

In [4]:
prompt = """
    Summarize the following conversation between two participants. Include the topic of conversation, 2-3 bullet points of discussion, and any follow-up action items. Output each of these pieces of information as its own section with a markdown heading. When referencing particular text in the conversation, specify the utterance with the utterance number in brackets, [123].
    
    {}
"""
gen(prompt.format(conv))

<s><|user|> Summarize the following conversation between two participants. Include the topic of conversation, 2-3 bullet points of discussion, and any follow-up action items. Output each of these pieces of information as its own section with a markdown heading. When referencing particular text in the conversation, specify the utterance with the utterance number in brackets, [123].
    
    1. Participant1: Hey, have you ever been fascinated by frogs?
2. Participant2: Actually, yes! I find them quite interesting. Why do you ask?
3. Participant1: I was just reading about their life cycle and it's quite fascinating.
4. Participant2: Oh, I'd love to hear more about it. What did you learn?
5. Participant1: Well, did you know that frogs start their life as eggs?
6. Participant2: Yes, I've heard about that. The eggs hatch into tadpoles, right?
7. Participant1: Exactly! Tadpoles live in water and breathe through gills.
8. Participant2: That's so cool! And then they undergo metamorphosis to bec

'# Conversation Topic\n- Participants discussing the fascinating aspects of frogs, including their life cycle, abilities, and communication methods.\n\n# Discussion Points\n- Frogs start their life as eggs, which hatch into tadpoles that live in water and breathe through gills.\n- During metamorphosis, tadpoles develop lungs and legs to adapt to life on land.\n- Frogs have powerful hind legs and elastic tendons that enable them to jump high and far.\n- Frogs communicate using vocalizations, body language, and chemical signals.\n- Some frogs can change their skin color.\n\n# Follow-up Actions\n- Participant1 will research nature reserves and zoos with a good frog exhibit.\n- Participant2 will look into local frog enthusiasts and experts who might be interested in joining the trip.\n- Both participants will reconvene next week to share their findings and plan the trip.'