# Import HuggingFace models into LangChain

We'll test incorporate models from Huggingfance using the HuggingFaceHub API calls. In this particular example, we'll test out the Flan2B models

Inspiration taken from: https://www.youtube.com/watch?v=VW5LBavIfY4&list=PL8motc6AQftk1Bs42EW45kwYbyJ4jOdiZ&index=5

### Initial loading and running of a HF-hosted model through Langchain

In [26]:
import os
from langchain.llms import HuggingFaceHub
from langchain.chains.conversation.memory import ConversationBufferMemory
from langchain.chains import ConversationChain

Unlike before, we need to use Huggingface tokens to get access its models. We can obtain these tokens freely after registering through HF: https://huggingface.co/docs/hub/security-tokens

In [6]:
os.environ["HUGGINGFACEHUB_API_TOKEN"] = ""

We'll use a couple of different versions of Flan models: UL2 and T5 models. Flan UL2 is based on the T5 architecture. More info at:
- UL2: https://huggingface.co/docs/transformers/model_doc/flan-ul2
- T5: https://huggingface.co/docs/transformers/model_doc/flan-t5

Recall that a higher temperature makes the model output more creative (too high and may hallucinate), and a lower temperature makes the output more deterministic

<u>Note:</u> Documentation on model card recommends to add "device_map='auto'" when running pytorch model. This is attempted by introducing argument under "kwargs"

In [21]:
flan_ul2 = HuggingFaceHub(
    repo_id="google/flan-ul2",
    model_kwargs={
        "temperature":0.1,
        "max_new_tokens":256,
        "device_map":"auto",
        "verbose":"True"
    }
)

flan_t5 = HuggingFaceHub(
    repo_id="google/flan-t5-xl",
    model_kwargs={
        "temperature":0,
        "device_map":"auto",
        "verbose":"True"
    }
)

Load conversation chains as beffore. We'll be using the ConversationBufferMemory

In [46]:
memory = ConversationBufferMemory()

Start the conversatin chain and interact with the LLM Agent as in the "Conversation_Memory" notebook.

In [47]:
conversation = ConversationChain(
    llm=flan_ul2,
    verbose=True,
    memory=memory
)

In [48]:
conversation.predict(input="Hi, I'm El Brian")



[1m> Entering new ConversationChain chain...[0m
Prompt after formatting:
[32;1m[1;3mThe following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:

Human: Hi, I'm El Brian
AI:[0m

[1m> Finished chain.[0m


"Hi, I'm El Brian. What can I help you with?"

It has the same functionality, like memory storage. Nothing's changed, just the LLM being used under the hood

In [49]:
print(conversation.memory.buffer)

Human: Hi, I'm El Brian
AI: Hi, I'm El Brian. What can I help you with?


In [50]:
conversation.predict(input="How are you doing this evening?")



[1m> Entering new ConversationChain chain...[0m
Prompt after formatting:
[32;1m[1;3mThe following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:
Human: Hi, I'm El Brian
AI: Hi, I'm El Brian. What can I help you with?
Human: How are you doing this evening?
AI:[0m

[1m> Finished chain.[0m


"I'm good. How about you?"

In [52]:
conversation.predict(input="I need some help with travel plans within the US in the summer")



[1m> Entering new ConversationChain chain...[0m
Prompt after formatting:
[32;1m[1;3mThe following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:
Human: Hi, I'm El Brian
AI: Hi, I'm El Brian. What can I help you with?
Human: How are you doing this evening?
AI: I'm good. How about you?
Human: I need some help with travel plans within the US in the summer
AI:[0m

[1m> Finished chain.[0m


'I can help you with that. What city are you interested in visiting?'

In [53]:
conversation.predict(input="The midwest, any recommendations for a low budget travel?")



[1m> Entering new ConversationChain chain...[0m
Prompt after formatting:
[32;1m[1;3mThe following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:
Human: Hi, I'm El Brian
AI: Hi, I'm El Brian. What can I help you with?
Human: How are you doing this evening?
AI: I'm good. How about you?
Human: I need some help with travel plans within the US in the summer
AI: I can help you with that. What city are you interested in visiting?
Human: The midwest, any recommendations for a low budget travel?
AI:[0m

[1m> Finished chain.[0m


'I can help you with that. What city are you interested in visiting?'

In [54]:
conversation.predict(input="I'm interested in travelling to North Carolina, any recommendations on where to stay?")



[1m> Entering new ConversationChain chain...[0m
Prompt after formatting:
[32;1m[1;3mThe following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:
Human: Hi, I'm El Brian
AI: Hi, I'm El Brian. What can I help you with?
Human: How are you doing this evening?
AI: I'm good. How about you?
Human: I need some help with travel plans within the US in the summer
AI: I can help you with that. What city are you interested in visiting?
Human: The midwest, any recommendations for a low budget travel?
AI: I can help you with that. What city are you interested in visiting?
Human: I'm interested in travelling to North Carolina, any recommendations on where to stay?
AI:[0m

[1m> Finished chain.[0m


'I can help you with that. What city are you interested in visiting?'

Similar responses to what we got before but more limited than gpt-3.5-turbo. For the purposes of this exercise, that is fine

### Token counting
Let's consider for a moment that we are intersted on learning more about the number of tokens being used by the tokenizer in this model

In [55]:
from transformers import AutoTokenizer

Loading the tokenizer for Flan UL2

In [56]:
tokenizer = AutoTokenizer.from_pretrained("google/flan-ul2")

Downloading (…)okenizer_config.json: 100%|██████████| 2.35k/2.35k [00:00<00:00, 1.46MB/s]
Downloading spiece.model: 100%|██████████| 792k/792k [00:00<00:00, 7.61MB/s]
Downloading (…)/main/tokenizer.json: 100%|██████████| 2.43M/2.43M [00:00<00:00, 16.1MB/s]
Downloading (…)in/added_tokens.json: 100%|██████████| 665/665 [00:00<00:00, 3.39MB/s]
Downloading (…)cial_tokens_map.json: 100%|██████████| 2.20k/2.20k [00:00<00:00, 7.59MB/s]


Tokenizing a sample input, we can see the input IDs used in the model

In [58]:
test_input = "Hi, I'm El Brian"

print(tokenizer([test_input]))

{'input_ids': [[2018, 6, 27, 31, 51, 1289, 7798, 1]], 'attention_mask': [[1, 1, 1, 1, 1, 1, 1, 1]]}


We can also take a look at the actual token break-down. This makes it easy for us to count the number of tokens in our input sentence 

In [59]:
print(tokenizer.tokenize(test_input))

['▁Hi', ',', '▁I', "'", 'm', '▁El', '▁Brian']


Now we can look at the memory, the text that is passed into the "History" portion of our LLM calls. So that now we can try to count at the tokens in it

In [60]:
memory.buffer

"Human: Hi, I'm El Brian\nAI: Hi, I'm El Brian. What can I help you with?\nHuman: How are you doing this evening?\nAI: I'm good. How about you?\nHuman: I need some help with travel plans within the US in the summer\nAI: I can help you with that. What city are you interested in visiting?\nHuman: The midwest, any recommendations for a low budget travel?\nAI: I can help you with that. What city are you interested in visiting?\nHuman: I'm interested in travelling to North Carolina, any recommendations on where to stay?\nAI: I can help you with that. What city are you interested in visiting?"

But before we count, we need to consider that there is also an actual prompt being pass into the LLM call, so we need to find a way to include that as well

We'll extract that by formatting the conversation prompt so that it includes the history (memory buffer). "the next input" here is just a dummy text to run the prompt and see what the entirity of text would have been like (this dummy text could be anything)

In [64]:
print(conversation.prompt.format(input="the next input", history=memory.buffer))

The following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:
Human: Hi, I'm El Brian
AI: Hi, I'm El Brian. What can I help you with?
Human: How are you doing this evening?
AI: I'm good. How about you?
Human: I need some help with travel plans within the US in the summer
AI: I can help you with that. What city are you interested in visiting?
Human: The midwest, any recommendations for a low budget travel?
AI: I can help you with that. What city are you interested in visiting?
Human: I'm interested in travelling to North Carolina, any recommendations on where to stay?
AI: I can help you with that. What city are you interested in visiting?
Human: the next input
AI:
