<a href="https://colab.research.google.com/github/MengOonLee/LLM/blob/main/References/LangChain/ipynb/T02_LangChain.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>

# LangChain

In [None]:
%%bash
pip install --no-cache-dir -qU \
    transformers langchain langgraph langchain-core langchain-community \
    langchain-huggingface huggingface_hub

## Download LLM

In [None]:
import os
import warnings
warnings.filterwarnings("ignore")
import dotenv
import huggingface_hub
import torch
import transformers
import time

_ = dotenv.load_dotenv(dotenv_path=".env", override=True)
huggingface_hub.login(token=os.environ["HF_TOKEN"])

model_name = "google/gemma-3-1b-it"
local_path = "./models"
os.makedirs(local_path, exist_ok=True)

start_time = time.time()
model = transformers.AutoModelForCausalLM.from_pretrained(
    pretrained_model_name_or_path=model_name,
    token=os.environ["HF_TOKEN"],
    device_map="auto", dtype=torch.bfloat16
)
model.save_pretrained(save_directory=local_path)

tokenizer = transformers.AutoTokenizer.from_pretrained(
    pretrained_model_name_or_path=model_name,
    token=os.environ["HF_TOKEN"]
)
tokenizer.save_pretrained(save_directory=local_path)
end_time = time.time() - start_time

print("Time taken: %.2f"%(end_time))

## Import LLM

In [1]:
import warnings
warnings.filterwarnings("ignore")
import langchain_huggingface

# Initialize LLM model
local_dir = "./models"
llm = langchain_huggingface.HuggingFacePipeline.from_model_id(
    task="text-generation", model_id=local_dir, pipeline_kwargs={
        "max_new_tokens": 128, "temperature": 0.3, "top_p": 0.3
    })
llm = langchain_huggingface.ChatHuggingFace(llm=llm)

ValueError: Could not import transformers python package. Please install it with `pip install transformers`.

## Prompt Template

In [None]:
from langchain_core.prompts import ChatPromptTemplate
from time import time

# Create a prompt template
prompt = ChatPromptTemplate.from_messages(messages=[
    ("system", "You're a helpful assistant that explains concepts clearly."),
    ("human", "Explain {topic} in simple terms.")
])

# Create a chain
chain = prompt | llm

# Run the chain
start_time = time()
result = chain.invoke(input={"topic": "quantum computing"})
end_time = time() - start_time
print("Time taken: %.2f"%(end_time))
print(result.content)

Time taken: 88.82
<bos><start_of_turn>user
You're a helpful assistant that explains concepts clearly.

Explain quantum computing in simple terms.<end_of_turn>
<start_of_turn>model
Okay, letâ€™s break down quantum computing! Itâ€™s a really fascinating and complex topic, but Iâ€™ll do my best to explain it in a way thatâ€™s easy to understand.

**The Basic Idea: Classical vs. Quantum**

Think about regular computers like the ones youâ€™re using now. They store information as **bits**. A bit is like a light switch: it can be either **on** (1) or **off** (0). Everything a computer does â€“ from playing games to running spreadsheets â€“ is ultimately built on manipulating these bits.

Quantum computers, on the other hand, use something


In [None]:
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.chat_history import InMemoryChatMessageHistory
from langchain_core.runnables.history import RunnableWithMessageHistory

# Create a prompt template
conversation_prompt = ChatPromptTemplate.from_messages(messages=[
    ("system", """You're a helpful assistant. Use the conversation history to provide
        context-aware responses."""),
     ("placeholder", "{chat_history}"),
     ("human", "{input}")
])

# Base chain
chain = conversation_prompt | llm

# In-memory store per session
store = {}
def get_session_history(session_id: str):
    if session_id not in store:
        store[session_id] = InMemoryChatMessageHistory()
    return store[session_id]

# Chain with memory
chain_with_history = RunnableWithMessageHistory(
    runnable=chain,
    get_session_history=get_session_history,
    input_messages_key="input",
    history_messages_key="chat_history"
)

# Helper function
def chat(message, session_id="default"):
    response = chain_with_history.invoke(
        input={"input": message},
        config={"configurable": {"session_id": session_id}}
    )
    return response.content

# Have a conversation
chat("Hi, my name is Meng")
chat("What's my name?")
print(chat("Tell me a joke about my name."))

<bos><start_of_turn>user
You're a helpful assistant. Use the conversation history to provide
        context-aware responses.

Hi, my name is Meng<end_of_turn>
<start_of_turn>model
Hi Meng! Itâ€™s nice to meet you. Iâ€™m here to help. ðŸ˜Š How can I assist you today?
<bos><start_of_turn>user
You're a helpful assistant. Use the conversation history to provide
        context-aware responses.

Hi, my name is Meng<end_of_turn>
<start_of_turn>model
<bos><start_of_turn>user
You're a helpful assistant. Use the conversation history to provide
        context-aware responses.

Hi, my name is Meng<end_of_turn>
<start_of_turn>model
Hi Meng! Itâ€™s nice to meet you. Iâ€™m here to help. ðŸ˜Š How can I assist you today?<end_of_turn>
<start_of_turn>user
What's my name?<end_of_turn>
<start_of_turn>model
You are Meng. ðŸ˜Š
<bos><start_of_turn>user
You're a helpful assistant. Use the conversation history to provide
        context-aware responses.

Hi, my name is Meng<end_of_turn>
<start_of_turn>model
