# LangChain: Memory

## Outline
* ConversationBufferMemory
* ConversationBufferWindowMemory
* ConversationTokenBufferMemory
* ConversationSummaryMemory

## ConversationBufferMemory

In [1]:
import os

from dotenv import load_dotenv, find_dotenv
_ = load_dotenv(find_dotenv()) # read local .env file

import warnings
warnings.filterwarnings('ignore')

Note: LLM's do not always produce the same results. When executing the code in your notebook, you may get slightly different answers that those in the lecture.

In [2]:
# account for deprecation of LLM model
import datetime
# Get the current date
current_date = datetime.datetime.now().date()

# Define the date after which the model should be set to "gpt-3.5-turbo"
target_date = datetime.date(2025, 8, 18)

# Set the model variable based on the current date
if current_date > target_date:
    llm_model = "gpt-3.5-turbo"
else:
    llm_model = "gpt-3.5-turbo-0301"

In [3]:
from langchain.chat_models import ChatOpenAI
from langchain.chains import ConversationChain
from langchain.memory import ConversationBufferMemory


### We are going to create a chatbot application using LangChain to learn Memory concept in LangChain

In [9]:
llm = ChatOpenAI(temperature=0.0, model=llm_model)
# We are setting up one specific type of memory here. 
memory = ConversationBufferMemory()
# Don't worry about Conversation Chain for now, we will learn this later 
conversation = ConversationChain(
    llm=llm, 
    memory = memory,
    verbose=True # try with False
)

In [10]:
conversation.predict(input="Hi, my name is Ankit")



[1m> Entering new ConversationChain chain...[0m
Prompt after formatting:
[32;1m[1;3mThe following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:

Human: Hi, my name is Ankit
AI:[0m

[1m> Finished chain.[0m


"Hello Ankit! It's nice to meet you. How can I assist you today?"

In [11]:
conversation.predict(input="What is 1+1?")



[1m> Entering new ConversationChain chain...[0m
Prompt after formatting:
[32;1m[1;3mThe following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:
Human: Hi, my name is Ankit
AI: Hello Ankit! It's nice to meet you. How can I assist you today?
Human: What is 1+1?
AI:[0m

[1m> Finished chain.[0m


'1 + 1 equals 2. Is there anything else you would like to know?'

In [12]:
conversation.predict(input="What is my name?")



[1m> Entering new ConversationChain chain...[0m
Prompt after formatting:
[32;1m[1;3mThe following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:
Human: Hi, my name is Ankit
AI: Hello Ankit! It's nice to meet you. How can I assist you today?
Human: What is 1+1?
AI: 1 + 1 equals 2. Is there anything else you would like to know?
Human: What is my name?
AI:[0m

[1m> Finished chain.[0m


'Your name is Ankit.'

In [13]:
print(memory.buffer)

Human: Hi, my name is Ankit
AI: Hello Ankit! It's nice to meet you. How can I assist you today?
Human: What is 1+1?
AI: 1 + 1 equals 2. Is there anything else you would like to know?
Human: What is my name?
AI: Your name is Ankit.


In [15]:
# This is what LangChain has remembered in its memoery so far. 
# Including what Human or AI has said. 
memory.load_memory_variables({})

{'history': "Human: Hi, my name is Ankit\nAI: Hello Ankit! It's nice to meet you. How can I assist you today?\nHuman: What is 1+1?\nAI: 1 + 1 equals 2. Is there anything else you would like to know?\nHuman: What is my name?\nAI: Your name is Ankit."}

### Task: run the code till here. 

In [16]:
memory = ConversationBufferMemory()

In [17]:
# this is how we can add more context to the memory 
memory.save_context({"input": "Hi"}, 
                    {"output": "What's up"})

In [18]:
print(memory.buffer)

Human: Hi
AI: What's up


In [19]:
memory.load_memory_variables({})

{'history': "Human: Hi\nAI: What's up"}

In [None]:
# adding more stuff in the memory 
memory.save_context({"input": "Not much, just hanging"}, 
                    {"output": "Cool"})

In [21]:
memory.load_memory_variables({})

{'history': "Human: Hi\nAI: What's up\nHuman: Not much, just hanging\nAI: Cool"}

### So when we use LLM for a conversation (go to slide #10)
- The LLM is stateless -- Meaning they don't remember the conversation you have done so far. 
- And Each transation in API endpoint is idependent 
- Chatbots appear to have memory only because there is the wrapper code which provides the full conversation that you have had so far as a `context` 
- Memory is used as additional context to the LLM such that it can generate the output as if it looks like a next message in the conversation
- As the conversation gets longer, required memory needs also increases - Thus, the cost of sending a lot of tokens to the LLM also increases significantly. 
- So, langchain provides a several kind of memory to store and accumulate the converations 

## ConversationBufferWindowMemory

In [22]:
# new type of memory 
from langchain.memory import ConversationBufferWindowMemory

In [None]:
# k=1 means wanted to remember just one conversation exchange
# This prevents the memory to grow without limits
# In practice, the k value would be set to a bit larger number
memory = ConversationBufferWindowMemory(k=1)               

In [25]:
memory.save_context({"input": "Hi"},
                    {"output": "What's up"})
memory.save_context({"input": "Not much, just hanging"},
                    {"output": "Cool"})

In [26]:
memory.load_memory_variables({})

{'history': 'Human: Not much, just hanging\nAI: Cool'}

In [None]:
llm = ChatOpenAI(temperature=0.0, model=llm_model)
memory = ConversationBufferWindowMemory(k=1)
conversation = ConversationChain(
    llm=llm, 
    memory = memory,
    verbose=False # Try changing this to True and re-run this snippet
)

In [28]:
conversation.predict(input="Hi, my name is Ankit")

"Hello Ankit! It's nice to meet you. How can I assist you today?"

In [29]:
conversation.predict(input="What is 1+1?")

'1+1 equals 2. Is there anything else you would like to know?'

In [30]:
conversation.predict(input="What is my name?")

"I'm sorry, I do not have access to personal information such as your name. Is there anything else you would like to know?"

## ConversationTokenBufferMemory

The memory will limit the number of tokens to be stored in memory

In [None]:
#!pip install tiktoken

In [33]:
from langchain.memory import ConversationTokenBufferMemory
from langchain.llms import OpenAI
llm = ChatOpenAI(temperature=0.0, model=llm_model)

In [None]:
# I would advise you to play with max_token_limit param to understand how a small number like 25 or 30 
# is different than a big number like 100 or 200
# check how the output differs

# important note: see llm is an argument here because different LLMs would have a different tokenizers 
memory = ConversationTokenBufferMemory(llm=llm, max_token_limit=50)
memory.save_context({"input": "AI is what?!"},
                    {"output": "Amazing!"})
memory.save_context({"input": "Backpropagation is what?"},
                    {"output": "Beautiful!"})
memory.save_context({"input": "Chatbots are what?"}, 
                    {"output": "Charming!"})

In [35]:
memory.load_memory_variables({})

{'history': 'AI: Amazing!\nHuman: Backpropagation is what?\nAI: Beautiful!\nHuman: Chatbots are what?\nAI: Charming!'}

## ConversationSummaryMemory
Use LLM to write the summary of the conversation so far and let that be the memory

In [38]:
from langchain.memory import ConversationSummaryBufferMemory

In [46]:
# create a long string
schedule = "There is a meeting at 8am with your product team. \
You will need your powerpoint presentation prepared. \
9am-12pm have time to work on your LangChain \
project which will go quickly because Langchain is such a powerful tool. \
At Noon, lunch at the italian resturant with a customer who is driving \
from over an hour away to meet you to understand the latest in AI. \
Be sure to bring your laptop to show the latest LLM demo."

# depending upon the max_token_limit it will decide if it has to
# store the entire conversation? or summarise the conversation using the 
# given LLM and store only the summary value. 
# check with max_token_limit = 400 and then with 100 or 50 
# see the difference in outputs in the `memory.load_memory_variables({})`
memory = ConversationSummaryBufferMemory(llm=llm, max_token_limit=50)
memory.save_context({"input": "Hello"}, {"output": "What's up"})
memory.save_context({"input": "Not much, just hanging"},
                    {"output": "Cool"})
memory.save_context({"input": "What is on the schedule today?"}, 
                    {"output": f"{schedule}"})

In [47]:
memory.load_memory_variables({})

{'history': "System: The human and AI exchange greetings and discuss the schedule for the day. The AI informs the human of a morning meeting with the product team, work on the LangChain project, and a lunch meeting with a customer interested in AI. The AI emphasizes the importance of being prepared for the day's events."}

In [48]:
conversation = ConversationChain(
    llm=llm, 
    memory = memory,
    verbose=True
)

In [49]:
conversation.predict(input="What would be a good demo to show?")



[1m> Entering new ConversationChain chain...[0m
Prompt after formatting:
[32;1m[1;3mThe following is a friendly conversation between a human and an AI. The AI is talkative and provides lots of specific details from its context. If the AI does not know the answer to a question, it truthfully says it does not know.

Current conversation:
System: The human and AI exchange greetings and discuss the schedule for the day. The AI informs the human of a morning meeting with the product team, work on the LangChain project, and a lunch meeting with a customer interested in AI. The AI emphasizes the importance of being prepared for the day's events.
Human: What would be a good demo to show?
AI:[0m

[1m> Finished chain.[0m


"For the morning meeting with the product team, a demo showcasing the latest features and updates on the LangChain project would be ideal. This could include a live demonstration of the language processing capabilities, data encryption features, and any new integrations that have been implemented. Additionally, for the lunch meeting with the customer interested in AI, a demo highlighting the AI algorithms and machine learning models used in our products could be impressive. It's important to tailor the demos to the specific audience and their interests to make a strong impact."

In [50]:
memory.load_memory_variables({})

{'history': "System: The human and AI exchange greetings and discuss the schedule for the day. The AI informs the human of a morning meeting with the product team, work on the LangChain project, and a lunch meeting with a customer interested in AI. The AI emphasizes the importance of being prepared for the day's events, suggesting tailored demos for the product team and customer meetings. For the product team, showcasing the latest features and updates on LangChain is recommended, while highlighting AI algorithms and machine learning models for the customer meeting."}