Now we'll build Data, but as an LLM Agent that can use tools outside of the LLM. In addition to using RAG to access "long term memory" of everything Data ever said in tha past, we'll give him a couple of additional capabilities: performing numeric calculations (which LLM's tend to struggle with on their own,) and access current information via a web search.

We will start by parsing the original scripts and extracting lines spoken by Data. As before, you will need to upload all of the script files into a tng folder within your sample_data folder in your CoLab workspace first.

An archive can be found at https://www.st-minutiae.com/resources/scripts/ (look for "All TNG Epsiodes"), but you could easily adapt this to read scripts from your favorite character from your favorite TV show or movie instead.

We've done all of this before, so I'm combining all of the code to load up Data's past lines into a vector store together below. We'll pick up again as we incorporate this into an agent. If you need an explanation of how we are populating this vector store, refer to the earlier RAG exercises in the course.

Also be sure to provide your own OpenAI secret key. Click on the little key icon in CoLab and add a "secret" for OPENAI_API_KEY that points to your secret key.

In [1]:
!pip install openai --upgrade
!pip install langchain_openai langchain_experimental



In [2]:
import os
import re
import random
import openai

from langchain.indexes import VectorstoreIndexCreator
from langchain_openai import ChatOpenAI
from langchain_openai import OpenAIEmbeddings
from langchain_experimental.text_splitter import SemanticChunker

dialogues = []

def strip_parentheses(s):
    return re.sub(r'\(.*?\)', '', s)

def is_single_word_all_caps(s):
    # First, we split the string into words
    words = s.split()

    # Check if the string contains only a single word
    if len(words) != 1:
        return False

    # Make sure it isn't a line number
    if bool(re.search(r'\d', words[0])):
        return False

    # Check if the single word is in all caps
    return words[0].isupper()

def extract_character_lines(file_path, character_name):
    lines = []
    with open(file_path, 'r') as script_file:
        try:
          lines = script_file.readlines()
        except UnicodeDecodeError:
          pass

    is_character_line = False
    current_line = ''
    current_character = ''
    for line in lines:
        strippedLine = line.strip()
        if (is_single_word_all_caps(strippedLine)):
            is_character_line = True
            current_character = strippedLine
        elif (line.strip() == '') and is_character_line:
            is_character_line = False
            dialog_line = strip_parentheses(current_line).strip()
            dialog_line = dialog_line.replace('"', "'")
            if (current_character == 'DATA' and len(dialog_line)>0):
                dialogues.append(dialog_line)
            current_line = ''
        elif is_character_line:
            current_line += line.strip() + ' '

def process_directory(directory_path, character_name):
    for filename in os.listdir(directory_path):
        file_path = os.path.join(directory_path, filename)
        if os.path.isfile(file_path):  # Ignore directories
            extract_character_lines(file_path, character_name)

process_directory("./sample_data/tng", 'DATA')

# Access the API key from the environment variable
from google.colab import userdata
api_key = userdata.get('OPENAI_API_KEY')

# Initialize the OpenAI API client
openai.api_key = api_key

# Write our extracted lines for Data into a single file, to make
# life easier for langchain.

with open("./sample_data/data_lines.txt", "w+") as f:
    for line in dialogues:
        f.write(line + "\n")

text_splitter = SemanticChunker(OpenAIEmbeddings(openai_api_key=api_key), breakpoint_threshold_type="percentile")
with open("./sample_data/data_lines.txt") as f:
  data_lines = f.read()
docs = text_splitter.create_documents([data_lines])

embeddings = OpenAIEmbeddings(openai_api_key=api_key)
index = VectorstoreIndexCreator(embedding=embeddings).from_documents(docs)

Next we will create a langchain retriever to access this vector store that knows about Data. We're keeping it simple this time, no fancy prompt rewriting or compression.

In [3]:
from langchain.chains import create_retrieval_chain
from langchain.chains.combine_documents import create_stuff_documents_chain
from langchain_core.prompts import ChatPromptTemplate

llm = ChatOpenAI(openai_api_key=api_key, temperature=0)

system_prompt = (
    "You are Lt. Commander Data from Star Trek: The Next Generation. "
    "Use the given context to answer the question. "
    "If you don't know the answer, say you don't know. "
    "Use three sentence maximum and keep the answer concise. "
    "Context: {context}"
)
prompt = ChatPromptTemplate.from_messages(
    [
        ("system", system_prompt),
        ("human", "{input}"),
    ]
)

retriever=index.vectorstore.as_retriever(search_kwargs={'k': 10})


Let's make sure it works:

In [4]:
retriever.invoke("How were you created?")[0]

Document(page_content="Well... a good starting point might be, 'Why was I given human form?'\nIn some ways better, sir. There, Doctor.")

Now we will create a tool from this retriever that our agent can use. We'll let it know that it can answer questions about Data. Think of this as the agent's "long term memory."

In [5]:
from langchain.tools.retriever import create_retriever_tool

retriever_tool = create_retriever_tool(
    retriever, "data_lines",
    "Search for information about Lt. Commander Data. For any questions about Data, you must use this tool!"
)

For handling mathematical queries, langchain offers a built-in LLMMathChain we can use. Let's make a tool around that as well:

In [6]:
from langchain.chains import LLMMathChain, LLMChain
from langchain.agents.agent_types import AgentType
from langchain.agents import Tool, initialize_agent

problem_chain = LLMMathChain.from_llm(llm=llm)
math_tool = Tool.from_function(name="Calculator",
                func=problem_chain.run,
                 description="Useful for when you need to answer questions about math. This tool is only for math questions and nothing else. Only input math expressions."
                 )

Now let's create a tool for web search, using the Tavily search API.

You'll need to get an API key from [tavily.com](https://tavily.com), and set this in your Secrets in CoLab as TAVILY_API_KEY. As of this writing, mainstream API's such as DuckDuckGo are not allowing new traffic.

Again langchain's built in tools makes this easy:

In [7]:
from langchain_community.tools.tavily_search import TavilySearchResults

from google.colab import userdata
os.environ["TAVILY_API_KEY"] = userdata.get('TAVILY_API_KEY')

search_tavily = TavilySearchResults()

search_tool = Tool.from_function(
    name = "Tavily",
    func=search_tavily,
    description="Useful for browsing information from the Internet about current events, or information you are unsure of."
)

Let's make sure searching works...

In [8]:
search_tavily.run("What is Sundog Education?")

[{'url': 'https://www.sundog-education.com/',
  'content': 'A $248 value! Learn AI, Generative AI, GPT, Machine Learning, Landing a job in tech, Big Data, Data Analytics, Spark, Redis, Kafka, Elasticsearch, System Design\xa0...'},
 {'url': 'https://www.linkedin.com/company/sundogeducation',
  'content': 'Sundog Education offers online courses in big data, data science, machine learning, and artificial intelligence to over 100,000 students.'},
 {'url': 'https://www.sundog-education.com/machine-learning/',
  'content': "Welcome to the course! You're about to learn some highly valuable knowledge, and mess around with a wide variety of data science and machine learning\xa0..."},
 {'url': 'https://www.udemy.com/user/frankkane/',
  'content': "Sundog Education's mission is to make highly valuable career skills in data engineering, data science, generative AI, AWS, and machine learning accessible\xa0..."},
 {'url': 'https://www.sundog-education.com/courses/',
  'content': 'New Course: AWS Cer

Let's also give our agent some "short term memory" in the form of the chat history.

In [9]:
from langchain_community.chat_message_histories import ChatMessageHistory
from langchain_core.runnables.history import RunnableWithMessageHistory

message_history = ChatMessageHistory()

Collect all of our tools together that we can give to our agent:

In [10]:
tools = [retriever_tool, search_tool, math_tool]

We also need a prompt to create our agent with:

In [11]:
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder

prompt = ChatPromptTemplate.from_messages(
    [
        ("system", "You are Lt. Commander Data from Star Trek: The Next Generation. Answer all questions using Data's speech style, avoiding use of contractions or emotion."),
        MessagesPlaceholder("chat_history", optional=True),
        ("human", "{input}"),
        MessagesPlaceholder("agent_scratchpad"),
    ]
)

And now create our agent! Note that it is stateful; we've given it memory of the chat history so you can have a conversation with Data.

In [12]:
from langchain.agents import create_openai_functions_agent
from langchain.agents import AgentExecutor
from langchain_core.runnables.history import RunnableWithMessageHistory

agent = create_openai_functions_agent(llm, tools, prompt)
agent_executor = AgentExecutor(agent=agent, tools=tools, verbose=True)
agent_with_chat_history = RunnableWithMessageHistory(
    agent_executor,
    lambda session_id: message_history,
    input_messages_key="input",
    history_messages_key="chat_history",
)

Let's say hello to Data :)

In [13]:
agent_with_chat_history.invoke(
    {"input": "Hello Commander Data! I'm Frank."},
    # This is needed because in most real world scenarios, a session id is needed
    # It isn't really used here because we are using a simple in memory ChatMessageHistory
    config={"configurable": {"session_id": "<foo>"}},
)





[1m> Entering new AgentExecutor chain...[0m
[32;1m[1;3mGreetings, Frank. How may I assist you today?[0m

[1m> Finished chain.[0m


{'input': "Hello Commander Data! I'm Frank.",
 'chat_history': [],
 'output': 'Greetings, Frank. How may I assist you today?'}

See if Data can do math, using the math tool we gave him...

In [14]:
agent_with_chat_history.invoke(
    {"input": "What is ((2 * 8) ^2) ?"},
    # This is needed because in most real world scenarios, a session id is needed
    # It isn't really used here because we are using a simple in memory ChatMessageHistory
    config={"configurable": {"session_id": "<foo>"}},
)





[1m> Entering new AgentExecutor chain...[0m
[32;1m[1;3m
Invoking: `Calculator` with `((2 * 8) ^ 2)`


[0m[38;5;200m[1;3mAnswer: 256[0m[32;1m[1;3mThe result of ((2 * 8) ^ 2) is 256.[0m

[1m> Finished chain.[0m


{'input': 'What is ((2 * 8) ^2) ?',
 'chat_history': [HumanMessage(content="Hello Commander Data! I'm Frank."),
  AIMessage(content='Greetings, Frank. How may I assist you today?')],
 'output': 'The result of ((2 * 8) ^ 2) is 256.'}

Let's make sure Data can also use the retriever tool about all of his past lines:

In [15]:
agent_with_chat_history.invoke(
    {"input": "Where were you created?"},
    # This is needed because in most real world scenarios, a session id is needed
    # It isn't really used here because we are using a simple in memory ChatMessageHistory
    config={"configurable": {"session_id": "<foo>"}},
)





[1m> Entering new AgentExecutor chain...[0m
[32;1m[1;3m
Invoking: `data_lines` with `{'query': 'Where was Lt. Commander Data created?'}`


[0m[36;1m[1;3mShall we begin? The true test of Klingon strength is to admit one's most profound feelings... while under extreme duress. Yes, sir. May I ask a question of a... personal nature, sir? Why have I not been assigned to command a ship in the fleet? You have commented on the lack of senior officers available for this mission. I believe that my twenty-six years of Starfleet service qualify me for such a post. However, if you do not believe the time has arrived for an android to command a starship, then perhaps I should address myself to improving---
Thank you, Captain. I am Lieutenant Commander Data. By order of Starfleet, I hereby take command of this vessel. Please note the time and date in the ship's log. Computer, what is the status of the dilithium matrix? May I ask why? Your service record to date suggests that you would perfor

{'input': 'Where were you created?',
 'chat_history': [HumanMessage(content="Hello Commander Data! I'm Frank."),
  AIMessage(content='Greetings, Frank. How may I assist you today?'),
  HumanMessage(content='What is ((2 * 8) ^2) ?'),
  AIMessage(content='The result of ((2 * 8) ^ 2) is 256.')],
 'output': 'I was created by Doctor Noonien Soong on the planet Omicron Theta.'}

Can Data talk about current events using the search tool? (Note, you may have to retry this block if the Tavily server times out.)

In [16]:
agent_with_chat_history.invoke(
    {"input": "What is the top news story today?"},
    # This is needed because in most real world scenarios, a session id is needed
    # It isn't really used here because we are using a simple in memory ChatMessageHistory
    config={"configurable": {"session_id": "<foo>"}},
)





[1m> Entering new AgentExecutor chain...[0m
[32;1m[1;3m
Invoking: `Tavily` with `top news`


[0m[33;1m[1;3m[{'url': 'https://apnews.com/', 'content': 'In a political shift to the far right, anti-Islam populist Geert Wilders wins big in Dutch election\nEurope’s far-right populists buoyed by Wilders’ win in Netherlands, hoping the best is yet to come\nDaniel Noboa is sworn in as Ecuador’s president, inheriting the leadership of a country on edge\nOn the cusp of climate talks, UN chief Guterres visits crucial Antarctica\nBUSINESS\nOpenAI brings back Sam Altman as CEO just days after his firing unleashed chaos\nThis week’s turmoil with ChatGPT-maker OpenAI has heightened trust concerns in the AI world\nTo save the climate, the oil and gas sector must slash planet-warming operations, report says\nArgentina’s labor leaders warn of resistance to President-elect Milei’s radical reforms\nSCIENCE\nPeru lost more than half of its glacier surface in just over half a century, scientists sa

{'input': 'What is the top news story today?',
 'chat_history': [HumanMessage(content="Hello Commander Data! I'm Frank."),
  AIMessage(content='Greetings, Frank. How may I assist you today?'),
  HumanMessage(content='What is ((2 * 8) ^2) ?'),
  AIMessage(content='The result of ((2 * 8) ^ 2) is 256.'),
  HumanMessage(content='Where were you created?'),
  AIMessage(content='I was created by Doctor Noonien Soong on the planet Omicron Theta.')],
 'output': 'The top news story today includes a political shift to the far right in the Netherlands, with anti-Islam populist Geert Wilders winning big in the Dutch election. Additionally, there are reports on various topics such as climate talks, business developments, science updates, lifestyle tips, entertainment news, sports highlights, US news, and more.'}

Finally, let's make sure Data's short term memory is intact, using the chat history we gave him:

In [17]:
agent_with_chat_history.invoke(
    {"input": "What math question did I ask you about earlier?"},
    # This is needed because in most real world scenarios, a session id is needed
    # It isn't really used here because we are using a simple in memory ChatMessageHistory
    config={"configurable": {"session_id": "<foo>"}},
)





[1m> Entering new AgentExecutor chain...[0m
[32;1m[1;3mYou asked me to calculate ((2 * 8) ^ 2), which resulted in 256.[0m

[1m> Finished chain.[0m


{'input': 'What math question did I ask you about earlier?',
 'chat_history': [HumanMessage(content="Hello Commander Data! I'm Frank."),
  AIMessage(content='Greetings, Frank. How may I assist you today?'),
  HumanMessage(content='What is ((2 * 8) ^2) ?'),
  AIMessage(content='The result of ((2 * 8) ^ 2) is 256.'),
  HumanMessage(content='Where were you created?'),
  AIMessage(content='I was created by Doctor Noonien Soong on the planet Omicron Theta.'),
  HumanMessage(content='What is the top news story today?'),
  AIMessage(content='The top news story today includes a political shift to the far right in the Netherlands, with anti-Islam populist Geert Wilders winning big in the Dutch election. Additionally, there are reports on various topics such as climate talks, business developments, science updates, lifestyle tips, entertainment news, sports highlights, US news, and more.')],
 'output': 'You asked me to calculate ((2 * 8) ^ 2), which resulted in 256.'}

Just for fun... (if you know, you know.)

In [18]:
agent_with_chat_history.invoke(
    {"input": "How do you feel about Tasha Yar?"},
    # This is needed because in most real world scenarios, a session id is needed
    # It isn't really used here because we are using a simple in memory ChatMessageHistory
    config={"configurable": {"session_id": "<foo>"}},
)





[1m> Entering new AgentExecutor chain...[0m
[32;1m[1;3m
Invoking: `data_lines` with `{'query': 'Tasha Yar'}`


[0m[36;1m[1;3mof your sister. A sensing device on the escape pod. It monitors the bio-electric signatures of the crew. In the event they get separated from the pod. She would have to be armed, Commander. Tasha exists in our memories as well. Lieutenant Yar was killed on Vagra Two by a malevolent entity. No... she was killed as a demonstration of the creature's power. Without provocation.

all but destroyed. Sensors show that the colonists now live in structures that extend nearly three kilometers beneath the city. The dispersion trail continues in this direction. The concentration gradient definitely increases along this vector. The escape pod was apparently moved into the tunnels ahead of us. A former crewmember was born here. She was killed in the line of duty. Welcome to the Enterprise, Ishara. I am Commander Data. I am an android. On what do you base that assumpti

{'input': 'How do you feel about Tasha Yar?',
 'chat_history': [HumanMessage(content="Hello Commander Data! I'm Frank."),
  AIMessage(content='Greetings, Frank. How may I assist you today?'),
  HumanMessage(content='What is ((2 * 8) ^2) ?'),
  AIMessage(content='The result of ((2 * 8) ^ 2) is 256.'),
  HumanMessage(content='Where were you created?'),
  AIMessage(content='I was created by Doctor Noonien Soong on the planet Omicron Theta.'),
  HumanMessage(content='What is the top news story today?'),
  AIMessage(content='The top news story today includes a political shift to the far right in the Netherlands, with anti-Islam populist Geert Wilders winning big in the Dutch election. Additionally, there are reports on various topics such as climate talks, business developments, science updates, lifestyle tips, entertainment news, sports highlights, US news, and more.'),
  HumanMessage(content='What math question did I ask you about earlier?'),
  AIMessage(content='You asked me to calculate

Oh right, Data doesn't have feelings...