# Multi source query RAG

This guide explains how to create a Retrieval-Augmented Generation (RAG) system for Q&A using multiple data sources:

1. Langchain Documentation: We scrape the documentation using WebBaseLoader and store it in a FAISS vector database.
2. Wikipedia: We retrieve relevant information using WikipediaApiWrapper.
3. arXiv: We fetch research papers using ArxivApiWrapper.

By integrating these sources, we can answer questions using a combination of structured data and a language model.

## Install dependencies

In [40]:
%pip install -qU \
    arxiv \
    wikipedia \
    langchain-community \
    faiss-cpu \
    langchain-openai \
    tiktoken \
    python-dotenv \
    langchainhub \
    langsmith

Note: you may need to restart the kernel to use updated packages.


In [41]:
from langchain_community.tools import WikipediaQueryRun
from langchain_community.utilities import WikipediaAPIWrapper

### Wikipedia API wrapper

In [42]:
api_wrapper=WikipediaAPIWrapper(top_k_results=1,doc_content_chars_max=200)
wiki=WikipediaQueryRun(api_wrapper=api_wrapper)

In [43]:
wiki.name

'wikipedia'

## Langchain documentation
1. Use WebBaseLoader to scrape the Langchain documentation from its website.
2. Store the scraped documents in a FAISS vector database for efficient retrieval.
3. Set up a retriever to query the FAISS database and fetch relevant information.

In [44]:
from langchain_community.document_loaders import WebBaseLoader
from langchain_community.vectorstores import FAISS
from langchain_openai import OpenAIEmbeddings
from langchain_text_splitters import RecursiveCharacterTextSplitter

loader=WebBaseLoader("https://docs.smith.langchain.com/")
docs=loader.load()
documents=RecursiveCharacterTextSplitter(chunk_size=1000,chunk_overlap=200).split_documents(docs)
vectordb=FAISS.from_documents(documents,OpenAIEmbeddings())
retriever=vectordb.as_retriever()
retriever

VectorStoreRetriever(tags=['FAISS', 'OpenAIEmbeddings'], vectorstore=<langchain_community.vectorstores.faiss.FAISS object at 0x7affdb0191b0>)

- Create retriever search tool from the langsmith retriever

In [45]:
from langchain.tools.retriever import create_retriever_tool
retriever_tool=create_retriever_tool(retriever,"langsmith_search",
                      "Search for information about LangSmith. For any questions about LangSmith, you must use this tool!")

In [46]:
retriever_tool.name

'langsmith_search'

### Arxiv Api Wrapper

In [47]:
## Arxiv Tool
from langchain_community.utilities import ArxivAPIWrapper
from langchain_community.tools import ArxivQueryRun

arxiv_wrapper=ArxivAPIWrapper(top_k_results=1, doc_content_chars_max=200)
arxiv=ArxivQueryRun(api_wrapper=arxiv_wrapper)
arxiv.name

'arxiv'

In [48]:
tools=[wiki,arxiv,retriever_tool]

In [49]:
tools

[WikipediaQueryRun(api_wrapper=WikipediaAPIWrapper(wiki_client=<module 'wikipedia' from '/home/smartin/envs/venv/lib/python3.10/site-packages/wikipedia/__init__.py'>, top_k_results=1, lang='en', load_all_available_meta=False, doc_content_chars_max=200)),
 ArxivQueryRun(api_wrapper=ArxivAPIWrapper(arxiv_search=<class 'arxiv.Search'>, arxiv_exceptions=(<class 'arxiv.ArxivError'>, <class 'arxiv.UnexpectedEmptyPageError'>, <class 'arxiv.HTTPError'>), top_k_results=1, ARXIV_MAX_QUERY_LENGTH=300, continue_on_failure=False, load_max_docs=100, load_all_available_meta=False, doc_content_chars_max=200, arxiv_result=<class 'arxiv.Result'>)),
 Tool(name='langsmith_search', description='Search for information about LangSmith. For any questions about LangSmith, you must use this tool!', args_schema=<class 'langchain_core.tools.RetrieverInput'>, func=functools.partial(<function _get_relevant_documents at 0x7affe2f3b520>, retriever=VectorStoreRetriever(tags=['FAISS', 'OpenAIEmbeddings'], vectorstore=<

In [50]:
from dotenv import load_dotenv

load_dotenv()
import os
os.environ['OPENAI_API_KEY']=os.getenv("OPENAI_API_KEY")
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-3.5-turbo-0125", temperature=0)

In [51]:

from langchain import hub
# Get the prompt to use - you can modify this!
prompt = hub.pull("hwchase17/openai-functions-agent")
prompt.messages

[SystemMessagePromptTemplate(prompt=PromptTemplate(input_variables=[], template='You are a helpful assistant')),
 MessagesPlaceholder(variable_name='chat_history', optional=True),
 HumanMessagePromptTemplate(prompt=PromptTemplate(input_variables=['input'], template='{input}')),
 MessagesPlaceholder(variable_name='agent_scratchpad')]

In [52]:
### Agents
from langchain.agents import create_openai_tools_agent
agent=create_openai_tools_agent(llm,tools,prompt)

In [53]:
## Agent Executer
from langchain.agents import AgentExecutor
agent_executor=AgentExecutor(agent=agent,tools=tools,verbose=True)
agent_executor

AgentExecutor(verbose=True, agent=RunnableMultiActionAgent(runnable=RunnableAssign(mapper={
  agent_scratchpad: RunnableLambda(lambda x: format_to_openai_tool_messages(x['intermediate_steps']))
})
| ChatPromptTemplate(input_variables=['agent_scratchpad', 'input'], input_types={'chat_history': typing.List[typing.Union[langchain_core.messages.ai.AIMessage, langchain_core.messages.human.HumanMessage, langchain_core.messages.chat.ChatMessage, langchain_core.messages.system.SystemMessage, langchain_core.messages.function.FunctionMessage, langchain_core.messages.tool.ToolMessage]], 'agent_scratchpad': typing.List[typing.Union[langchain_core.messages.ai.AIMessage, langchain_core.messages.human.HumanMessage, langchain_core.messages.chat.ChatMessage, langchain_core.messages.system.SystemMessage, langchain_core.messages.function.FunctionMessage, langchain_core.messages.tool.ToolMessage]]}, metadata={'lc_hub_owner': 'hwchase17', 'lc_hub_repo': 'openai-functions-agent', 'lc_hub_commit_hash': 'a165

In [57]:
agent_executor.invoke({"input":"Tell me about Machine Learning"}, verbose=False)



[1m> Entering new AgentExecutor chain...[0m
[32;1m[1;3m
Invoking: `wikipedia` with `{'query': 'Machine Learning'}`


[0m[36;1m[1;3mPage: Machine learning
Summary: Machine learning (ML) is a field of study in artificial intelligence concerned with the development and study of statistical algorithms that can learn from data and gen[0m[32;1m[1;3mMachine learning (ML) is a field of study in artificial intelligence concerned with the development and study of statistical algorithms that can learn from data and generate insights or predictions. It involves creating and using algorithms that allow computers to learn from and make predictions or decisions based on data. Machine learning is widely used in various applications such as image recognition, natural language processing, and recommendation systems.[0m

[1m> Finished chain.[0m


{'input': 'Tell me about Machine Learning',
 'output': 'Machine learning (ML) is a field of study in artificial intelligence concerned with the development and study of statistical algorithms that can learn from data and generate insights or predictions. It involves creating and using algorithms that allow computers to learn from and make predictions or decisions based on data. Machine learning is widely used in various applications such as image recognition, natural language processing, and recommendation systems.'}

In [58]:
agent_executor.invoke({"input":"What's the paper 2407.01140 about?"})



[1m> Entering new AgentExecutor chain...[0m
[32;1m[1;3m
Invoking: `arxiv` with `{'query': '2407.01140'}`


[0m[33;1m[1;3mPublished: 2024-07-01
Title: Amenable covers and relative bounded cohomology
Authors: Pietro Capovilla
Summary: We establish a relative version of Gromov's Vanishing Theorem in the presence
of amenabl[0m[32;1m[1;3mThe paper with identifier 2407.01140 is titled "Amenable covers and relative bounded cohomology" by Pietro Capovilla. It discusses a relative version of Gromov's Vanishing Theorem in the presence of amenability.[0m

[1m> Finished chain.[0m


{'input': "What's the paper 2407.01140 about?",
 'output': 'The paper with identifier 2407.01140 is titled "Amenable covers and relative bounded cohomology" by Pietro Capovilla. It discusses a relative version of Gromov\'s Vanishing Theorem in the presence of amenability.'}