Lesson 17 / 31
Building a RAG Chain
Wire retriever, prompt, model and parser together.
Context plus question into the prompt
A RAG chain feeds two things into the prompt: the retrieved context and the original question. In LCEL this is a dict at the start of the chain: {"context": retriever | format_docs, "question": RunnablePassthrough()} then | prompt | model | parser. The formatting step numbers the passages so the model can cite [1]. Instruct the model to answer only from the context and to say it does not know otherwise. To return sources alongside the answer, keep the retrieved Documents in the output (using a parallel branch) so your UI can link to them.
A complete RAG chain, run
I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. The printed prompt shows the retrieved passage numbered [1] in the context, followed by the question. The fake model returns "24 days [1]". Retrieval here is by a meaningless fake embedding and a one-document match on identical text, so only the wiring is being demonstrated; swap in a real embedding model and chat model for real answers.
from langchain_core.documents import Document
from langchain_core.vectorstores import InMemoryVectorStore
from langchain_core.embeddings import DeterministicFakeEmbedding
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser
from langchain_core.runnables import RunnablePassthrough, RunnableLambda
from langchain_core.language_models.fake_chat_models import FakeListChatModel
store = InMemoryVectorStore(DeterministicFakeEmbedding(size=32))
store.add_documents([Document(page_content="Employees get 24 days of paid leave per year."),
Document(page_content="Hotels are capped at 6000 rupees per night.")])
retriever = store.as_retriever(search_kwargs={"k": 1})
def fmt(docs): return "\n".join(f"[{i}] {d.page_content}" for i, d in enumerate(docs, 1))
prompt = ChatPromptTemplate.from_template("Answer from the context only.\nContext:\n{context}\nQuestion: {question}")
model = FakeListChatModel(responses=["24 days [1]"])
chain = ({"context": retriever | RunnableLambda(fmt), "question": RunnablePassthrough()}
| prompt | model | StrOutputParser())
question = "Employees get 24 days of paid leave per year."
print(prompt.invoke({"context": fmt(retriever.invoke(question)), "question": question}).to_string())
print("answer:", chain.invoke(question))
Output:
Human: Answer from the context only. Context: [1] Employees get 24 days of paid leave per year. Question: Employees get 24 days of paid leave per year. answer: 24 days [1]
Quick check: What two things does the RAG prompt receive?
- Two models
- A GPU and a CPU
- Retrieved context and the original question
- Only the answer
Answer
Retrieved context and the original question — The model needs both evidence and the question to answer faithfully.