Lesson 18 / 27
Retrieval-Augmented Generation (RAG)
Ground answers in your documents using embeddings and search.
Look it up, then answer
RAG reduces hallucination and bypasses the knowledge cutoff by adding relevant text to the prompt. Steps: (1) split documents into chunks; (2) store each chunk's embedding in a search index; (3) at question time, embed the question and fetch the nearest chunks; (4) put them in the prompt with the instruction "answer only from this context and cite it"; (5) generate. Quality depends mostly on retrieval: chunk size, metadata filters, hybrid keyword + vector search and re-ranking. Always evaluate retrieval separately from generation.
Planning chunks, run
I ran this plain-Python (standard library only) example. A 3,000-word document is about 4,000 tokens by the rough 0.75 words-per-token rule. With a 4,096-token window and 500 tokens reserved for the answer, you need 2 chunks.
def count_words_as_tokens(text):
# crude rule of thumb: 1 token is about 0.75 English word
return round(len(text.split()) / 0.75)
doc = "word " * 3000
print("approx tokens:", count_words_as_tokens(doc))
window = 4096
chunks = -(-count_words_as_tokens(doc) // (window - 500))
print("chunks needed with 500 tokens reserved for the answer:", chunks)
Output:
approx tokens: 4000 chunks needed with 500 tokens reserved for the answer: 2
Test retrieval on its own
For 50 sample questions, check whether the right chunk appears in the top 5 results. If it does not, no prompt will save the answer.
Quick check: What is the main idea of RAG?
- Retrieve relevant text and put it in the prompt before answering
- Retrain the model for each question
- Delete the context
- Use a larger tokenizer
Answer
Retrieve relevant text and put it in the prompt before answering — Grounding in retrieved text improves accuracy and allows citations.