# Retrieval-Augmented Generation (RAG) — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/u-rag

> Ground answers in your documents using embeddings and search.

## Look it up, then answer

**RAG** reduces hallucination and bypasses the knowledge cutoff by adding relevant text to the prompt. Steps: (1) split documents into **chunks**; (2) store each chunk's **embedding** in a search index; (3) at question time, embed the question and fetch the nearest chunks; (4) put them in the prompt with the instruction "answer only from this context and cite it"; (5) generate. Quality depends mostly on **retrieval**: chunk size, metadata filters, hybrid keyword + vector search and re-ranking. Always evaluate retrieval separately from generation.

## Planning chunks, run

I ran this plain-Python (standard library only) example. A 3,000-word document is about 4,000 tokens by the rough 0.75 words-per-token rule. With a 4,096-token window and 500 tokens reserved for the answer, you need 2 chunks.

```python
def count_words_as_tokens(text):
    # crude rule of thumb: 1 token is about 0.75 English word
    return round(len(text.split()) / 0.75)

doc = "word " * 3000
print("approx tokens:", count_words_as_tokens(doc))
window = 4096
chunks = -(-count_words_as_tokens(doc) // (window - 500))
print("chunks needed with 500 tokens reserved for the answer:", chunks)

```

Output:

```
approx tokens: 4000
chunks needed with 500 tokens reserved for the answer: 2
```

## Test retrieval on its own

For 50 sample questions, check whether the right chunk appears in the top 5 results. If it does not, no prompt will save the answer.

**Quiz:** What is the main idea of RAG?

- [x] Retrieve relevant text and put it in the prompt before answering
- [ ] Retrain the model for each question
- [ ] Delete the context
- [ ] Use a larger tokenizer

*Answer:* Retrieve relevant text and put it in the prompt before answering. Grounding in retrieved text improves accuracy and allows citations.
