# Retrieval-Augmented Generation (RAG) — Large Language Models

Source: https://www.geekswithgeeks.com/hi/llms/u-rag

> Embeddings और search से उत्तरों को अपने दस्तावेज़ों पर आधारित करें।

## पहले खोजें, फिर उत्तर दें

**RAG** प्रासंगिक पाठ को prompt में जोड़कर hallucination घटाता है और knowledge cutoff को दरकिनार करता है। चरण: (1) दस्तावेज़ों को **chunks** में बाँटें; (2) हर chunk का **embedding** search index में रखें; (3) सवाल के समय सवाल को embed करके निकटतम chunks लाएँ; (4) उन्हें "सिर्फ़ इस संदर्भ से उत्तर दें और स्रोत बताएँ" निर्देश के साथ prompt में रखें; (5) generate करें। गुणवत्ता अधिकतर **retrieval** पर निर्भर है: chunk आकार, metadata filters, hybrid keyword + vector search और re-ranking। Retrieval का मूल्यांकन generation से अलग करें।

## Chunks की योजना, चलाकर

मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। 3,000 शब्दों का दस्तावेज़ मोटे 0.75 शब्द-प्रति-token नियम से लगभग 4,000 tokens है। 4,096-token window और उत्तर के लिए आरक्षित 500 tokens के साथ 2 chunks चाहिए।

```python
def count_words_as_tokens(text):
    # crude rule of thumb: 1 token is about 0.75 English word
    return round(len(text.split()) / 0.75)

doc = "word " * 3000
print("approx tokens:", count_words_as_tokens(doc))
window = 4096
chunks = -(-count_words_as_tokens(doc) // (window - 500))
print("chunks needed with 500 tokens reserved for the answer:", chunks)

```

Output:

```
approx tokens: 4000
chunks needed with 500 tokens reserved for the answer: 2
```

## Retrieval का परीक्षण अलग से करें

50 नमूना सवालों के लिए जाँचें कि सही chunk शीर्ष 5 परिणामों में आता है या नहीं। न आए तो कोई prompt उत्तर नहीं बचाएगा।

**Quiz:** RAG का मुख्य विचार क्या है?

- [x] प्रासंगिक पाठ खोजकर उत्तर से पहले prompt में रखना
- [ ] हर सवाल पर मॉडल को दोबारा प्रशिक्षित करना
- [ ] Context हटाना
- [ ] बड़ा tokenizer उपयोग करना

*Answer:* प्रासंगिक पाठ खोजकर उत्तर से पहले prompt में रखना. Retrieved पाठ पर आधार उत्तरों की सटीकता बढ़ाता है और स्रोत देना संभव करता है।
