# Caching: सटीक, Normalised और Semantic — LLM Engineering की बुनियाद

Source: https://www.geekswithgeeks.com/hi/llm-engineering/p-cache

> एक ही सवाल के लिए दो बार भुगतान से बचें।

## सबसे सस्ती call वह है जो आप नहीं करते

असली traffic ख़ुद को दोहराता है। **सटीक caching** पूरे अनुरोध पर key करती है और केवल समान पाठ पर hit देती है। **Normalised caching** पहले छोटे अक्षर, trim और विराम चिह्न हटाती है, इसलिए मामूली भिन्नताएँ एक ही प्रविष्टि पर hit देती हैं। **Semantic caching** सवाल के embeddings की तुलना करती है और पर्याप्त समान सवालों के लिए उत्तर पुनः उपयोग करती है; वह पुनर्कथन पकड़ती है पर **सूक्ष्म रूप से अलग सवाल** का उत्तर लौटाने का जोखिम रखती है, इसलिए कड़ी समानता सीमा उपयोग करें, उसे सुरक्षित श्रेणियों (FAQs, वैयक्तिक या समय-संवेदी उत्तर नहीं) तक सीमित करें और cache key में user का tenant व अनुमतियाँ शामिल करें। लंबे स्थिर prefixes के लिए provider की **prompt caching** भी उपयोग करें। उचित **expiry** रखें, और दस्तावेज़ या prompts बदलने पर अमान्य करें। **Hit rate** मापें और जाँचें कि cached उत्तर अब भी सही हैं।

## पर्याप्त तेज़, पर्याप्त सस्ता, पर्याप्त बड़ा

दोहराया काम cache करें, आसान अनुरोध सस्ते मॉडलों को भेजें, tail latency सँभालें और सरल अंकगणित से क्षमता आँकें।

![पाँच उत्तोलक: cache, route, hedge, आकार, क़ीमत।](assets/figures/llm-engineering/section-4-map.svg) — चित्र 4.1 — Cache, route, hedge, आकार और क़ीमत।

## सटीक बनाम normalised cache hits, चलाकर

मैंने यह सादे Python 3 (सिर्फ़ standard library) से चलाया। 7 अनुरोधों में सिर्फ़ 1 ठीक उसी पाठ का दोहराव है, पर अक्षर-रूप और विराम चिह्न normalise करने के बाद 4 cache को hit करते हैं, 4 मॉडल calls बचाते हुए, मान ली $0.004 प्रति call पर $0.028 में से $0.016 की बचत।

```python
import re

def normalise(q): return re.sub(r"[^a-z0-9 ]", "", q.lower()).strip()
stream = ["What is your refund policy?", "what is your refund policy", "How do I reset my password?",
          "What is your refund policy?!", "how do i reset my password", "Where is my order?", "What is your refund policy?"]

exact, norm = {}, {}
hits_exact = hits_norm = 0
for q in stream:
    hits_exact += q in exact; exact[q] = True
    k = normalise(q); hits_norm += k in norm; norm[k] = True
n = len(stream)
print(f"requests: {n} | exact-text cache hits: {hits_exact} | normalised-key cache hits: {hits_norm}")
price = 0.004                                               # example cost of one model call, USD
print(f"model calls avoided with the normalised cache: {hits_norm} -> saves ${hits_norm * price:.3f} of ${n * price:.3f}")

```

Output:

```
requests: 7 | exact-text cache hits: 1 | normalised-key cache hits: 4
model calls avoided with the normalised cache: 4 -> saves $0.016 of $0.028
```

## Cache key में अनुमतियाँ शामिल करें

वरना एक user का cached उत्तर दूसरे को परोसा जा सकता है।

**Quiz:** Semantic caching का मुख्य जोखिम क्या है?

- [ ] यह prompt हटाती है
- [ ] यह embeddings अनुपयोगी बनाती है
- [ ] यह हमेशा ज़्यादा ख़र्चती है
- [x] समान पर अलग सवाल का उत्तर लौटाना

*Answer:* समान पर अलग सवाल का उत्तर लौटाना. कड़ी सीमाएँ उपयोग करें और इसे सुरक्षित, अवैयक्तिक श्रेणियों तक सीमित रखें।
