# Caching: Exact, Normalised and Semantic — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/p-cache

> Avoid paying twice for the same question.

## The cheapest call is the one you do not make

Real traffic repeats itself. **Exact caching** keys on the full request and only hits for identical text. **Normalised caching** first lowercases, trims and removes punctuation, so trivial variations hit the same entry. **Semantic caching** compares embeddings of the question and reuses an answer for sufficiently similar ones; it catches paraphrases but risks returning an answer to a **subtly different question**, so use a strict similarity threshold, restrict it to safe categories (FAQs, not personalised or time-sensitive answers) and include the user's tenant and permissions in the cache key. Also use the provider's **prompt caching** for long stable prefixes. Set sensible **expiry**, and invalidate when documents or prompts change. Measure the **hit rate** and check that cached answers are still correct.

## Fast enough, cheap enough, big enough

Cache repeated work, route easy requests to cheaper models, tame tail latency and size capacity with simple arithmetic.

![Five levers: cache, route, hedge, size, price.](assets/figures/llm-engineering/section-4-map.svg) — Figure 4.1 — Cache, route, hedge, size and price.

## Exact versus normalised cache hits, run

I ran this with plain Python 3 (standard library only). Of 7 requests, only 1 is a repeat of the exact same text, but after normalising case and punctuation 4 hit the cache, avoiding 4 model calls and saving $0.016 of $0.028 at an assumed $0.004 per call.

```python
import re

def normalise(q): return re.sub(r"[^a-z0-9 ]", "", q.lower()).strip()
stream = ["What is your refund policy?", "what is your refund policy", "How do I reset my password?",
          "What is your refund policy?!", "how do i reset my password", "Where is my order?", "What is your refund policy?"]

exact, norm = {}, {}
hits_exact = hits_norm = 0
for q in stream:
    hits_exact += q in exact; exact[q] = True
    k = normalise(q); hits_norm += k in norm; norm[k] = True
n = len(stream)
print(f"requests: {n} | exact-text cache hits: {hits_exact} | normalised-key cache hits: {hits_norm}")
price = 0.004                                               # example cost of one model call, USD
print(f"model calls avoided with the normalised cache: {hits_norm} -> saves ${hits_norm * price:.3f} of ${n * price:.3f}")

```

Output:

```
requests: 7 | exact-text cache hits: 1 | normalised-key cache hits: 4
model calls avoided with the normalised cache: 4 -> saves $0.016 of $0.028
```

## Include permissions in the cache key

Otherwise one user's cached answer could be served to another.

**Quiz:** What is the main risk of semantic caching?

- [ ] It deletes the prompt
- [ ] It makes embeddings unusable
- [ ] It always costs more
- [x] Returning an answer to a similar but different question

*Answer:* Returning an answer to a similar but different question. Use strict thresholds and limit it to safe, non-personalised categories.
