Lesson 14 / 29
Caching: Exact, Normalised and Semantic
Avoid paying twice for the same question.
The cheapest call is the one you do not make
Real traffic repeats itself. Exact caching keys on the full request and only hits for identical text. Normalised caching first lowercases, trims and removes punctuation, so trivial variations hit the same entry. Semantic caching compares embeddings of the question and reuses an answer for sufficiently similar ones; it catches paraphrases but risks returning an answer to a subtly different question, so use a strict similarity threshold, restrict it to safe categories (FAQs, not personalised or time-sensitive answers) and include the user's tenant and permissions in the cache key. Also use the provider's prompt caching for long stable prefixes. Set sensible expiry, and invalidate when documents or prompts change. Measure the hit rate and check that cached answers are still correct.
Fast enough, cheap enough, big enough
Cache repeated work, route easy requests to cheaper models, tame tail latency and size capacity with simple arithmetic.
Exact versus normalised cache hits, run
I ran this with plain Python 3 (standard library only). Of 7 requests, only 1 is a repeat of the exact same text, but after normalising case and punctuation 4 hit the cache, avoiding 4 model calls and saving $0.016 of $0.028 at an assumed $0.004 per call.
import re
def normalise(q): return re.sub(r"[^a-z0-9 ]", "", q.lower()).strip()
stream = ["What is your refund policy?", "what is your refund policy", "How do I reset my password?",
"What is your refund policy?!", "how do i reset my password", "Where is my order?", "What is your refund policy?"]
exact, norm = {}, {}
hits_exact = hits_norm = 0
for q in stream:
hits_exact += q in exact; exact[q] = True
k = normalise(q); hits_norm += k in norm; norm[k] = True
n = len(stream)
print(f"requests: {n} | exact-text cache hits: {hits_exact} | normalised-key cache hits: {hits_norm}")
price = 0.004 # example cost of one model call, USD
print(f"model calls avoided with the normalised cache: {hits_norm} -> saves ${hits_norm * price:.3f} of ${n * price:.3f}")
Output:
requests: 7 | exact-text cache hits: 1 | normalised-key cache hits: 4 model calls avoided with the normalised cache: 4 -> saves $0.016 of $0.028
Include permissions in the cache key
Otherwise one user's cached answer could be served to another.
Quick check: What is the main risk of semantic caching?
- It deletes the prompt
- It makes embeddings unusable
- It always costs more
- Returning an answer to a similar but different question
Answer
Returning an answer to a similar but different question — Use strict thresholds and limit it to safe, non-personalised categories.