Lesson 24 / 27
Latency, Cost and Caching
Make RAG fast and affordable.
Where the time and money go
Query-time latency is the sum of query rewrite, embedding the question, retrieval, reranking, and generation; generation usually dominates and grows with prompt size. Ways to cut it: stream the answer, retrieve fewer and better chunks, rerank only a small shortlist, use a smaller model for easy questions (and a larger one only when needed), cache embeddings of frequent queries and whole answers for identical questions, use prompt caching for the fixed instruction prefix where supported, and run retrieval steps in parallel. Track p50/p95 latency, tokens per request and cost per answered question.
Measure cost per answered question
Divide total monthly cost by questions actually answered, not by requests, so refusals and retries are counted honestly.
Quick check: Which step usually dominates RAG latency?
- Computing a hash
- Loading the page
- LLM generation
- Sorting five rows
Answer
LLM generation — Generating tokens takes longer than the retrieval steps, especially with long prompts.