Lesson 24 / 27

Latency, Cost and Caching

Make RAG fast and affordable.

Where the time and money go

Query-time latency is the sum of query rewrite, embedding the question, retrieval, reranking, and generation; generation usually dominates and grows with prompt size. Ways to cut it: stream the answer, retrieve fewer and better chunks, rerank only a small shortlist, use a smaller model for easy questions (and a larger one only when needed), cache embeddings of frequent queries and whole answers for identical questions, use prompt caching for the fixed instruction prefix where supported, and run retrieval steps in parallel. Track p50/p95 latency, tokens per request and cost per answered question.

Measure cost per answered question

Divide total monthly cost by questions actually answered, not by requests, so refusals and retries are counted honestly.

Quick check: Which step usually dominates RAG latency?

  • Computing a hash
  • Loading the page
  • LLM generation
  • Sorting five rows
Answer

LLM generation — Generating tokens takes longer than the retrieval steps, especially with long prompts.