# Latency, Cost and Caching — Retrieval-Augmented Generation (RAG)

Source: https://www.geekswithgeeks.com/en/rag/p-perf

> Make RAG fast and affordable.

## Where the time and money go

Query-time latency is the sum of query rewrite, embedding the question, retrieval, reranking, and generation; generation usually dominates and grows with prompt size. Ways to cut it: **stream** the answer, retrieve fewer and better chunks, rerank only a small shortlist, use a **smaller model** for easy questions (and a larger one only when needed), **cache** embeddings of frequent queries and whole answers for identical questions, use **prompt caching** for the fixed instruction prefix where supported, and run retrieval steps **in parallel**. Track p50/p95 latency, tokens per request and cost per answered question.

## Measure cost per answered question

Divide total monthly cost by questions actually answered, not by requests, so refusals and retries are counted honestly.

**Quiz:** Which step usually dominates RAG latency?

- [ ] Computing a hash
- [ ] Loading the page
- [x] LLM generation
- [ ] Sorting five rows

*Answer:* LLM generation. Generating tokens takes longer than the retrieval steps, especially with long prompts.
