Lesson 22 / 25
Caching, Batching and Trimming
Cut cost with prompt caching, response caching, batch APIs and shorter context.
Do less work, or do it cheaper
Prompt caching lets the provider reuse a repeated prefix (long instructions, tool definitions, documents) at a reduced price and lower latency, if the start of the prompt stays identical. Response caching stores answers to identical or near-identical questions so you skip the model call entirely (only for questions whose answers do not depend on the user or time). Batch APIs run non-urgent jobs asynchronously at a discount. Trimming removes unnecessary context: shorter instructions, summarised history, fewer retrieved chunks, and a capped output length.
Prompt caching arithmetic, run
I ran this with an illustrative assumption that cached tokens are billed at 10% of the normal input price. For 1,000 input tokens at 3 per million, the cost is 0.003 without caching and 0.00084 when 80% of the prompt is a cache hit. Check your provider for the real cache pricing and rules.
def cached_cost(in_tok, cached_frac, p_in, read_mult=0.1):
return round(in_tok / 1e6 * p_in * ((1 - cached_frac) + cached_frac * read_mult), 6)
print(cached_cost(1000, 0.0, 3.0), cached_cost(1000, 0.8, 3.0))
Output:
0.003 0.00084
Put stable text first
Caching only works on an unchanged prefix. Place long fixed instructions and documents at the start and the changing user question at the end; adding a timestamp at the top breaks the cache.
Quick check: Where should the changing user question go for prompt caching to work well?
- Nowhere
- At the very start
- Inside the API key
- At the end, after the stable prefix
Answer
At the end, after the stable prefix — Only an identical prefix can be reused, so volatile content belongs after it.