Lesson 21 / 27
Reducing Cost and Latency: Caching, Batching, Model Choice
Apply the biggest levers first.
Send less, reuse more, pick the right size
The biggest levers, roughly in order: (1) send fewer tokens: shorter system prompts, only relevant documents, trimmed history, concise tool definitions; (2) cap output with max_tokens and by asking for brevity; (3) use a smaller, cheaper model for easy tasks (classification, extraction) and a stronger one only where needed, with an evaluation set to confirm quality; (4) prompt caching: both providers can reuse a long unchanging prefix (system prompt, reference documents, tool definitions) at a lower price and with lower latency on repeat requests, if you place the stable content first; (5) batch APIs for non-urgent bulk work, typically at a discount; (6) response caching of identical requests in your own system; and (7) streaming for perceived speed. Always re-run your quality tests after each change.
What caching a long prefix could save, run
I ran this plain-Python (standard library only) example. With invented prices and a cached-prefix rate of 10%, 1,000 calls that reuse a 6,000-token prefix drop from $21.30 to $5.12, about 76% less. Real discounts, minimum prefix sizes and cache lifetimes differ by provider, so treat this as the shape of the saving, not a quote.
# Prompt caching idea: a long, unchanging prefix (system prompt + docs) can be billed at a lower rate on repeat calls.
prefix_tokens, question_tokens, output_tokens, calls = 6000, 100, 200, 1000
price_in, price_out, cached_factor = 3.00, 15.00, 0.10 # example numbers, NOT real prices
no_cache = calls * ((prefix_tokens + question_tokens) * price_in + output_tokens * price_out) / 1e6
first = (prefix_tokens + question_tokens) * price_in + output_tokens * price_out
rest = (calls - 1) * (prefix_tokens * price_in * cached_factor + question_tokens * price_in + output_tokens * price_out)
with_cache = (first + rest) / 1e6
print("without caching: $%.2f" % no_cache)
print("with caching : $%.2f" % with_cache)
print("saving : %.0f%%" % (100 * (1 - with_cache / no_cache)))
Output:
without caching: $21.30 with caching : $5.12 saving : 76%
Put stable content first
Caching works on a matching prefix. Put the unchanging instructions and documents at the start and the changing user question at the end.
Quick check: How should a prompt be arranged to benefit from prompt caching?
- Changing content first
- Stable content first, changing content last
- Random order
- Caching needs no arrangement
Answer
Stable content first, changing content last — Caches match on an identical prefix, so the shared part must come first.