Lesson 21 / 27

Reducing Cost and Latency: Caching, Batching, Model Choice

Apply the biggest levers first.

Send less, reuse more, pick the right size

The biggest levers, roughly in order: (1) send fewer tokens: shorter system prompts, only relevant documents, trimmed history, concise tool definitions; (2) cap output with max_tokens and by asking for brevity; (3) use a smaller, cheaper model for easy tasks (classification, extraction) and a stronger one only where needed, with an evaluation set to confirm quality; (4) prompt caching: both providers can reuse a long unchanging prefix (system prompt, reference documents, tool definitions) at a lower price and with lower latency on repeat requests, if you place the stable content first; (5) batch APIs for non-urgent bulk work, typically at a discount; (6) response caching of identical requests in your own system; and (7) streaming for perceived speed. Always re-run your quality tests after each change.

What caching a long prefix could save, run

I ran this plain-Python (standard library only) example. With invented prices and a cached-prefix rate of 10%, 1,000 calls that reuse a 6,000-token prefix drop from $21.30 to $5.12, about 76% less. Real discounts, minimum prefix sizes and cache lifetimes differ by provider, so treat this as the shape of the saving, not a quote.

# Prompt caching idea: a long, unchanging prefix (system prompt + docs) can be billed at a lower rate on repeat calls.
prefix_tokens, question_tokens, output_tokens, calls = 6000, 100, 200, 1000
price_in, price_out, cached_factor = 3.00, 15.00, 0.10      # example numbers, NOT real prices

no_cache = calls * ((prefix_tokens + question_tokens) * price_in + output_tokens * price_out) / 1e6
first = (prefix_tokens + question_tokens) * price_in + output_tokens * price_out
rest = (calls - 1) * (prefix_tokens * price_in * cached_factor + question_tokens * price_in + output_tokens * price_out)
with_cache = (first + rest) / 1e6
print("without caching: $%.2f" % no_cache)
print("with caching   : $%.2f" % with_cache)
print("saving         : %.0f%%" % (100 * (1 - with_cache / no_cache)))

Output:

without caching: $21.30
with caching   : $5.12
saving         : 76%

Put stable content first

Caching works on a matching prefix. Put the unchanging instructions and documents at the start and the changing user question at the end.

Quick check: How should a prompt be arranged to benefit from prompt caching?

  • Changing content first
  • Stable content first, changing content last
  • Random order
  • Caching needs no arrangement
Answer

Stable content first, changing content last — Caches match on an identical prefix, so the shared part must come first.