# Reducing Cost and Latency: Caching, Batching, Model Choice — Claude API / OpenAI API Basics

Source: https://www.geekswithgeeks.com/en/llm-apis/c-reduce

> Apply the biggest levers first.

## Send less, reuse more, pick the right size

The biggest levers, roughly in order: (1) **send fewer tokens**: shorter system prompts, only relevant documents, trimmed history, concise tool definitions; (2) **cap output** with `max_tokens` and by asking for brevity; (3) **use a smaller, cheaper model** for easy tasks (classification, extraction) and a stronger one only where needed, with an evaluation set to confirm quality; (4) **prompt caching**: both providers can reuse a long unchanging prefix (system prompt, reference documents, tool definitions) at a lower price and with lower latency on repeat requests, if you place the stable content first; (5) **batch APIs** for non-urgent bulk work, typically at a discount; (6) **response caching** of identical requests in your own system; and (7) **streaming** for perceived speed. Always re-run your quality tests after each change.

## What caching a long prefix could save, run

I ran this plain-Python (standard library only) example. With invented prices and a cached-prefix rate of 10%, 1,000 calls that reuse a 6,000-token prefix drop from $21.30 to $5.12, about 76% less. Real discounts, minimum prefix sizes and cache lifetimes differ by provider, so treat this as the shape of the saving, not a quote.

```python
# Prompt caching idea: a long, unchanging prefix (system prompt + docs) can be billed at a lower rate on repeat calls.
prefix_tokens, question_tokens, output_tokens, calls = 6000, 100, 200, 1000
price_in, price_out, cached_factor = 3.00, 15.00, 0.10      # example numbers, NOT real prices

no_cache = calls * ((prefix_tokens + question_tokens) * price_in + output_tokens * price_out) / 1e6
first = (prefix_tokens + question_tokens) * price_in + output_tokens * price_out
rest = (calls - 1) * (prefix_tokens * price_in * cached_factor + question_tokens * price_in + output_tokens * price_out)
with_cache = (first + rest) / 1e6
print("without caching: $%.2f" % no_cache)
print("with caching   : $%.2f" % with_cache)
print("saving         : %.0f%%" % (100 * (1 - with_cache / no_cache)))

```

Output:

```
without caching: $21.30
with caching   : $5.12
saving         : 76%
```

## Put stable content first

Caching works on a matching prefix. Put the unchanging instructions and documents at the start and the changing user question at the end.

**Quiz:** How should a prompt be arranged to benefit from prompt caching?

- [ ] Changing content first
- [x] Stable content first, changing content last
- [ ] Random order
- [ ] Caching needs no arrangement

*Answer:* Stable content first, changing content last. Caches match on an identical prefix, so the shared part must come first.
