# Caching, Batching and Trimming — AI Safety, Evaluation and Cost Control

Source: https://www.geekswithgeeks.com/en/ai-safety/cost-caching-trimming

> Cut cost with prompt caching, response caching, batch APIs and shorter context.

## Do less work, or do it cheaper

**Prompt caching** lets the provider reuse a repeated prefix (long instructions, tool definitions, documents) at a reduced price and lower latency, if the start of the prompt stays identical. **Response caching** stores answers to identical or near-identical questions so you skip the model call entirely (only for questions whose answers do not depend on the user or time). **Batch APIs** run non-urgent jobs asynchronously at a discount. **Trimming** removes unnecessary context: shorter instructions, summarised history, fewer retrieved chunks, and a capped output length.

## Prompt caching arithmetic, run

I ran this with an illustrative assumption that cached tokens are billed at 10% of the normal input price. For 1,000 input tokens at 3 per million, the cost is 0.003 without caching and 0.00084 when 80% of the prompt is a cache hit. Check your provider for the real cache pricing and rules.

```python
def cached_cost(in_tok, cached_frac, p_in, read_mult=0.1):
    return round(in_tok / 1e6 * p_in * ((1 - cached_frac) + cached_frac * read_mult), 6)

print(cached_cost(1000, 0.0, 3.0), cached_cost(1000, 0.8, 3.0))
```

Output:

```
0.003 0.00084
```

## Put stable text first

Caching only works on an unchanged prefix. Place long fixed instructions and documents at the start and the changing user question at the end; adding a timestamp at the top breaks the cache.

**Quiz:** Where should the changing user question go for prompt caching to work well?

- [ ] Nowhere
- [ ] At the very start
- [ ] Inside the API key
- [x] At the end, after the stable prefix

*Answer:* At the end, after the stable prefix. Only an identical prefix can be reused, so volatile content belongs after it.
