# Cost and Latency — Advanced Agent Workflows and Skills

Source: https://www.geekswithgeeks.com/en/agent-workflows/rel-cost-latency

> Reduce spend and wait time with caching, smaller models and trimmed context.

## Where the tokens go

Cost grows with tokens in and out and with the number of model calls. Cut it by sending less context, caching repeated prefixes with **prompt caching**, using a smaller model for easy steps, and running independent calls in parallel to reduce waiting. Measure first so you optimise the step that really dominates.

## Count before you cut

Log tokens and time per step. Often one verbose tool result or one oversized file accounts for most of the bill, and trimming that single thing saves more than any clever prompt.

**Quiz:** What should you do before optimising agent cost?

- [ ] Guess which step is expensive
- [x] Measure tokens and time per step
- [ ] Remove all tools
- [ ] Switch providers immediately

*Answer:* Measure tokens and time per step. Measurements show where the cost really is, so effort goes to the right place.
