Lesson 13 / 29
Agent Limits, Costs and Stop Conditions
Bound steps, tokens, time and money.
Every loop needs a budget
An agent makes many model calls, each re-reading a growing message history, so cost and latency grow faster than linearly with the number of steps. Set several limits: the recursion limit, a maximum number of tool calls in state, a token budget (stop or summarise when history gets long), a wall-clock timeout, and a cost cap per run. Decide the stop behaviour in advance: return the best partial answer, ask the user, or escalate. Track tokens and cost per run in your traces, and review the longest runs regularly: they reveal prompt problems and missing tools.
Quick check: Why does agent cost grow faster than the number of steps?
- Tokens get cheaper with time
- Each step re-reads a longer message history
- Tools are free
- Graphs forbid caching
Answer
Each step re-reads a longer message history — Growing context makes later calls more expensive than earlier ones.