# Setting max_tokens Wisely — Agent Loops, Stop Conditions and Token Budgets

Source: https://www.geekswithgeeks.com/en/agent-loops/budget-max-tokens

> Choose an output cap that fits the task and handle truncation.

## A cap, not a target

`max_tokens` is the most the model may generate in one reply. It does not make the model write that much. Too low and answers or tool calls get cut off; very high wastes your reserve and allows runaway output. Pick a value that fits the expected answer (a short tool call needs far less than a long report).

## Different caps for different turns

Use a small cap while the agent is acting through tools and a larger one for the final write-up.

```python
ACTING_CAP = 1_024      # tool calls and short thoughts
FINAL_CAP  = 4_096      # final summary

reply = client.messages.create(
    model=MODEL, max_tokens=ACTING_CAP, tools=TOOLS, messages=messages)
```

**Quiz:** What does setting max_tokens higher do?

- [ ] Forces the model to write that many tokens
- [x] Allows longer replies but does not force them
- [ ] Makes the model faster
- [ ] Increases the context window

*Answer:* Allows longer replies but does not force them. It raises the ceiling on reply length; the model still stops when it is done.
