Lesson 11 / 25

Setting max_tokens Wisely

Choose an output cap that fits the task and handle truncation.

A cap, not a target

max_tokens is the most the model may generate in one reply. It does not make the model write that much. Too low and answers or tool calls get cut off; very high wastes your reserve and allows runaway output. Pick a value that fits the expected answer (a short tool call needs far less than a long report).

Different caps for different turns

Use a small cap while the agent is acting through tools and a larger one for the final write-up.

ACTING_CAP = 1_024      # tool calls and short thoughts
FINAL_CAP  = 4_096      # final summary

reply = client.messages.create(
    model=MODEL, max_tokens=ACTING_CAP, tools=TOOLS, messages=messages)

Quick check: What does setting max_tokens higher do?

  • Forces the model to write that many tokens
  • Allows longer replies but does not force them
  • Makes the model faster
  • Increases the context window
Answer

Allows longer replies but does not force them — It raises the ceiling on reply length; the model still stops when it is done.