Lesson 2 / 25

Reading stop_reason

Handle the common stop reasons correctly: end of turn, tool use, token limit and stop sequence.

Why the model stopped

Each response says why generation ended. Common values in the Claude Messages API are end_turn (it finished its answer), tool_use (it wants a tool run), max_tokens (it hit your output limit mid-answer) and stop_sequence (it produced a sequence you asked it to stop at). Your loop should handle each one on purpose, never by accident.

Handling each reason

The max_tokens case matters: the answer or a tool call may be cut off, so it is not safe to treat the reply as complete.

match reply.stop_reason:
    case "end_turn":
        return final_text(reply)
    case "tool_use":
        messages.append(run_tools(reply))
    case "max_tokens":
        raise OutputTruncated("reply cut off; raise max_tokens or ask for less")
    case _:
        raise UnexpectedStop(reply.stop_reason)

Never assume end_turn

A loop that treats every non-tool reply as a finished answer will happily return half a sentence when max_tokens is hit. Check the reason explicitly.

Quick check: What does stop_reason "max_tokens" tell you?

  • The reply finished normally
  • The reply may be cut off at your output limit
  • The model refused the task
  • The API key expired
Answer

The reply may be cut off at your output limit — It means generation hit the output cap, so the content can be incomplete.