Lesson 7 / 27

Parameters: max_tokens, Temperature, Top-p, Stop

Control length, randomness and stopping.

A few knobs matter most

max_tokens caps the length of the reply (and so the maximum output cost); if the model reaches it, the reply is cut off and the stop reason says so, so check for it. temperature controls randomness: low values (near 0) give focused, repeatable-ish output for extraction and classification; higher values give more varied text for brainstorming; it never guarantees identical or correct output. top_p (nucleus sampling) limits choices to the most likely tokens; change either temperature or top_p, not both. stop sequences end generation when a given string appears. Some newer models restrict or ignore some sampling parameters, so consult the documentation for the model you use.

Parameter cheat sheet

Starting values to adapt to your task.

Task                          temperature   max_tokens        notes
extract fields / classify     0 - 0.2       small (50-200)    validate the output in code
summarise a document          0.2 - 0.5     medium            ask for a length in the prompt too
customer reply draft          0.3 - 0.7     medium            human review for sensitive cases
brainstorm names / ideas      0.8 - 1.0     medium            expect variety, check for repetition

Always: check stop_reason / finish_reason for truncation ("max_tokens" / "length").

Set max_tokens deliberately

Too low truncates answers; too high allows long, costly replies. Pick a limit from the longest answer you actually expect.

Quick check: What does a stop reason of max_tokens (or length) tell you?

  • The reply was cut off at your limit
  • The reply is perfect
  • The key is invalid
  • The model refused
Answer

The reply was cut off at your limit — Truncated output may be incomplete or invalid JSON; handle it explicitly.