# Parameters: max_tokens, Temperature, Top-p, Stop — Claude API / OpenAI API Basics

Source: https://www.geekswithgeeks.com/en/llm-apis/m-params

> Control length, randomness and stopping.

## A few knobs matter most

**max_tokens** caps the length of the reply (and so the maximum output cost); if the model reaches it, the reply is cut off and the stop reason says so, so check for it. **temperature** controls randomness: low values (near 0) give focused, repeatable-ish output for extraction and classification; higher values give more varied text for brainstorming; it never guarantees identical or correct output. **top_p** (nucleus sampling) limits choices to the most likely tokens; change either temperature or top_p, not both. **stop sequences** end generation when a given string appears. Some newer models restrict or ignore some sampling parameters, so consult the documentation for the model you use.

## Parameter cheat sheet

Starting values to adapt to your task.

```text
Task                          temperature   max_tokens        notes
extract fields / classify     0 - 0.2       small (50-200)    validate the output in code
summarise a document          0.2 - 0.5     medium            ask for a length in the prompt too
customer reply draft          0.3 - 0.7     medium            human review for sensitive cases
brainstorm names / ideas      0.8 - 1.0     medium            expect variety, check for repetition

Always: check stop_reason / finish_reason for truncation ("max_tokens" / "length").
```

## Set max_tokens deliberately

Too low truncates answers; too high allows long, costly replies. Pick a limit from the longest answer you actually expect.

**Quiz:** What does a stop reason of max_tokens (or length) tell you?

- [x] The reply was cut off at your limit
- [ ] The reply is perfect
- [ ] The key is invalid
- [ ] The model refused

*Answer:* The reply was cut off at your limit. Truncated output may be incomplete or invalid JSON; handle it explicitly.
