Lesson 2 / 29
How Models Read Your Prompt
Understand tokens, context and why wording changes results.
Tokens in, probabilities out
The model splits your prompt into tokens, reads all of them at once within its context window, and then generates its reply one token at a time, each chosen from a probability distribution shaped by everything before it. Consequences: (1) it only knows what is in the prompt plus its training, so you must supply missing facts; (2) tiny wording changes can shift probabilities and change answers; (3) text near the beginning and end is often used more reliably than details buried in a very long middle; (4) replies are not guaranteed identical across runs unless sampling is fixed.
Say what you left unsaid
If a human would need to ask a clarifying question, the model needs the answer in the prompt. Include audience, constraints and definitions.
Quick check: Why must you include facts the model may not know?
- It only has the prompt plus its training, not your private data
- Models forget their training
- Facts shorten the reply
- It is a rule of HTTP
Answer
It only has the prompt plus its training, not your private data — Missing facts lead to guesses, so supply them in the prompt.