Lesson 7 / 27
Sampling: Greedy, Top-k and Top-p
Choose the next token with truncation strategies.
Cut off the unlikely tail
Picking from the full distribution sometimes selects a very unlikely token and derails the text. Greedy always takes the top token (deterministic but repetitive). Top-k keeps only the k most likely tokens. Top-p (nucleus) keeps the smallest set whose probabilities add up to p. In both cases the remaining probabilities are renormalised and a token is drawn at random. These settings and temperature together shape creativity versus reliability.
Top-k and top-p, run
I ran this plain-Python (standard library only) example. Top-k with k=3 keeps cat, dog, fish. Top-p with p=0.8 keeps just cat and dog because 0.503 + 0.305 passes 0.8. Draws use a fixed random seed.
import math, random
def softmax(xs):
m = max(xs); e = [math.exp(x - m) for x in xs]; s = sum(e)
return [v / s for v in e]
words = ["cat", "dog", "fish", "bird", "rock", "tree"]
probs = softmax([3.0, 2.5, 1.5, 1.0, -1.0, -2.0])
def top_k(words, probs, k):
pairs = sorted(zip(words, probs), key=lambda p: -p[1])[:k]
z = sum(p for _, p in pairs)
return [(w, p / z) for w, p in pairs]
def top_p(words, probs, p):
pairs = sorted(zip(words, probs), key=lambda x: -x[1])
keep, total = [], 0.0
for w, pr in pairs:
keep.append((w, pr)); total += pr
if total >= p:
break
z = sum(pr for _, pr in keep)
return [(w, pr / z) for w, pr in keep]
print("full ", [(w, round(p, 3)) for w, p in zip(words, probs)])
print("top-k3", [(w, round(p, 3)) for w, p in top_k(words, probs, 3)])
print("top-p0.8", [(w, round(p, 3)) for w, p in top_p(words, probs, 0.8)])
random.seed(1)
kept = top_p(words, probs, 0.8)
draws = random.choices([w for w, _ in kept], [p for _, p in kept], k=10)
print(draws)
Output:
full [('cat', 0.503), ('dog', 0.305), ('fish', 0.112), ('bird', 0.068), ('rock', 0.009), ('tree', 0.003)]
top-k3 [('cat', 0.547), ('dog', 0.331), ('fish', 0.122)]
top-p0.8 [('cat', 0.622), ('dog', 0.378)]
['cat', 'dog', 'dog', 'cat', 'cat', 'cat', 'dog', 'dog', 'cat', 'cat']Change one sampling knob at a time
Providers usually recommend adjusting either temperature or top-p, not both, so you can tell which change caused a difference.
Quick check: What does top-p sampling keep?
- The smallest set of tokens whose probabilities sum to at least p
- Exactly p tokens
- Only the last token
- Every token
Answer
The smallest set of tokens whose probabilities sum to at least p — The size of the kept set adapts to how confident the model is.