Lesson 7 / 27

Sampling: Greedy, Top-k and Top-p

Choose the next token with truncation strategies.

Cut off the unlikely tail

Picking from the full distribution sometimes selects a very unlikely token and derails the text. Greedy always takes the top token (deterministic but repetitive). Top-k keeps only the k most likely tokens. Top-p (nucleus) keeps the smallest set whose probabilities add up to p. In both cases the remaining probabilities are renormalised and a token is drawn at random. These settings and temperature together shape creativity versus reliability.

Top-k and top-p, run

I ran this plain-Python (standard library only) example. Top-k with k=3 keeps cat, dog, fish. Top-p with p=0.8 keeps just cat and dog because 0.503 + 0.305 passes 0.8. Draws use a fixed random seed.

import math, random

def softmax(xs):
    m = max(xs); e = [math.exp(x - m) for x in xs]; s = sum(e)
    return [v / s for v in e]

words = ["cat", "dog", "fish", "bird", "rock", "tree"]
probs = softmax([3.0, 2.5, 1.5, 1.0, -1.0, -2.0])

def top_k(words, probs, k):
    pairs = sorted(zip(words, probs), key=lambda p: -p[1])[:k]
    z = sum(p for _, p in pairs)
    return [(w, p / z) for w, p in pairs]

def top_p(words, probs, p):
    pairs = sorted(zip(words, probs), key=lambda x: -x[1])
    keep, total = [], 0.0
    for w, pr in pairs:
        keep.append((w, pr)); total += pr
        if total >= p:
            break
    z = sum(pr for _, pr in keep)
    return [(w, pr / z) for w, pr in keep]

print("full  ", [(w, round(p, 3)) for w, p in zip(words, probs)])
print("top-k3", [(w, round(p, 3)) for w, p in top_k(words, probs, 3)])
print("top-p0.8", [(w, round(p, 3)) for w, p in top_p(words, probs, 0.8)])

random.seed(1)
kept = top_p(words, probs, 0.8)
draws = random.choices([w for w, _ in kept], [p for _, p in kept], k=10)
print(draws)

Output:

full   [('cat', 0.503), ('dog', 0.305), ('fish', 0.112), ('bird', 0.068), ('rock', 0.009), ('tree', 0.003)]
top-k3 [('cat', 0.547), ('dog', 0.331), ('fish', 0.122)]
top-p0.8 [('cat', 0.622), ('dog', 0.378)]
['cat', 'dog', 'dog', 'cat', 'cat', 'cat', 'dog', 'dog', 'cat', 'cat']

Change one sampling knob at a time

Providers usually recommend adjusting either temperature or top-p, not both, so you can tell which change caused a difference.

Quick check: What does top-p sampling keep?

  • The smallest set of tokens whose probabilities sum to at least p
  • Exactly p tokens
  • Only the last token
  • Every token
Answer

The smallest set of tokens whose probabilities sum to at least p — The size of the kept set adapts to how confident the model is.