Lesson 9 / 27
Loss and Perplexity
Measure how well a language model predicts text.
How surprised is the model
Training minimises cross-entropy loss: the average of -log(probability the model gave to the true next token). Lower is better. Perplexity is exp(loss) and can be read as "the model is about as unsure as if choosing uniformly among this many tokens". A perplexity of 1 means perfect prediction; 10 means as uncertain as a fair 10-way choice. Perplexity is useful to compare models on the same text and tokenizer, but a low value does not by itself prove the model is helpful or truthful.
Perplexity of the bigram model, run
I ran this plain-Python (standard library only) example. The training sentence scores 1.817 and another seen sentence 2.048. A sentence with an unseen word pair scores infinity, which is why real systems smooth probabilities, and a uniform 10-way guess scores exactly 10.
import math
text = "the cat sat on the mat . the dog sat on the log . the cat saw the dog ."
tokens = text.split()
from collections import defaultdict, Counter
counts = defaultdict(Counter)
for a, b in zip(tokens, tokens[1:]):
counts[a][b] += 1
def prob(a, b):
total = sum(counts[a].values())
return counts[a][b] / total if total else 0.0
def perplexity(seq):
logp = 0.0
for a, b in zip(seq, seq[1:]):
p = prob(a, b)
logp += math.log(p) if p > 0 else float("-inf")
return math.exp(-logp / (len(seq) - 1))
print("seen sentence :", round(perplexity("the cat sat on the mat .".split()), 3))
print("other sentence:", round(perplexity("the cat saw the dog .".split()), 3))
print("unseen pair :", perplexity("the sat cat on the mat .".split()))
print("uniform over 10:", round(math.exp(-math.log(1 / 10)), 3))
Output:
seen sentence : 1.817 other sentence: 2.048 unseen pair : inf uniform over 10: 10.0
Compare perplexity fairly
Only compare perplexities measured on the same text with the same tokenizer; different tokenizers split text differently and change the numbers.
Quick check: A perplexity of 1 means…
- The text has one token
- The model knows nothing
- The model predicts the text perfectly
- The model is biased
Answer
The model predicts the text perfectly — Perplexity 1 corresponds to assigning probability 1 to every true next token.