# Loss and Perplexity — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/n-perplexity

> Measure how well a language model predicts text.

## How surprised is the model

Training minimises **cross-entropy loss**: the average of `-log(probability the model gave to the true next token)`. Lower is better. **Perplexity** is `exp(loss)` and can be read as "the model is about as unsure as if choosing uniformly among this many tokens". A perplexity of 1 means perfect prediction; 10 means as uncertain as a fair 10-way choice. Perplexity is useful to compare models on the same text and tokenizer, but a low value does not by itself prove the model is helpful or truthful.

## Perplexity of the bigram model, run

I ran this plain-Python (standard library only) example. The training sentence scores 1.817 and another seen sentence 2.048. A sentence with an unseen word pair scores infinity, which is why real systems smooth probabilities, and a uniform 10-way guess scores exactly 10.

```python
import math

text = "the cat sat on the mat . the dog sat on the log . the cat saw the dog ."
tokens = text.split()
from collections import defaultdict, Counter
counts = defaultdict(Counter)
for a, b in zip(tokens, tokens[1:]):
    counts[a][b] += 1

def prob(a, b):
    total = sum(counts[a].values())
    return counts[a][b] / total if total else 0.0

def perplexity(seq):
    logp = 0.0
    for a, b in zip(seq, seq[1:]):
        p = prob(a, b)
        logp += math.log(p) if p > 0 else float("-inf")
    return math.exp(-logp / (len(seq) - 1))

print("seen sentence  :", round(perplexity("the cat sat on the mat .".split()), 3))
print("other sentence:", round(perplexity("the cat saw the dog .".split()), 3))
print("unseen pair   :", perplexity("the sat cat on the mat .".split()))
print("uniform over 10:", round(math.exp(-math.log(1 / 10)), 3))

```

Output:

```
seen sentence  : 1.817
other sentence: 2.048
unseen pair   : inf
uniform over 10: 10.0
```

## Compare perplexity fairly

Only compare perplexities measured on the same text with the same tokenizer; different tokenizers split text differently and change the numbers.

**Quiz:** A perplexity of 1 means…

- [ ] The text has one token
- [ ] The model knows nothing
- [x] The model predicts the text perfectly
- [ ] The model is biased

*Answer:* The model predicts the text perfectly. Perplexity 1 corresponds to assigning probability 1 to every true next token.
