Lesson 8 / 27
A Tiny Language Model You Can Run
Build a bigram model to see next-word prediction end to end.
Counting is already a language model
A bigram model predicts the next word from only the previous word by counting how often each word follows each other word in a training text. It is a real, if tiny, language model: counts become probabilities, and generating text is a loop of sampling from them. LLMs replace the counting table with a neural network that looks at thousands of previous tokens, but the interface is the same: previous text in, distribution over the next token out.
Count and generate, run
I ran this plain-Python (standard library only) example. After "the" the model has seen cat twice, dog twice, mat once and log once. With seed 3 it generated "the cat saw the dog . the cat saw".
import random
from collections import defaultdict, Counter
text = "the cat sat on the mat . the dog sat on the log . the cat saw the dog ."
tokens = text.split()
counts = defaultdict(Counter)
for a, b in zip(tokens, tokens[1:]):
counts[a][b] += 1
print("after 'the':", dict(counts["the"]))
print("after 'sat':", dict(counts["sat"]))
random.seed(3)
word, out = "the", ["the"]
for _ in range(8):
nxt = counts[word]
word = random.choices(list(nxt), list(nxt.values()))[0]
out.append(word)
print(" ".join(out))
Output:
after 'the': {'cat': 2, 'mat': 1, 'dog': 2, 'log': 1}
after 'sat': {'on': 2}
the cat saw the dog . the cat sawWhy n-grams hit a wall
A bigram sees one previous word, so it forgets everything earlier. Longer n-grams need exponentially more data. Neural networks generalise from similar contexts instead.
Quick check: What does a bigram model use to predict the next word?
- Images
- A database of facts
- Counts of which word followed the previous word
- A compiler
Answer
Counts of which word followed the previous word — It estimates probabilities from word-pair frequencies.