# A Tiny Language Model You Can Run — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/n-bigram

> Build a bigram model to see next-word prediction end to end.

## Counting is already a language model

A **bigram model** predicts the next word from only the previous word by counting how often each word follows each other word in a training text. It is a real, if tiny, language model: counts become probabilities, and generating text is a loop of sampling from them. LLMs replace the counting table with a neural network that looks at thousands of previous tokens, but the interface is the same: previous text in, distribution over the next token out.

## Count and generate, run

I ran this plain-Python (standard library only) example. After "the" the model has seen cat twice, dog twice, mat once and log once. With seed 3 it generated "the cat saw the dog . the cat saw".

```python
import random
from collections import defaultdict, Counter

text = "the cat sat on the mat . the dog sat on the log . the cat saw the dog ."
tokens = text.split()
counts = defaultdict(Counter)
for a, b in zip(tokens, tokens[1:]):
    counts[a][b] += 1

print("after 'the':", dict(counts["the"]))
print("after 'sat':", dict(counts["sat"]))

random.seed(3)
word, out = "the", ["the"]
for _ in range(8):
    nxt = counts[word]
    word = random.choices(list(nxt), list(nxt.values()))[0]
    out.append(word)
print(" ".join(out))

```

Output:

```
after 'the': {'cat': 2, 'mat': 1, 'dog': 2, 'log': 1}
after 'sat': {'on': 2}
the cat saw the dog . the cat saw
```

## Why n-grams hit a wall

A bigram sees one previous word, so it forgets everything earlier. Longer n-grams need exponentially more data. Neural networks generalise from similar contexts instead.

**Quiz:** What does a bigram model use to predict the next word?

- [ ] Images
- [ ] A database of facts
- [x] Counts of which word followed the previous word
- [ ] A compiler

*Answer:* Counts of which word followed the previous word. It estimates probabilities from word-pair frequencies.
