# Tokens and Byte-Pair Encoding — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/f-tokens

> Understand how text is split into tokens and how BPE builds a vocabulary.

## Pieces, not words

Models do not read letters or whole words but **tokens**: frequent chunks such as `the`, ` ing`, `Hel` or a single character. Most modern tokenizers use a form of **byte-pair encoding (BPE)**: start with single characters, repeatedly merge the most frequent adjacent pair into a new symbol, and stop at a chosen vocabulary size (often tens of thousands to a few hundred thousand tokens). Common words become one token; rare words split into pieces. For English a token is roughly three-quarters of a word, but other scripts, including Hindi, often need more tokens for the same meaning, which affects cost and context length.

## Learning merges, run

I ran this plain-Python (standard library only) example. On a tiny corpus the first merges are `e+s` then `es+t` (from newest/widest) and `l+o`, `lo+w` (from low/lower). The final list shows each word as the tokens BPE would use.

```python
from collections import Counter

def get_pairs(words):
    pairs = Counter()
    for w, f in words.items():
        for a, b in zip(w, w[1:]):
            pairs[(a, b)] += f
    return pairs

def merge(words, pair):
    out = {}
    for w, f in words.items():
        new, i = [], 0
        while i < len(w):
            if i < len(w) - 1 and (w[i], w[i + 1]) == pair:
                new.append(w[i] + w[i + 1]); i += 2
            else:
                new.append(w[i]); i += 1
        out[tuple(new)] = f
    return out

corpus = {"low": 5, "lower": 2, "newest": 6, "widest": 3}
words = {tuple(w): f for w, f in corpus.items()}
for step in range(5):
    pair = get_pairs(words).most_common(1)[0][0]
    words = merge(words, pair)
    print(step + 1, "merge", pair)
print(list(words))

```

Output:

```
1 merge ('e', 's')
2 merge ('es', 't')
3 merge ('l', 'o')
4 merge ('lo', 'w')
5 merge ('n', 'e')
[('low',), ('low', 'e', 'r'), ('ne', 'w', 'est'), ('w', 'i', 'd', 'est')]
```

## Count tokens, not words

Limits and prices are in tokens. Use your provider's tokenizer to count them; do not estimate from word counts when precision matters.

**Quiz:** What does BPE repeatedly do?

- [ ] Sort the alphabet
- [ ] Delete rare letters
- [ ] Translate words
- [x] Merge the most frequent adjacent pair into a new token

*Answer:* Merge the most frequent adjacent pair into a new token. Merging frequent pairs builds a vocabulary of useful chunks.
