# Causal Masking and Generation — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/t-causal

> Stop tokens from seeing the future during training and generation.

## No peeking ahead

A text-generating (decoder) model must predict token t using only tokens before it. A **causal mask** sets the scores for future positions to minus infinity before softmax, so their weights become exactly 0. This also lets the model be trained on every position of a sequence in parallel: each position is a separate prediction task, all computed in one pass. At generation time tokens are produced one by one, each step reading the earlier ones.

## A causal mask, run

I ran this plain-Python (standard library only) example. Row 0 can only attend to itself (1.0), row 1 to tokens 0 and 1, and row 2 to all three. The upper-right triangle is always 0.

```python
import math

def softmax(xs):
    m = max(xs); e = [math.exp(x - m) for x in xs]; s = sum(e)
    return [v / s for v in e]

scores = [[0.5, 1.0, 2.0], [1.5, 0.2, 0.1], [0.3, 0.9, 1.2]]
for i, row in enumerate(scores):
    masked = [s if j <= i else float("-inf") for j, s in enumerate(row)]
    print(i, [round(w, 3) for w in softmax(masked)])

```

Output:

```
0 [1.0, 0.0, 0.0]
1 [0.786, 0.214, 0.0]
2 [0.189, 0.345, 0.466]
```

## Why training is parallel but generation is not

In training the whole text is known, so every position is predicted at once. In generation each new token depends on the one before, so steps run one after another.

**Quiz:** What does the causal mask prevent?

- [ ] Embedding
- [ ] Tokenisation
- [ ] Softmax
- [x] A token attending to later tokens

*Answer:* A token attending to later tokens. The model must predict the next token without seeing it.
