Lesson 12 / 27

Layers, Feed-Forward Blocks and Positions

See how attention, MLP, residuals and position information fit together.

Stacking the same block

A transformer is a stack of identical blocks. Each block has (1) multi-head self-attention to mix information between tokens, (2) a feed-forward network (MLP) applied to each token separately to transform it, (3) residual connections that add the block's input back to its output, and (4) normalisation layers for stable training. Because attention alone ignores order, positional information (such as rotary position embeddings) is added so the model knows word order. Stacking dozens of blocks lets early layers capture simple patterns and later layers more abstract ones. A final linear layer produces the logits over the vocabulary.

Block diagram

Read from bottom to top; the same block repeats N times.

logits over vocabulary
      ^
[ final norm + linear ]
      ^
+-- block x N ------------------+
| x + MultiHeadAttention(norm(x)) |
| x + FeedForward(norm(x))        |
+--------------------------------+
      ^
[ token embeddings + positions ]
      ^
token ids

Depth and width

Models differ in number of layers (depth), hidden size (width) and attention heads. Together with vocabulary size they determine the parameter count.

Quick check: Why are positions added to a transformer?

  • To reduce the vocabulary
  • Attention alone does not know word order
  • To compress images
  • To encrypt text
Answer

Attention alone does not know word order — Without positional information the same words in any order look identical.