# Layers, Feed-Forward Blocks and Positions — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/t-architecture

> See how attention, MLP, residuals and position information fit together.

## Stacking the same block

A transformer is a stack of identical blocks. Each block has (1) **multi-head self-attention** to mix information between tokens, (2) a **feed-forward network (MLP)** applied to each token separately to transform it, (3) **residual connections** that add the block's input back to its output, and (4) **normalisation** layers for stable training. Because attention alone ignores order, **positional information** (such as rotary position embeddings) is added so the model knows word order. Stacking dozens of blocks lets early layers capture simple patterns and later layers more abstract ones. A final linear layer produces the logits over the vocabulary.

## Block diagram

Read from bottom to top; the same block repeats N times.

```text
logits over vocabulary
      ^
[ final norm + linear ]
      ^
+-- block x N ------------------+
| x + MultiHeadAttention(norm(x)) |
| x + FeedForward(norm(x))        |
+--------------------------------+
      ^
[ token embeddings + positions ]
      ^
token ids
```

## Depth and width

Models differ in number of layers (depth), hidden size (width) and attention heads. Together with vocabulary size they determine the parameter count.

**Quiz:** Why are positions added to a transformer?

- [ ] To reduce the vocabulary
- [x] Attention alone does not know word order
- [ ] To compress images
- [ ] To encrypt text

*Answer:* Attention alone does not know word order. Without positional information the same words in any order look identical.
