# Self-Attention: Queries, Keys, Values — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/t-attention

> Compute attention weights and weighted sums step by step.

## Who should I listen to

Each token produces three vectors: a **query** (what I am looking for), a **key** (what I offer) and a **value** (the information I carry). The score between token i and token j is the dot product of query i and key j, divided by `sqrt(d)` to keep numbers in a stable range. Softmax turns the scores into **attention weights** that sum to 1, and the output for token i is the weighted sum of all values. In a real model the Q, K and V vectors come from learned matrix multiplications of the embeddings, and many **heads** run in parallel to capture different relationships.

## Tokens look at each other

Attention lets each token gather information from the others; stacked layers refine it.

![Four ideas: attend, mask, stack, cache.](assets/figures/llms/section-3-map.svg) — Figure 3.1 — Attend, mask, stack and cache.

## Attention by hand, run

I ran this plain-Python (standard library only) example. Token 0 splits its attention 40% / 20% / 40% over tokens 0, 1 and 2 and outputs [3.0, 4.0], the weighted blend of the value vectors.

```python
import math

def softmax(xs):
    m = max(xs); e = [math.exp(x - m) for x in xs]; s = sum(e)
    return [v / s for v in e]

def dot(a, b): return sum(x * y for x, y in zip(a, b))

# 3 tokens, each with a 2-d query/key/value vector
Q = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
K = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
V = [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]]
d = 2

for i, q in enumerate(Q):
    scores = [dot(q, k) / math.sqrt(d) for k in K]
    weights = softmax(scores)
    out = [sum(w * v[j] for w, v in zip(weights, V)) for j in range(2)]
    print(i, "weights", [round(w, 3) for w in weights], "output", [round(o, 3) for o in out])

```

Output:

```
0 weights [0.401, 0.198, 0.401] output [3.0, 4.0]
1 weights [0.198, 0.401, 0.401] output [3.407, 4.407]
2 weights [0.248, 0.248, 0.503] output [3.51, 4.51]
```

## Heads look for different things

Multi-head attention runs several attention computations in parallel, each free to learn a different relationship such as syntax or reference.

**Quiz:** In attention, what do the softmax weights multiply?

- [x] The value vectors
- [ ] The tokenizer
- [ ] The GPU
- [ ] The prompt

*Answer:* The value vectors. The output is a weighted sum of values, with weights from query-key scores.
