Lesson 10 / 27

Self-Attention: Queries, Keys, Values

Compute attention weights and weighted sums step by step.

Who should I listen to

Each token produces three vectors: a query (what I am looking for), a key (what I offer) and a value (the information I carry). The score between token i and token j is the dot product of query i and key j, divided by sqrt(d) to keep numbers in a stable range. Softmax turns the scores into attention weights that sum to 1, and the output for token i is the weighted sum of all values. In a real model the Q, K and V vectors come from learned matrix multiplications of the embeddings, and many heads run in parallel to capture different relationships.

Tokens look at each other

Attention lets each token gather information from the others; stacked layers refine it.

Four ideas: attend, mask, stack, cache.
Figure 3.1 — Attend, mask, stack and cache.

Attention by hand, run

I ran this plain-Python (standard library only) example. Token 0 splits its attention 40% / 20% / 40% over tokens 0, 1 and 2 and outputs [3.0, 4.0], the weighted blend of the value vectors.

import math

def softmax(xs):
    m = max(xs); e = [math.exp(x - m) for x in xs]; s = sum(e)
    return [v / s for v in e]

def dot(a, b): return sum(x * y for x, y in zip(a, b))

# 3 tokens, each with a 2-d query/key/value vector
Q = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
K = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
V = [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]]
d = 2

for i, q in enumerate(Q):
    scores = [dot(q, k) / math.sqrt(d) for k in K]
    weights = softmax(scores)
    out = [sum(w * v[j] for w, v in zip(weights, V)) for j in range(2)]
    print(i, "weights", [round(w, 3) for w in weights], "output", [round(o, 3) for o in out])

Output:

0 weights [0.401, 0.198, 0.401] output [3.0, 4.0]
1 weights [0.198, 0.401, 0.401] output [3.407, 4.407]
2 weights [0.248, 0.248, 0.503] output [3.51, 4.51]

Heads look for different things

Multi-head attention runs several attention computations in parallel, each free to learn a different relationship such as syntax or reference.

Quick check: In attention, what do the softmax weights multiply?

  • The value vectors
  • The tokenizer
  • The GPU
  • The prompt
Answer

The value vectors — The output is a weighted sum of values, with weights from query-key scores.