Lesson 10 / 27
Self-Attention: Queries, Keys, Values
Compute attention weights and weighted sums step by step.
Who should I listen to
Each token produces three vectors: a query (what I am looking for), a key (what I offer) and a value (the information I carry). The score between token i and token j is the dot product of query i and key j, divided by sqrt(d) to keep numbers in a stable range. Softmax turns the scores into attention weights that sum to 1, and the output for token i is the weighted sum of all values. In a real model the Q, K and V vectors come from learned matrix multiplications of the embeddings, and many heads run in parallel to capture different relationships.
Tokens look at each other
Attention lets each token gather information from the others; stacked layers refine it.
Attention by hand, run
I ran this plain-Python (standard library only) example. Token 0 splits its attention 40% / 20% / 40% over tokens 0, 1 and 2 and outputs [3.0, 4.0], the weighted blend of the value vectors.
import math
def softmax(xs):
m = max(xs); e = [math.exp(x - m) for x in xs]; s = sum(e)
return [v / s for v in e]
def dot(a, b): return sum(x * y for x, y in zip(a, b))
# 3 tokens, each with a 2-d query/key/value vector
Q = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
K = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
V = [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]]
d = 2
for i, q in enumerate(Q):
scores = [dot(q, k) / math.sqrt(d) for k in K]
weights = softmax(scores)
out = [sum(w * v[j] for w, v in zip(weights, V)) for j in range(2)]
print(i, "weights", [round(w, 3) for w in weights], "output", [round(o, 3) for o in out])
Output:
0 weights [0.401, 0.198, 0.401] output [3.0, 4.0] 1 weights [0.198, 0.401, 0.401] output [3.407, 4.407] 2 weights [0.248, 0.248, 0.503] output [3.51, 4.51]
Heads look for different things
Multi-head attention runs several attention computations in parallel, each free to learn a different relationship such as syntax or reference.
Quick check: In attention, what do the softmax weights multiply?
- The value vectors
- The tokenizer
- The GPU
- The prompt
Answer
The value vectors — The output is a weighted sum of values, with weights from query-key scores.