पाठ 10 / 27

Self-Attention: Queries, Keys, Values

Attention weights और भारित योग चरण-दर-चरण निकालें।

मैं किसकी सुनूँ

हर token तीन vectors बनाता है: query (मैं क्या खोज रहा हूँ), key (मैं क्या देता हूँ) और value (मैं कौन-सी जानकारी लिए हूँ)। token i और token j के बीच का score query i और key j का dot product है, जिसे संख्याओं को स्थिर सीमा में रखने को sqrt(d) से भाग दिया जाता है। Softmax scores को attention weights बनाता है जिनका योग 1 है, और token i का आउटपुट सभी values का भारित योग है। असली मॉडल में Q, K, V vectors embeddings के सीखे हुए matrix गुणनफल से आते हैं, और अलग-अलग संबंध पकड़ने को कई heads समानांतर चलते हैं।

Tokens एक-दूसरे को देखते हैं

Attention हर token को दूसरों से जानकारी जुटाने देता है; परतें उसे निखारती हैं।

चार विचार: attend, mask, stack, cache।
चित्र 3.1 — Attend, mask, stack और cache।

हाथ से attention, चलाकर

मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। Token 0 अपना attention tokens 0, 1, 2 पर 40% / 20% / 40% बाँटता है और आउटपुट [3.0, 4.0] देता है, value vectors का भारित मिश्रण।

import math

def softmax(xs):
    m = max(xs); e = [math.exp(x - m) for x in xs]; s = sum(e)
    return [v / s for v in e]

def dot(a, b): return sum(x * y for x, y in zip(a, b))

# 3 tokens, each with a 2-d query/key/value vector
Q = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
K = [[1.0, 0.0], [0.0, 1.0], [1.0, 1.0]]
V = [[1.0, 2.0], [3.0, 4.0], [5.0, 6.0]]
d = 2

for i, q in enumerate(Q):
    scores = [dot(q, k) / math.sqrt(d) for k in K]
    weights = softmax(scores)
    out = [sum(w * v[j] for w, v in zip(weights, V)) for j in range(2)]
    print(i, "weights", [round(w, 3) for w in weights], "output", [round(o, 3) for o in out])

Output:

0 weights [0.401, 0.198, 0.401] output [3.0, 4.0]
1 weights [0.198, 0.401, 0.401] output [3.407, 4.407]
2 weights [0.248, 0.248, 0.503] output [3.51, 4.51]

Heads अलग चीज़ें खोजते हैं

Multi-head attention कई attention गणनाएँ समानांतर चलाता है, हर एक अलग संबंध, जैसे व्याकरण या संदर्भ, सीखने को स्वतंत्र।

त्वरित जाँच: Attention में softmax weights किससे गुणा होते हैं?

  • Value vectors
  • Tokenizer
  • GPU
  • Prompt
Answer

Value vectors — आउटपुट values का भारित योग है, weights query-key scores से आते हैं।