Lesson 13 / 27
Context Window and the KV Cache
Understand why context length is limited and how caching speeds generation.
Memory is the limit
The context window is the maximum number of tokens (prompt plus generated output) the model can consider at once. Attention compares every token with every earlier one, so naive cost grows with the square of the length. To avoid recomputing, serving systems store each earlier token's keys and values in a KV cache, so each new token only computes its own. The cache uses GPU memory that grows with sequence length, number of layers, heads and the number of simultaneous users, often becoming the limiting factor for long contexts and high throughput.
Memory arithmetic, run
I ran this plain-Python (standard library only) example. Weights alone take 2 bytes per parameter in fp16: 14 GB for a 7B model. The KV-cache line uses an example configuration (32 layers, 8 KV heads, head size 128, fp16): one 8192-token sequence needs about 1.07 GB. Real models differ; the formula is the point.
def params_gb(params_b, bytes_per):
return params_b * 1e9 * bytes_per / 1e9
for name, b in (("7B", 7), ("13B", 13), ("70B", 70)):
print(name, "fp16", params_gb(b, 2), "GB | int8", params_gb(b, 1), "GB | 4-bit", params_gb(b, 0.5), "GB")
# KV cache: 2 (K and V) * layers * kv_heads * head_dim * bytes * tokens
layers, kv_heads, head_dim, bytes_per, tokens = 32, 8, 128, 2, 8192
kv = 2 * layers * kv_heads * head_dim * bytes_per * tokens
print("KV cache for one 8192-token sequence:", round(kv / 1e9, 3), "GB")
Output:
7B fp16 14.0 GB | int8 7.0 GB | 4-bit 3.5 GB 13B fp16 26.0 GB | int8 13.0 GB | 4-bit 6.5 GB 70B fp16 140.0 GB | int8 70.0 GB | 4-bit 35.0 GB KV cache for one 8192-token sequence: 1.074 GB
Longer context is not free
A bigger window costs money and latency, and models can attend less reliably to details buried in the middle of very long prompts. Send what is relevant, not everything.
Quick check: What does a KV cache store?
- The user's password
- Keys and values of earlier tokens, to avoid recomputing them
- The training data
- The tokenizer rules
Answer
Keys and values of earlier tokens, to avoid recomputing them — Caching earlier keys and values makes each new token cheaper to generate.