Lesson 27 / 27
Revision: Cheat Sheet and Self-Check
Review the key ideas of the whole course.
Cheat sheet
Core: an LLM predicts the next token; text is split into BPE tokens; tokens become embeddings; softmax with temperature turns logits into probabilities; greedy, top-k and top-p choose the token. Transformer: self-attention (query, key, value; softmax of scaled dot products), causal mask, stacked blocks with MLP, residuals, normalisation and positions; KV cache trades memory for speed; context window is the hard limit. Training: pretraining on next-token loss (perplexity = exp(loss)), scaling laws, SFT, LoRA, RLHF/DPO. Using: clear prompts with examples, RAG for current and citable facts, tools with validation, count cost in tokens. Quality: fixed eval set, groundedness, verify citations, low temperature for facts. Safety: prompt injection, least privilege, privacy, bias, human approval. Deploy: memory = parameters × bytes, quantization, streaming, hosted versus open models.
Quick check: A model gives wrong answers about last week's policy change. Which fix fits best?
- Shorten the vocabulary
- Raise the temperature
- Fine-tune on one example
- Add retrieval so current policy text is placed in the prompt
Answer
Add retrieval so current policy text is placed in the prompt — Fresh, citable facts belong in the prompt via retrieval, not in the weights.
Quick check: Why does softmax with a very low temperature give nearly the same token every time?
- It deletes other tokens from the vocabulary
- It makes the distribution very sharp around the top logit
- It trains the model again
- It adds more layers
Answer
It makes the distribution very sharp around the top logit — Dividing logits by a small number magnifies differences, concentrating the probability.
Quick check: What does a causal mask guarantee?
- The output is always true
- A position cannot use information from later positions
- The prompt is shorter
- The model is quantized
Answer
A position cannot use information from later positions — Future scores are set to minus infinity so their attention weights are 0.