Lesson 27 / 27

Revision: Cheat Sheet and Self-Check

Review the key ideas of the whole course.

Cheat sheet

Core: an LLM predicts the next token; text is split into BPE tokens; tokens become embeddings; softmax with temperature turns logits into probabilities; greedy, top-k and top-p choose the token. Transformer: self-attention (query, key, value; softmax of scaled dot products), causal mask, stacked blocks with MLP, residuals, normalisation and positions; KV cache trades memory for speed; context window is the hard limit. Training: pretraining on next-token loss (perplexity = exp(loss)), scaling laws, SFT, LoRA, RLHF/DPO. Using: clear prompts with examples, RAG for current and citable facts, tools with validation, count cost in tokens. Quality: fixed eval set, groundedness, verify citations, low temperature for facts. Safety: prompt injection, least privilege, privacy, bias, human approval. Deploy: memory = parameters × bytes, quantization, streaming, hosted versus open models.

Quick check: A model gives wrong answers about last week's policy change. Which fix fits best?

  • Shorten the vocabulary
  • Raise the temperature
  • Fine-tune on one example
  • Add retrieval so current policy text is placed in the prompt
Answer

Add retrieval so current policy text is placed in the prompt — Fresh, citable facts belong in the prompt via retrieval, not in the weights.

Quick check: Why does softmax with a very low temperature give nearly the same token every time?

  • It deletes other tokens from the vocabulary
  • It makes the distribution very sharp around the top logit
  • It trains the model again
  • It adds more layers
Answer

It makes the distribution very sharp around the top logit — Dividing logits by a small number magnifies differences, concentrating the probability.

Quick check: What does a causal mask guarantee?

  • The output is always true
  • A position cannot use information from later positions
  • The prompt is shorter
  • The model is quantized
Answer

A position cannot use information from later positions — Future scores are set to minus infinity so their attention weights are 0.