# Revision: Cheat Sheet and Self-Check — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/v-revision

> Review the key ideas of the whole course.

## Cheat sheet

**Core**: an LLM predicts the next token; text is split into BPE tokens; tokens become embeddings; softmax with temperature turns logits into probabilities; greedy, top-k and top-p choose the token. **Transformer**: self-attention (query, key, value; softmax of scaled dot products), causal mask, stacked blocks with MLP, residuals, normalisation and positions; KV cache trades memory for speed; context window is the hard limit. **Training**: pretraining on next-token loss (perplexity = exp(loss)), scaling laws, SFT, LoRA, RLHF/DPO. **Using**: clear prompts with examples, RAG for current and citable facts, tools with validation, count cost in tokens. **Quality**: fixed eval set, groundedness, verify citations, low temperature for facts. **Safety**: prompt injection, least privilege, privacy, bias, human approval. **Deploy**: memory = parameters × bytes, quantization, streaming, hosted versus open models.

**Quiz:** A model gives wrong answers about last week's policy change. Which fix fits best?

- [ ] Shorten the vocabulary
- [ ] Raise the temperature
- [ ] Fine-tune on one example
- [x] Add retrieval so current policy text is placed in the prompt

*Answer:* Add retrieval so current policy text is placed in the prompt. Fresh, citable facts belong in the prompt via retrieval, not in the weights.

**Quiz:** Why does softmax with a very low temperature give nearly the same token every time?

- [ ] It deletes other tokens from the vocabulary
- [x] It makes the distribution very sharp around the top logit
- [ ] It trains the model again
- [ ] It adds more layers

*Answer:* It makes the distribution very sharp around the top logit. Dividing logits by a small number magnifies differences, concentrating the probability.

**Quiz:** What does a causal mask guarantee?

- [ ] The output is always true
- [x] A position cannot use information from later positions
- [ ] The prompt is shorter
- [ ] The model is quantized

*Answer:* A position cannot use information from later positions. Future scores are set to minus infinity so their attention weights are 0.
