# Revision: Cheat Sheet and Self-Check — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/z-revision

> Review the key ideas of the whole course.

## Cheat sheet

**Mindset**: LLM output is non-deterministic and hard to specify, so engineer around it: measure, version, gate, observe, bound. **Define**: metrics and targets (quality, faithfulness, validity, latency, cost, safety, business impact) before building; staged path prototype, eval set, iterate, harden, pilot, rollout, operate. **Evaluate**: harness with programmatic scorers first; categories and held-out set; confidence intervals (bootstrap) because small sets are noisy; paired sign test on disagreements; LLM-as-judge needs calibration (kappa, not raw agreement) and order swapping; avoid leakage (n-gram overlap checks). **Engineer**: prompt+model+settings as a versioned bundle with a release ID; CI gate on quality, safety, latency, cost; structured output with validation and a bounded repair loop; prompting, then RAG, then fine-tuning (LoRA trains a fraction of the weights). **Optimise**: exact, normalised and semantic caching; cascades/routers evaluated end to end; percentiles, timeouts, streaming, hedging; Little's law for capacity; cost per resolved task. **Operate**: traces with spans; drift (PSI) and quality sampling; layered failure handling and failure drills; hosted vs self-hosted by numbers; memory, batching, quantisation; shadow, canary, ramp, one-step rollback. **Govern**: safety essentials, privacy and data-flow maps, system card, roles, blameless reviews.

**Quiz:** Prompt B passes 3 more cases than prompt A out of 40, and A passes 2 that B fails. Is B clearly better?

- [ ] Yes, always pick the higher score
- [x] No: 5 vs 2 disagreements is within noise; collect more cases before deciding
- [ ] Yes, because B is newer
- [ ] Neither can be evaluated

*Answer:* No: 5 vs 2 disagreements is within noise; collect more cases before deciding. Small gaps on small sets are indistinguishable from chance.

**Quiz:** Your LLM judge agrees with humans 70% of the time, but 70% of the labels are "good". What should you check?

- [ ] Whether the judge is longer than the rubric
- [ ] Nothing, 70% is great
- [ ] The font of the prompt
- [x] Cohen's kappa, since raw agreement is inflated by the label imbalance

*Answer:* Cohen's kappa, since raw agreement is inflated by the label imbalance. A judge that always says "good" would also score 70% raw agreement.

**Quiz:** Requests are 40 per second and each takes 5 seconds on average. About how many are in flight at once?

- [x] 200
- [ ] 8
- [ ] 45
- [ ] 0.125

*Answer:* 200. In-flight = rate x latency = 40 x 5.
