Lesson 29 / 29
Revision: Cheat Sheet and Self-Check
Review the key ideas of the whole course.
Cheat sheet
Mindset: LLM output is non-deterministic and hard to specify, so engineer around it: measure, version, gate, observe, bound. Define: metrics and targets (quality, faithfulness, validity, latency, cost, safety, business impact) before building; staged path prototype, eval set, iterate, harden, pilot, rollout, operate. Evaluate: harness with programmatic scorers first; categories and held-out set; confidence intervals (bootstrap) because small sets are noisy; paired sign test on disagreements; LLM-as-judge needs calibration (kappa, not raw agreement) and order swapping; avoid leakage (n-gram overlap checks). Engineer: prompt+model+settings as a versioned bundle with a release ID; CI gate on quality, safety, latency, cost; structured output with validation and a bounded repair loop; prompting, then RAG, then fine-tuning (LoRA trains a fraction of the weights). Optimise: exact, normalised and semantic caching; cascades/routers evaluated end to end; percentiles, timeouts, streaming, hedging; Little's law for capacity; cost per resolved task. Operate: traces with spans; drift (PSI) and quality sampling; layered failure handling and failure drills; hosted vs self-hosted by numbers; memory, batching, quantisation; shadow, canary, ramp, one-step rollback. Govern: safety essentials, privacy and data-flow maps, system card, roles, blameless reviews.
Quick check: Prompt B passes 3 more cases than prompt A out of 40, and A passes 2 that B fails. Is B clearly better?
- Yes, always pick the higher score
- No: 5 vs 2 disagreements is within noise; collect more cases before deciding
- Yes, because B is newer
- Neither can be evaluated
Answer
No: 5 vs 2 disagreements is within noise; collect more cases before deciding — Small gaps on small sets are indistinguishable from chance.
Quick check: Your LLM judge agrees with humans 70% of the time, but 70% of the labels are "good". What should you check?
- Whether the judge is longer than the rubric
- Nothing, 70% is great
- The font of the prompt
- Cohen's kappa, since raw agreement is inflated by the label imbalance
Answer
Cohen's kappa, since raw agreement is inflated by the label imbalance — A judge that always says "good" would also score 70% raw agreement.
Quick check: Requests are 40 per second and each takes 5 seconds on average. About how many are in flight at once?
- 200
- 8
- 45
- 0.125
Answer
200 — In-flight = rate x latency = 40 x 5.