Lesson 25 / 25
Revision: Cheat Sheet and Self-Check
Review the safeguards, metrics and cost levers from the whole course.
Cheat sheet
Failure modes: hallucination, harmful content, privacy leak, bias, misuse, over-reliance. Safeguards: layers around the model; moderation with thresholds; mask PII; ground answers and allow "I don't know"; measure under- and over-refusal. Misuse: structural defences against injection; rate limits by tokens. Responsibility: slice metrics, data minimisation, system card, human oversight for high stakes. Evaluation: realistic held-out set; precision, recall, F1; confidence intervals; validate judges with kappa; A/B with z-test; regression in CI. Monitoring: versioned traces, percentiles, alerts. Cost: tokens × price; route; cache; trim; budgets and separate keys.
Questions interviewers ask
Be ready to explain: how you would reduce hallucinations in a Q&A bot, the precision-recall trade-off in a moderation filter, why accuracy can mislead, how you would validate an LLM judge, how to know if version B is truly better, and three ways to cut an LLM bill without hurting quality.
Quick check: A harmful-content filter is 99% accurate because only 1% of messages are harmful and it never flags anything. What is the problem?
- It is too fast
- Its precision is too high
- Its recall is zero for the rare harmful class
- Nothing is wrong
Answer
Its recall is zero for the rare harmful class — Accuracy hides the failure on the rare class; recall exposes it.
Quick check: Two versions score 90% and 88% on 25 test cases. What should you conclude?
- There is no real evidence of a difference; gather more cases
- Version A is clearly better
- Version B is clearly better
- Scores cannot be compared
Answer
There is no real evidence of a difference; gather more cases — With so few cases the confidence intervals overlap heavily, so the gap is within noise.
Quick check: Which change usually cuts LLM cost without lowering quality when done carefully?
- Disabling monitoring
- Removing the evaluation set
- Doubling the context length
- Routing easy requests to a smaller model after testing quality
Answer
Routing easy requests to a smaller model after testing quality — Routing, caching and trimming reduce spend, and your evaluation set confirms quality holds.