Lesson 25 / 25

Revision: Cheat Sheet and Self-Check

Review the safeguards, metrics and cost levers from the whole course.

Cheat sheet

Failure modes: hallucination, harmful content, privacy leak, bias, misuse, over-reliance. Safeguards: layers around the model; moderation with thresholds; mask PII; ground answers and allow "I don't know"; measure under- and over-refusal. Misuse: structural defences against injection; rate limits by tokens. Responsibility: slice metrics, data minimisation, system card, human oversight for high stakes. Evaluation: realistic held-out set; precision, recall, F1; confidence intervals; validate judges with kappa; A/B with z-test; regression in CI. Monitoring: versioned traces, percentiles, alerts. Cost: tokens × price; route; cache; trim; budgets and separate keys.

Questions interviewers ask

Be ready to explain: how you would reduce hallucinations in a Q&A bot, the precision-recall trade-off in a moderation filter, why accuracy can mislead, how you would validate an LLM judge, how to know if version B is truly better, and three ways to cut an LLM bill without hurting quality.

Quick check: A harmful-content filter is 99% accurate because only 1% of messages are harmful and it never flags anything. What is the problem?

  • It is too fast
  • Its precision is too high
  • Its recall is zero for the rare harmful class
  • Nothing is wrong
Answer

Its recall is zero for the rare harmful class — Accuracy hides the failure on the rare class; recall exposes it.

Quick check: Two versions score 90% and 88% on 25 test cases. What should you conclude?

  • There is no real evidence of a difference; gather more cases
  • Version A is clearly better
  • Version B is clearly better
  • Scores cannot be compared
Answer

There is no real evidence of a difference; gather more cases — With so few cases the confidence intervals overlap heavily, so the gap is within noise.

Quick check: Which change usually cuts LLM cost without lowering quality when done carefully?

  • Disabling monitoring
  • Removing the evaluation set
  • Doubling the context length
  • Routing easy requests to a smaller model after testing quality
Answer

Routing easy requests to a smaller model after testing quality — Routing, caching and trimming reduce spend, and your evaluation set confirms quality holds.