Lesson 28 / 29
Case Study: Taking a Support Assistant From Demo to Production
Walk through the full lifecycle for one feature.
The journey
A team builds an assistant that classifies support tickets, drafts replies from the help centre, and hands hard cases to humans. Define: targets written down (category accuracy at least 92%, faithfulness at least 95%, valid output at least 99.5%, p95 latency at most 3 s, cost at most 1.50 per resolved ticket, zero critical safety failures). Evaluation: 200 labelled tickets (English, Hindi, Hinglish, including adversarial and unanswerable), a harness with programmatic checks plus a calibrated judge (kappa measured on 80 human-labelled cases, order swapped), a held-out slice nobody tunes on, and a leakage check. Iterate: prompt versions compared by paired sign test and confidence intervals; a cascade routes easy tickets to a small model after an end-to-end check; a semantic-safe FAQ cache. Harden: schema validation with a two-step repair loop and human fallback, timeouts, retries with backoff, spend caps, security controls, tracing with a release ID on every request. Release: CI gate (quality, attack suite, latency, cost), shadow week, 2% canary, ramp with rollback rules. Operate: dashboards, PSI drift checks, weekly sampled reviews, feedback folded into the evaluation set, a system card, an incident runbook, and quarterly re-evaluation of model choice.
Measure, gate, observe, improve
A measured, versioned, observable, safely released system improves steadily instead of drifting into surprises.
The practice on one page
Each line maps to a section of this course.
Define written targets: quality, faithfulness, validity, latency, cost, safety (Sec 1)
Evaluate 200 labelled cases, harness, calibrated judge, held-out slice, leakage check, intervals (Sec 2)
Engineer versioned bundle + release ID, CI gate, schema validation + repair, FT only if needed (Sec 3)
Optimise cache, cascade (checked end to end), hedging/timeouts, capacity + unit economics (Sec 4)
Protect traces, PSI drift, layered failure handling, failure drills (Sec 5)
Release shadow -> canary -> ramp with rollback rules; hosted vs self-host decision by numbers (Sec 6)
Govern safety checklist, privacy map, system card, roles, blameless reviews (Sec 7)Write down targets before iterating
Targets chosen after seeing results are easy to bend.
Quick check: Why is the cascade checked end to end before adoption?
- Cascades never work
- Cost savings can hide quality loss in hard categories
- It is required by APIs
- To increase token use
Answer
Cost savings can hide quality loss in hard categories — Always compare quality and cost with the strong-only baseline.