Lesson 24 / 25

Case Study: Launching a Support Assistant

Plan the safety, evaluation, monitoring and cost for a customer-support assistant before launch.

A launch checklist

Safety: PII masking, input and output moderation with a review band, answers grounded in help-centre articles with citations, "I don't know" allowed, no tools that change accounts, rate limits per user, injection tests. Evaluation: 250 labelled cases with Hindi and English slices; ship only if grounded-correct is at least 90% with the lower confidence bound above 85%, and no slice more than 8 points behind. Monitoring: trace per request, p95 latency, thumbs-down rate and moderation-block alerts. Cost: a monthly budget with alerts, prompt caching on the stable instructions, a small model for FAQ-type questions and a stronger one for the rest.

Safe, measured and affordable

A real launch needs safeguards, an evaluation gate, monitoring and a cost plan together.

Four launch gates: safe, tested, watched, budgeted.
Figure 8.1 — Safe, tested, watched and budgeted.

The launch gate on one page

Each line maps to a section of this course. Numbers are examples to adapt.

Safeguards   PII mask, moderation + review band, grounding check      (Sec 1-3)
Responsible   system card, AI disclosure, human handoff, Hindi slice tested (Sec 4)
Evaluation    250 cases; >= 90% (lower bound > 85%); slice gap <= 8 pts   (Sec 5)
Regression    eval runs in CI on every prompt/model change                 (Sec 5)
Monitoring    traces, p95 latency, feedback, moderation + cost alerts      (Sec 6)
Cost          budget + alerts, caching, routing, per-user token limits     (Sec 7)

Launch small, then widen

Release to a small share of users first, watch the dashboards and the feedback, fix what you find, and then widen. A staged rollout limits the impact of anything your tests missed.

Quick check: Why release to a small share of users first?

  • Because small launches are free
  • To avoid collecting feedback
  • To limit the impact of problems your tests missed
  • To skip evaluation
Answer

To limit the impact of problems your tests missed — Staged rollouts catch real-world problems while few people are affected.