# Case Study: Taking a Support Assistant From Demo to Production — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/z-case

> Walk through the full lifecycle for one feature.

## The journey

A team builds an assistant that classifies support tickets, drafts replies from the help centre, and hands hard cases to humans. **Define**: targets written down (category accuracy at least 92%, faithfulness at least 95%, valid output at least 99.5%, p95 latency at most 3 s, cost at most 1.50 per resolved ticket, zero critical safety failures). **Evaluation**: 200 labelled tickets (English, Hindi, Hinglish, including adversarial and unanswerable), a harness with programmatic checks plus a calibrated judge (kappa measured on 80 human-labelled cases, order swapped), a held-out slice nobody tunes on, and a leakage check. **Iterate**: prompt versions compared by paired sign test and confidence intervals; a cascade routes easy tickets to a small model after an end-to-end check; a semantic-safe FAQ cache. **Harden**: schema validation with a two-step repair loop and human fallback, timeouts, retries with backoff, spend caps, security controls, tracing with a release ID on every request. **Release**: CI gate (quality, attack suite, latency, cost), shadow week, 2% canary, ramp with rollback rules. **Operate**: dashboards, PSI drift checks, weekly sampled reviews, feedback folded into the evaluation set, a system card, an incident runbook, and quarterly re-evaluation of model choice.

## Measure, gate, observe, improve

A measured, versioned, observable, safely released system improves steadily instead of drifting into surprises.

![Four habits: measure, gate, observe, learn.](assets/figures/llm-engineering/section-8-map.svg) — Figure 8.1 — Measure, gate, observe and learn.

## The practice on one page

Each line maps to a section of this course.

```text
Define      written targets: quality, faithfulness, validity, latency, cost, safety                (Sec 1)
Evaluate    200 labelled cases, harness, calibrated judge, held-out slice, leakage check, intervals  (Sec 2)
Engineer    versioned bundle + release ID, CI gate, schema validation + repair, FT only if needed (Sec 3)
Optimise    cache, cascade (checked end to end), hedging/timeouts, capacity + unit economics      (Sec 4)
Protect     traces, PSI drift, layered failure handling, failure drills                           (Sec 5)
Release     shadow -> canary -> ramp with rollback rules; hosted vs self-host decision by numbers  (Sec 6)
Govern      safety checklist, privacy map, system card, roles, blameless reviews                  (Sec 7)
```

## Write down targets before iterating

Targets chosen after seeing results are easy to bend.

**Quiz:** Why is the cascade checked end to end before adoption?

- [ ] Cascades never work
- [x] Cost savings can hide quality loss in hard categories
- [ ] It is required by APIs
- [ ] To increase token use

*Answer:* Cost savings can hide quality loss in hard categories. Always compare quality and cost with the strong-only baseline.
