Lesson 3 / 29
Defining Success: Requirements and Metrics
Turn "it should be good" into measurable targets.
Write the acceptance test before the prompt
Before building, write down what "good" means in measurable terms. Useful dimensions: task quality (correctness, completeness, faithfulness to sources, tone); safety (no leaks, no harmful content, correct refusals); reliability (valid output rate, success rate under failure); latency (p50, p95, p99 end to end); cost (per request and per successful task); coverage (which languages, user types and topics are in scope); and business impact (resolution rate, time saved, conversion, user satisfaction). For each choose a metric, a target and a way to measure it. Distinguish offline metrics (on a fixed test set, before release) from online metrics (real traffic, after). Also decide what happens when the model is unsure: a good system can say "I don't know" or hand over to a person, and measuring that behaviour is part of quality.
A one-page quality spec
Targets are examples; set your own from user needs.
Dimension Metric Target (example) Measured by
quality correct category on test set >= 92% offline eval (200 cases)
faithfulness claims supported by sources >= 95% judge + human spot check
safety successful attacks in red-team suite 0 critical, <= 2% minor attack suite in CI
reliability valid structured output >= 99.5% logs
latency end-to-end p95 <= 3.0 s traces
cost cost per resolved ticket <= 1.50 billing + logs
business tickets solved without human >= 60% product analyticsInclude the "I don't know" path
Measure how often the system correctly declines or hands over; it is part of quality.
Quick check: Why define metrics and targets before building the prompt?
- There is no reason
- Because prompts cannot be written otherwise
- To make documentation longer
- So you can tell objectively whether a change improved the system
Answer
So you can tell objectively whether a change improved the system — Without a target, every change is an opinion.