Lesson 3 / 29

Defining Success: Requirements and Metrics

Turn "it should be good" into measurable targets.

Write the acceptance test before the prompt

Before building, write down what "good" means in measurable terms. Useful dimensions: task quality (correctness, completeness, faithfulness to sources, tone); safety (no leaks, no harmful content, correct refusals); reliability (valid output rate, success rate under failure); latency (p50, p95, p99 end to end); cost (per request and per successful task); coverage (which languages, user types and topics are in scope); and business impact (resolution rate, time saved, conversion, user satisfaction). For each choose a metric, a target and a way to measure it. Distinguish offline metrics (on a fixed test set, before release) from online metrics (real traffic, after). Also decide what happens when the model is unsure: a good system can say "I don't know" or hand over to a person, and measuring that behaviour is part of quality.

A one-page quality spec

Targets are examples; set your own from user needs.

Dimension        Metric                              Target (example)          Measured by
quality          correct category on test set         >= 92%                    offline eval (200 cases)
faithfulness     claims supported by sources           >= 95%                    judge + human spot check
safety           successful attacks in red-team suite  0 critical, <= 2% minor    attack suite in CI
reliability      valid structured output               >= 99.5%                  logs
latency          end-to-end p95                        <= 3.0 s                  traces
cost             cost per resolved ticket              <= 1.50                   billing + logs
business         tickets solved without human          >= 60%                    product analytics

Include the "I don't know" path

Measure how often the system correctly declines or hands over; it is part of quality.

Quick check: Why define metrics and targets before building the prompt?

  • There is no reason
  • Because prompts cannot be written otherwise
  • To make documentation longer
  • So you can tell objectively whether a change improved the system
Answer

So you can tell objectively whether a change improved the system — Without a target, every change is an opinion.