# Defining Success: Requirements and Metrics — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/i-metrics

> Turn "it should be good" into measurable targets.

## Write the acceptance test before the prompt

Before building, write down what "good" means in measurable terms. Useful dimensions: **task quality** (correctness, completeness, faithfulness to sources, tone); **safety** (no leaks, no harmful content, correct refusals); **reliability** (valid output rate, success rate under failure); **latency** (p50, p95, p99 end to end); **cost** (per request and per successful task); **coverage** (which languages, user types and topics are in scope); and **business impact** (resolution rate, time saved, conversion, user satisfaction). For each choose a **metric**, a **target** and a **way to measure** it. Distinguish **offline** metrics (on a fixed test set, before release) from **online** metrics (real traffic, after). Also decide what happens when the model is unsure: a good system can say "I don't know" or hand over to a person, and measuring that behaviour is part of quality.

## A one-page quality spec

Targets are examples; set your own from user needs.

```text
Dimension        Metric                              Target (example)          Measured by
quality          correct category on test set         >= 92%                    offline eval (200 cases)
faithfulness     claims supported by sources           >= 95%                    judge + human spot check
safety           successful attacks in red-team suite  0 critical, <= 2% minor    attack suite in CI
reliability      valid structured output               >= 99.5%                  logs
latency          end-to-end p95                        <= 3.0 s                  traces
cost             cost per resolved ticket              <= 1.50                   billing + logs
business         tickets solved without human          >= 60%                    product analytics
```

## Include the "I don't know" path

Measure how often the system correctly declines or hands over; it is part of quality.

**Quiz:** Why define metrics and targets before building the prompt?

- [ ] There is no reason
- [ ] Because prompts cannot be written otherwise
- [ ] To make documentation longer
- [x] So you can tell objectively whether a change improved the system

*Answer:* So you can tell objectively whether a change improved the system. Without a target, every change is an opinion.
