Lesson 13 / 25

Building an Evaluation Set

Collect realistic, labelled cases that cover normal use, edge cases and known failures.

Data that looks like real use

An evaluation set is a collection of inputs with the expected answer or a way to judge the output. Build it from real or realistic traffic (with personal data removed), include normal requests, tricky edge cases, adversarial inputs and every failure you have ever fixed. Start with 50 to 200 cases, label them carefully, keep them separate from examples you use to tune prompts (so you do not fool yourself), and grow the set over time.

Measure before you trust

Evaluation turns "it seems fine" into numbers you can compare, track and defend.

Four steps: dataset, metric, uncertainty, comparison.
Figure 5.1 — Dataset, metric, uncertainty and comparison.

An evaluation case as data

Storing cases as data makes them easy to review, version and re-run. Tags let you report results by slice.

{
  "id": "warranty-001",
  "input": "How long is the warranty on the X200?",
  "expected": "12 months",
  "check": "contains",
  "tags": ["en", "factual", "policy"]
}

Keep a held-out set

If you keep adjusting the prompt until the same cases pass, your score stops reflecting real quality. Hold back some cases that you never tune on and use them only for final checks.

Quick check: Why keep some evaluation cases separate from tuning?

  • To save disk space
  • So the final score is not inflated by tuning to the same cases
  • Because they are secret
  • It makes tests slower
Answer

So the final score is not inflated by tuning to the same cases — Tuning on the test cases leads to an optimistic score that will not hold in production.