Lesson 20 / 25
Building an Eval Set
Collect realistic tasks with clear success checks and re-run them on every change.
Tests for behaviour
Collect 20 to 50 realistic tasks, including the hard and the ambiguous ones, plus every real failure you have seen. For each, define a check: an exact value, a record in the database, a test that passes, or a rubric scored by a second model for open-ended output. Run the set whenever you change a prompt, a tool description or the model, and compare scores, cost and tool calls.
An eval case as data
Storing cases as data makes them easy to review, version and extend.
{
"id": "refund-basic",
"input": "Refund order 10482, the item arrived broken.",
"must_call": ["get_order", "refund_order"],
"must_not_call": ["delete_account"],
"max_tool_calls": 6,
"check": "order 10482 status == refunded"
}Include the boring cases
Hard edge cases are tempting, but most real traffic is ordinary. If the eval set holds only tricky inputs, a change that breaks everyday requests can slip through.
Quick check: When should the eval set be re-run?
- Only at launch
- After any change to prompts, tool descriptions or the model
- Never
- Only when users complain
Answer
After any change to prompts, tool descriptions or the model — Any of those changes can alter behaviour, so each needs a regression check.