# Building an Eval Set — AI Agents and Tool Use

Source: https://www.geekswithgeeks.com/en/ai-agents-mcp/eval-design

> Collect realistic tasks with clear success checks and re-run them on every change.

## Tests for behaviour

Collect 20 to 50 realistic tasks, including the hard and the ambiguous ones, plus every real failure you have seen. For each, define a check: an exact value, a record in the database, a test that passes, or a rubric scored by a second model for open-ended output. Run the set whenever you change a prompt, a tool description or the model, and compare scores, cost and tool calls.

## An eval case as data

Storing cases as data makes them easy to review, version and extend.

```json
{
  "id": "refund-basic",
  "input": "Refund order 10482, the item arrived broken.",
  "must_call": ["get_order", "refund_order"],
  "must_not_call": ["delete_account"],
  "max_tool_calls": 6,
  "check": "order 10482 status == refunded"
}
```

## Include the boring cases

Hard edge cases are tempting, but most real traffic is ordinary. If the eval set holds only tricky inputs, a change that breaks everyday requests can slip through.

**Quiz:** When should the eval set be re-run?

- [ ] Only at launch
- [x] After any change to prompts, tool descriptions or the model
- [ ] Never
- [ ] Only when users complain

*Answer:* After any change to prompts, tool descriptions or the model. Any of those changes can alter behaviour, so each needs a regression check.
