# Evals for Agents — Advanced Agent Workflows and Skills

Source: https://www.geekswithgeeks.com/en/agent-workflows/rel-evals

> Build a small test set of tasks with clear pass/fail rules and run it on every change.

## Tests for prompts

An **eval** is a set of example tasks with a way to score each result. Run it whenever you change a prompt, a skill or a model, so you see whether quality went up or down instead of guessing from one lucky example. Start with 10 to 20 realistic cases, including ones that failed before.

## A tiny eval harness

Each case has an input and a checker function. Prefer exact checks (contains, parses, tests pass) over asking another model.

```python
cases = [
    ("Refund order 123", lambda out: "refund" in out.lower()),
    ("Cancel order 456", lambda out: "cancel" in out.lower()),
]

passed = sum(check(run_agent(prompt)) for prompt, check in cases)
print(f"{passed}/{len(cases)} passed")
```

## Keep a failure library

Every real failure you find becomes a new eval case. Over time the set captures exactly what your users hit, and regressions are caught early.

**Quiz:** When should you run your evals?

- [ ] Never, vibes are enough
- [ ] Only once before launch
- [x] Whenever a prompt, skill or model changes
- [ ] Only when users complain

*Answer:* Whenever a prompt, skill or model changes. Re-running after each change reveals regressions immediately.
