Lesson 21 / 25
Evals for Agents
Build a small test set of tasks with clear pass/fail rules and run it on every change.
Tests for prompts
An eval is a set of example tasks with a way to score each result. Run it whenever you change a prompt, a skill or a model, so you see whether quality went up or down instead of guessing from one lucky example. Start with 10 to 20 realistic cases, including ones that failed before.
A tiny eval harness
Each case has an input and a checker function. Prefer exact checks (contains, parses, tests pass) over asking another model.
cases = [
("Refund order 123", lambda out: "refund" in out.lower()),
("Cancel order 456", lambda out: "cancel" in out.lower()),
]
passed = sum(check(run_agent(prompt)) for prompt, check in cases)
print(f"{passed}/{len(cases)} passed")Keep a failure library
Every real failure you find becomes a new eval case. Over time the set captures exactly what your users hit, and regressions are caught early.
Quick check: When should you run your evals?
- Never, vibes are enough
- Only once before launch
- Whenever a prompt, skill or model changes
- Only when users complain
Answer
Whenever a prompt, skill or model changes — Re-running after each change reveals regressions immediately.