Lesson 21 / 25

Evals for Agents

Build a small test set of tasks with clear pass/fail rules and run it on every change.

Tests for prompts

An eval is a set of example tasks with a way to score each result. Run it whenever you change a prompt, a skill or a model, so you see whether quality went up or down instead of guessing from one lucky example. Start with 10 to 20 realistic cases, including ones that failed before.

A tiny eval harness

Each case has an input and a checker function. Prefer exact checks (contains, parses, tests pass) over asking another model.

cases = [
    ("Refund order 123", lambda out: "refund" in out.lower()),
    ("Cancel order 456", lambda out: "cancel" in out.lower()),
]

passed = sum(check(run_agent(prompt)) for prompt, check in cases)
print(f"{passed}/{len(cases)} passed")

Keep a failure library

Every real failure you find becomes a new eval case. Over time the set captures exactly what your users hit, and regressions are caught early.

Quick check: When should you run your evals?

  • Never, vibes are enough
  • Only once before launch
  • Whenever a prompt, skill or model changes
  • Only when users complain
Answer

Whenever a prompt, skill or model changes — Re-running after each change reveals regressions immediately.