Lesson 26 / 29
Evaluating Agents: Trajectories, Tools and Outcomes
Judge not only the final answer but the path taken.
Right answer, wrong way is a failure too
Agents can reach the right answer by an unsafe or wasteful path, or the wrong answer after a good one. Evaluate at three levels: outcome (is the final answer or end state correct and grounded?), trajectory (were the right tools called, in a sensible order, with correct arguments, in few steps?) and efficiency and safety (tokens, cost, latency, no forbidden actions, approvals respected). Build a dataset of realistic tasks with expected outcomes and, where useful, expected tool sequences; include adversarial and unanswerable cases. Use traces (for example LangSmith or any OpenTelemetry-compatible tool) to inspect failures, and re-run the suite on every prompt, tool or model change.
Quick check: Why evaluate the trajectory and not just the final answer?
- An agent can be right for the wrong or unsafe reasons
- Trajectories are shorter
- Final answers are never useful
- It avoids tracing
Answer
An agent can be right for the wrong or unsafe reasons — The path reveals wasted steps, wrong tools and policy violations.