Lesson 19 / 25
Evaluating Trajectories
Score the sequence of tool calls against an expected path.
The path matters
Two runs can both produce the right final answer, yet one took 3 tool calls and the other took 12, or one used a dangerous tool. A trajectory evaluation records the tool calls and scores them: the right tools, in a sensible order, without repeats or forbidden calls. Combine it with an outcome check (was the final result right?) and cost (tokens and time).
Measure the journey, not just the answer
Agents can reach the right answer the wrong way, so evaluate the path as well as the outcome.
A simple trajectory score
This ran as shown. The extra search call pushes the second score down to 0.25 because it shifts every later position.
def tool_sequence_score(expected, actual):
hits = sum(1 for e, a in zip(expected, actual) if e == a)
return round(hits / max(len(expected), len(actual)), 2)
print(tool_sequence_score(["search", "read", "write"], ["search", "read", "write"]))
print(tool_sequence_score(["search", "read", "write"], ["search", "search", "read", "write"]))
Output:
1.0 0.25
Do not over-constrain
There may be several valid paths. Score "uses forbidden tool" or "exceeds 8 calls" as hard failures, and treat exact order as a soft signal rather than demanding one fixed sequence.
Quick check: Why evaluate the tool-call path and not only the final answer?
- Paths are always identical
- It makes agents faster by itself
- Final answers cannot be checked
- A right answer can hide wasteful or unsafe steps
Answer
A right answer can hide wasteful or unsafe steps — Path scoring catches wasted cost, repeated calls and risky actions that a correct answer can mask.