# Evaluating Trajectories — AI Agents and Tool Use

Source: https://www.geekswithgeeks.com/en/ai-agents-mcp/eval-trajectories

> Score the sequence of tool calls against an expected path.

## The path matters

Two runs can both produce the right final answer, yet one took 3 tool calls and the other took 12, or one used a dangerous tool. A **trajectory evaluation** records the tool calls and scores them: the right tools, in a sensible order, without repeats or forbidden calls. Combine it with an outcome check (was the final result right?) and cost (tokens and time).

## Measure the journey, not just the answer

Agents can reach the right answer the wrong way, so evaluate the path as well as the outcome.

![Three things to score: outcome, path, cost.](assets/figures/ai-agents-mcp/section-6-map.svg) — Figure 6.1 — Outcome, path and cost.

## A simple trajectory score

This ran as shown. The extra `search` call pushes the second score down to 0.25 because it shifts every later position.

```python
def tool_sequence_score(expected, actual):
    hits = sum(1 for e, a in zip(expected, actual) if e == a)
    return round(hits / max(len(expected), len(actual)), 2)

print(tool_sequence_score(["search", "read", "write"], ["search", "read", "write"]))
print(tool_sequence_score(["search", "read", "write"], ["search", "search", "read", "write"]))
```

Output:

```
1.0
0.25
```

## Do not over-constrain

There may be several valid paths. Score "uses forbidden tool" or "exceeds 8 calls" as hard failures, and treat exact order as a soft signal rather than demanding one fixed sequence.

**Quiz:** Why evaluate the tool-call path and not only the final answer?

- [ ] Paths are always identical
- [ ] It makes agents faster by itself
- [ ] Final answers cannot be checked
- [x] A right answer can hide wasteful or unsafe steps

*Answer:* A right answer can hide wasteful or unsafe steps. Path scoring catches wasted cost, repeated calls and risky actions that a correct answer can mask.
