Lesson 22 / 25
Testing Servers and Agents
Unit-test tool functions and add a small end-to-end eval for the agent.
Two layers of tests
Because MCP tools are plain functions, unit-test them directly like any code: normal input, bad input, edge cases. Then add a small end-to-end eval: a handful of realistic user requests, run through the real agent, with checks such as "called search_issues with a sensible keyword". Re-run it whenever you change a tool description, because that changes model behaviour.
A pytest unit test
Tests call the function directly, with no model or protocol involved, so they are fast and deterministic.
import pytest
from server import search_issues
def test_limit_must_be_in_range():
with pytest.raises(ValueError):
search_issues("crash", limit=0)
def test_returns_at_most_limit(fake_tracker):
assert len(search_issues("crash", limit=3)) <= 3Quick check: Why re-run evals after editing a tool description?
- Descriptions affect how the model chooses and uses the tool
- Python requires it
- Descriptions are compiled
- It resets the server
Answer
Descriptions affect how the model chooses and uses the tool — The description is part of the prompt the model sees, so changing it can change behaviour.