Lesson 22 / 25

Testing Servers and Agents

Unit-test tool functions and add a small end-to-end eval for the agent.

Two layers of tests

Because MCP tools are plain functions, unit-test them directly like any code: normal input, bad input, edge cases. Then add a small end-to-end eval: a handful of realistic user requests, run through the real agent, with checks such as "called search_issues with a sensible keyword". Re-run it whenever you change a tool description, because that changes model behaviour.

A pytest unit test

Tests call the function directly, with no model or protocol involved, so they are fast and deterministic.

import pytest
from server import search_issues

def test_limit_must_be_in_range():
    with pytest.raises(ValueError):
        search_issues("crash", limit=0)

def test_returns_at_most_limit(fake_tracker):
    assert len(search_issues("crash", limit=3)) <= 3

Quick check: Why re-run evals after editing a tool description?

  • Descriptions affect how the model chooses and uses the tool
  • Python requires it
  • Descriptions are compiled
  • It resets the server
Answer

Descriptions affect how the model chooses and uses the tool — The description is part of the prompt the model sees, so changing it can change behaviour.