# Testing and Observing Multi-Agent Systems — MCP & Agent-to-Agent Protocols

Source: https://www.geekswithgeeks.com/en/mcp-a2a/t-test

> Test contracts with fakes and trace requests across hops.

## Contracts first, then journeys

Test in layers. (1) **Contract tests** for each MCP server (tool list, schemas, error codes) and each Agent Card/endpoint, using scripted clients and **stub remote agents** that return canned task sequences, including `input-required`, failure and timeout. (2) **Routing and policy tests** with fake cards and fake tool calls, covering allow-lists, approval gates and the "no agent found" path. (3) **End-to-end evaluations** with real models on realistic scenarios, scoring outcome, tool/delegation trajectory, cost and safety. For observability, propagate a **trace id** across MCP and A2A calls, log delegations and tool calls with redaction, and monitor failure rate, loop or hop counts, latency, cost per request and approval backlog. Without tracing, a failure three hops away is nearly impossible to diagnose.

## A stub remote agent for tests (illustrative)

It plays back a scripted task sequence so the client's handling of input-required and failure can be tested offline. Not run here.

```python
class StubRemoteAgent:
    """Plays back scripted task states so client logic can be tested offline."""
    def __init__(self, script):
        self.script = list(script)
    def next_status(self):
        return self.script.pop(0)

def test_client_asks_user_when_input_required():
    agent = StubRemoteAgent(["working", "input-required", "working", "completed"])
    seen = [agent.next_status() for _ in range(4)]
    assert "input-required" in seen            # the client under test must surface a question

def test_client_gives_up_after_failure():
    agent = StubRemoteAgent(["working", "failed"])
    assert [agent.next_status() for _ in range(2)][-1] == "failed"
```

**Quiz:** Why propagate a trace id across MCP and A2A calls?

- [ ] To encrypt tokens
- [ ] To speed up the network
- [ ] To avoid logging
- [x] To follow one request through every hop when diagnosing failures

*Answer:* To follow one request through every hop when diagnosing failures. A shared id links logs from different services into one story.
