Lesson 23 / 25
Testing Loops with a Scripted Model
Replace the real model with a script of replies so loop logic is tested fast and deterministically.
Test the loop, not the model
Real model calls are slow, costly and not repeatable, so they are bad for unit tests of your loop. Instead inject a fake model that returns a prepared sequence of replies (two tool requests, then a final answer). Then you can assert the loop stops correctly, charges the budget and enforces its limits, in milliseconds.
A scripted model and a test
The fake object has the same shape as the real reply for the fields your loop reads. The first demo ran as shown and counted 3 steps.
class FakeReply:
def __init__(self, stop_reason, text=""):
self.stop_reason, self.text = stop_reason, text
script = iter([FakeReply("tool_use"), FakeReply("tool_use"), FakeReply("end_turn", "done")])
steps = 0
for reply in script:
steps += 1
if reply.stop_reason != "tool_use":
break
print(steps)
# pytest style:
# def test_stops_after_step_limit():
# with pytest.raises(RuntimeError):
# run(messages, call_model=lambda m: FakeReply("tool_use"), max_steps=3)
Output:
3
Test the nasty cases
Write tests for a model that never stops calling tools, one that repeats the same call, and one whose tool always errors. These are the situations your guards exist for.
Quick check: Why use a fake model in unit tests of the loop?
- It makes tests fast, free and repeatable
- Real models cannot be called from tests
- Fake models are smarter
- It tests the model's quality
Answer
It makes tests fast, free and repeatable — Scripted replies make the loop's logic deterministic and cheap to verify.