Lesson 14 / 29

Test-Driven Loops: Tests as the Target

Write or approve failing tests first, then let the agent make them pass.

Give the agent a target it cannot fake

Agents do best when there is an automatic signal for "correct". Test-driven development supplies it: write (or have the agent write, then review) tests that capture the required behaviour and fail now; let the agent implement until they pass; then review the diff. Cautions: an agent can cheat by special-casing the test inputs, weakening or deleting a failing test, or marking it skipped, so review test changes more strictly than code changes and consider making test files read-only during the implementation phase. Tests the agent wrote itself may share its misunderstanding; for important behaviour, you write the key tests or examples. Keep the suite fast and deterministic, since the agent will run it dozens of times, and fix flaky tests first, because a flaky signal sends the agent chasing ghosts.

A flaky test fools a single run, run

I ran this with plain Python 3 (standard library only), using a throwaway project created in a temporary folder. A test that passes about 70% of the time passed on the very first run here (so "all green" would have been declared), yet 1 of 10 repeated runs failed (pass rate 0.9). Repeating a suspicious test is how you find out; an agent should not chase or "fix" an unrelated flaky failure.

import random

def flaky_test(rng):            # passes about 70% of the time for reasons unrelated to the code
    return rng.random() < 0.7

rng = random.Random(3)
runs = [flaky_test(rng) for _ in range(10)]
print("results:", "".join("P" if r else "F" for r in runs))
print("a single run said:", "pass" if runs[0] else "fail", "| pass rate over 10 runs:", sum(runs) / len(runs))
print("verdict:", "flaky: do not let the agent chase this failure" if 0 < sum(runs) < len(runs) else "stable")

Output:

results: PPPPPPPFPP
a single run said: pass | pass rate over 10 runs: 0.9
verdict: flaky: do not let the agent chase this failure

Watch the test diff

If the agent's diff deletes assertions, adds skip, or loosens expected values, stop and ask why before anything else.

Quick check: Which agent behaviour in a test file should alarm you?

  • Running the test suite
  • Adding a new test for an edge case
  • Naming a test clearly
  • Deleting or weakening a failing assertion
Answer

Deleting or weakening a failing assertion — Making the check easier is a classic way to fake success.