Lesson 3 / 29
The Agent Loop in Action
Run a miniature coding agent that fixes a bug and watch each step.
Observe, decide, act, check
The loop is simple. (1) Observe: run the tests and see what fails. (2) Decide: read the relevant code and form a hypothesis. (3) Act: make a small edit. (4) Check: run the tests again. (5) Stop when the checks pass, or when a limit is reached. The example below is a miniature, real agent loop: a throwaway project has a bug in apply_discount (it divides the percentage by 10 instead of 100); the harness provides run_tests, read_file and edit_file tools; and a scripted stand-in plays the model, proposing the five actions a real model would plausibly choose. Nothing here calls a real model, so it shows the mechanics (tools, observations, checks, a hard step cap), not the intelligence. A real agent chooses its own actions, which is exactly why the verification step matters.
A miniature coding agent, run
I ran this with plain Python 3 (standard library only), using a throwaway project created in a temporary folder. Step 1 runs the tests and finds test_discount_ten_percent failing. Steps 2 and 3 read the file and apply an exact-match edit changing / 10 to / 100. Step 4 re-runs the tests, which now pass, and step 5 finishes with a summary. A final independent test run confirms the fix. The model here is a fixed script, so this demonstrates the harness, not a real model.
import os, tempfile, textwrap
def make_project(root):
files = {
"shop/__init__.py": "",
"shop/pricing.py": textwrap.dedent("""
def apply_discount(price, percent):
\"\"\"Return price after a percentage discount.\"\"\"
return price - price * percent / 10
def add_tax(price, rate=0.18):
return round(price * (1 + rate), 2)
"""),
"shop/cart.py": textwrap.dedent("""
from shop.pricing import apply_discount, add_tax
class Cart:
def __init__(self):
self.items = []
def add(self, name, price):
self.items.append((name, price))
def total(self, discount_percent=0):
subtotal = sum(p for _, p in self.items)
return add_tax(apply_discount(subtotal, discount_percent))
"""),
"tests/__init__.py": "",
"tests/test_pricing.py": textwrap.dedent("""
import unittest
from shop.pricing import apply_discount, add_tax
class PricingTests(unittest.TestCase):
def test_discount_ten_percent(self):
self.assertEqual(apply_discount(200, 10), 180)
def test_discount_zero(self):
self.assertEqual(apply_discount(200, 0), 200)
def test_tax(self):
self.assertEqual(add_tax(100), 118.0)
"""),
}
for rel, text in files.items():
path = os.path.join(root, rel)
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "w") as f:
f.write(text.lstrip("\n"))
import subprocess, sys, re
def run_tests(root):
p = subprocess.run([sys.executable, "-m", "unittest", "discover", "-s", "tests", "-t", "."],
cwd=root, capture_output=True, text=True)
failed = re.findall(r"^(?:FAIL|ERROR): (\S+)", p.stderr, re.M)
return {"passed": p.returncode == 0, "failed": failed}
def read_file(root, rel): return open(os.path.join(root, rel)).read()
def edit_file(root, rel, old, new):
path = os.path.join(root, rel); text = open(path).read()
if text.count(old) != 1: return "error: old text must appear exactly once"
open(path, "w").write(text.replace(old, new)); return "ok"
# A scripted stand-in for the model: it proposes the next action given what it has observed so far.
SCRIPT = [
("run_tests", {}),
("read_file", {"rel": "shop/pricing.py"}),
("edit_file", {"rel": "shop/pricing.py", "old": "price * percent / 10", "new": "price * percent / 100"}),
("run_tests", {}),
("finish", {"message": "Fixed apply_discount: percent was divided by 10 instead of 100."}),
]
def agent(root, max_steps=8):
script = iter(SCRIPT)
for step in range(1, max_steps + 1): # a hard cap on steps
name, args = next(script)
if name == "run_tests": obs = run_tests(root)
elif name == "read_file": obs = "(%d chars)" % len(read_file(root, **args))
elif name == "edit_file": obs = edit_file(root, **args)
else:
print(f"step {step}: finish -> {args['message']}"); return
print(f"step {step}: {name}({', '.join(f'{k}=...' for k in args)}) -> {obs}")
with tempfile.TemporaryDirectory() as root:
make_project(root)
agent(root)
print("final test run:", run_tests(root))
Output:
step 1: run_tests() -> {'passed': False, 'failed': ['test_discount_ten_percent']}
step 2: read_file(rel=...) -> (200 chars)
step 3: edit_file(rel=..., old=..., new=...) -> ok
step 4: run_tests() -> {'passed': True, 'failed': []}
step 5: finish -> Fixed apply_discount: percent was divided by 10 instead of 100.
final test run: {'passed': True, 'failed': []}Start every run from green
Make sure tests pass before the agent starts; otherwise you cannot tell which failures it caused.
Quick check: Why does the loop end with an independent test run?
- Because the model cannot read files
- To make the run longer
- Tests are optional decoration
- The agent's own claim of success is not proof; tests are
Answer
The agent's own claim of success is not proof; tests are — Verification should not depend on the agent's self-report.