Lesson 7 / 29
Running Commands and Tests, and Reading the Results
Turn noisy command output into a concise observation.
The feedback that makes agents work
Running code is what separates an agent from a text generator. The harness provides a shell/command tool (often with a timeout, output limits, and a working directory fixed to the project) and usually convenient wrappers for tests, linters, type checkers and builds. Raw output can be huge: a good tool returns a concise observation: exit code, how many tests ran, which failed, and the first assertion message and stack frame, with the full log available on request. The agent uses this to decide the next step. Commands can also cause harm (deleting files, installing packages, network calls), so which commands run automatically, which need approval and which are denied must be set by policy (see the guardrails course), not trusted to the model.
A test-runner tool that returns a concise observation, run
I ran this with plain Python 3 (standard library only), using a throwaway project created in a temporary folder. Running the sample project's three tests, the tool reports exit code 1, "3 tests", exactly which test failed (test_discount_ten_percent) and the assertion message 0.0 != 180. That is enough for a model to find the bug without reading pages of log.
import os, tempfile, textwrap
def make_project(root):
files = {
"shop/__init__.py": "",
"shop/pricing.py": textwrap.dedent("""
def apply_discount(price, percent):
\"\"\"Return price after a percentage discount.\"\"\"
return price - price * percent / 10
def add_tax(price, rate=0.18):
return round(price * (1 + rate), 2)
"""),
"shop/cart.py": textwrap.dedent("""
from shop.pricing import apply_discount, add_tax
class Cart:
def __init__(self):
self.items = []
def add(self, name, price):
self.items.append((name, price))
def total(self, discount_percent=0):
subtotal = sum(p for _, p in self.items)
return add_tax(apply_discount(subtotal, discount_percent))
"""),
"tests/__init__.py": "",
"tests/test_pricing.py": textwrap.dedent("""
import unittest
from shop.pricing import apply_discount, add_tax
class PricingTests(unittest.TestCase):
def test_discount_ten_percent(self):
self.assertEqual(apply_discount(200, 10), 180)
def test_discount_zero(self):
self.assertEqual(apply_discount(200, 0), 200)
def test_tax(self):
self.assertEqual(add_tax(100), 118.0)
"""),
}
for rel, text in files.items():
path = os.path.join(root, rel)
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, "w") as f:
f.write(text.lstrip("\n"))
import subprocess, sys, re
def run_tests(root):
p = subprocess.run([sys.executable, "-m", "unittest", "discover", "-s", "tests", "-t", "."],
cwd=root, capture_output=True, text=True)
out = p.stderr
summary = re.search(r"Ran (\d+) tests? in", out).group(1) + " tests"
failed = re.findall(r"^(?:FAIL|ERROR): (\S+)", out, re.M)
first = re.search(r"AssertionError: .*", out)
return {"exit_code": p.returncode, "summary": summary, "failed": failed, "first_error": first.group(0) if first else None}
with tempfile.TemporaryDirectory() as root:
make_project(root)
print(run_tests(root))
Output:
{'exit_code': 1, 'summary': '3 tests', 'failed': ['test_discount_ten_percent'], 'first_error': 'AssertionError: 0.0 != 180'}Always set a timeout
A hung test or an interactive prompt can freeze an agent forever. Give every command a time limit and a non-interactive mode.
Quick check: What should a test tool return to the model?
- A concise summary: exit code, failed tests and the first error
- Nothing at all
- Only the word "done"
- The entire log every time
Answer
A concise summary: exit code, failed tests and the first error — Concise, relevant observations save context and speed up fixes.