Lesson 13 / 25
Generating Tests That Mean Something
Check that generated tests would actually fail if the code were wrong.
Vacuous tests pass for any code
Assistants can generate many tests fast, but some assert almost nothing: they only check that a function returns something, or they copy the implementation's own logic. A quick check is to break the code on purpose and see whether the tests fail. A test that stays green on broken code adds false confidence. Also look for missing cases: empty input, boundaries, invalid input and error paths.
Prove it, then explain it
Assistants write tests and docs quickly, but you must make sure they check something real.
A weak test vs a strong test
I ran this. Against the correct function both pass. After is_even is broken to always return True, the weak test still passes and only the strong test fails.
def is_even(n): return n % 2 == 0
def weak_test(): assert is_even(2) in (True, False) # always passes
def strong_test(): assert is_even(2) is True and is_even(3) is False
def broken_is_even(n): return True
is_even = broken_is_even
for name, fn in (("weak", weak_test), ("strong", strong_test)):
try: fn(); print(name, "passes on BROKEN code")
except AssertionError: print(name, "fails on broken code")
Output:
weak passes on BROKEN code strong fails on broken code
Mutation by hand
Change a > to >=, flip a condition or return a constant, then run the tests. If nothing fails, the tests are not protecting that line. Tools called mutation testers automate this idea.
Quick check: How can you check that a generated test is meaningful?
- Count how many lines it has
- Break the code on purpose and see if it fails
- Check that it has a long name
- Delete it
Answer
Break the code on purpose and see if it fails — A meaningful test fails when the behaviour it protects is broken.