Lesson 24 / 29

Building a Test Set and Harness

Compare prompt versions on fixed examples with a score.

Replace impressions with numbers

Collect 30 to 200 realistic test cases with expected outputs or grading rules, including normal cases, edge cases, adversarial inputs and examples where the right answer is "unknown". Write a small harness that runs a prompt version over all cases and reports a score and the failures. Use exact or rule-based checks where possible (label equals expected, JSON valid, contains required fields) and rubrics or a judge model for open-ended text, spot-checked by humans. Change one thing at a time and re-run the whole set, because a fix for one case often breaks another.

Measure before you ship

A test set, versioning and cost tracking turn prompting into engineering.

Four habits: test, version, cost, monitor.
Figure 7.1 — Test, version, cost and monitor.

Comparing two prompts, run

I ran this plain-Python (standard library only) example. Two simple functions stand in for two prompt versions. Version A gets 3 of 4 right and fails on "Refund my invoice"; version B, which also knows "refund" and "invoice", gets 4 of 4. The harness shows exactly which case changed.

cases = [
    ("I was charged twice", "billing"),
    ("App crashes on start", "technical"),
    ("Where is your office", "other"),
    ("Refund my invoice", "billing"),
]

def prompt_a(text):          # keyword-only "model" standing in for prompt A
    t = text.lower()
    return "billing" if "charged" in t else "technical" if "crash" in t else "other"

def prompt_b(text):          # an improved version that also knows "refund"/"invoice"
    t = text.lower()
    if any(w in t for w in ("charged", "refund", "invoice")): return "billing"
    return "technical" if "crash" in t else "other"

for name, fn in (("A", prompt_a), ("B", prompt_b)):
    ok = [fn(t) == y for t, y in cases]
    print(name, f"{sum(ok)}/{len(cases)}", "failed:", [c[0] for c, o in zip(cases, ok) if not o])

Output:

A 3/4 failed: ['Refund my invoice']
B 4/4 failed: []

Add every production failure

When a real user hits a bad answer, add that input to the test set so it can never silently regress.

Quick check: Why re-run the whole test set after every prompt change?

  • To avoid versioning
  • To make prompts longer
  • A fix for one case can break others
  • It is required by the tokenizer
Answer

A fix for one case can break others — Regression testing catches unintended side effects.