# Building a Test Set and Harness — Prompt Engineering

Source: https://www.geekswithgeeks.com/en/prompt-engineering/v-testset

> Compare prompt versions on fixed examples with a score.

## Replace impressions with numbers

Collect 30 to 200 **realistic test cases** with expected outputs or grading rules, including normal cases, edge cases, adversarial inputs and examples where the right answer is "unknown". Write a small **harness** that runs a prompt version over all cases and reports a score and the failures. Use **exact or rule-based checks** where possible (label equals expected, JSON valid, contains required fields) and **rubrics or a judge model** for open-ended text, spot-checked by humans. Change one thing at a time and re-run the whole set, because a fix for one case often breaks another.

## Measure before you ship

A test set, versioning and cost tracking turn prompting into engineering.

![Four habits: test, version, cost, monitor.](assets/figures/prompt-engineering/section-7-map.svg) — Figure 7.1 — Test, version, cost and monitor.

## Comparing two prompts, run

I ran this plain-Python (standard library only) example. Two simple functions stand in for two prompt versions. Version A gets 3 of 4 right and fails on "Refund my invoice"; version B, which also knows "refund" and "invoice", gets 4 of 4. The harness shows exactly which case changed.

```python
cases = [
    ("I was charged twice", "billing"),
    ("App crashes on start", "technical"),
    ("Where is your office", "other"),
    ("Refund my invoice", "billing"),
]

def prompt_a(text):          # keyword-only "model" standing in for prompt A
    t = text.lower()
    return "billing" if "charged" in t else "technical" if "crash" in t else "other"

def prompt_b(text):          # an improved version that also knows "refund"/"invoice"
    t = text.lower()
    if any(w in t for w in ("charged", "refund", "invoice")): return "billing"
    return "technical" if "crash" in t else "other"

for name, fn in (("A", prompt_a), ("B", prompt_b)):
    ok = [fn(t) == y for t, y in cases]
    print(name, f"{sum(ok)}/{len(cases)}", "failed:", [c[0] for c, o in zip(cases, ok) if not o])

```

Output:

```
A 3/4 failed: ['Refund my invoice']
B 4/4 failed: []
```

## Add every production failure

When a real user hits a bad answer, add that input to the test set so it can never silently regress.

**Quiz:** Why re-run the whole test set after every prompt change?

- [ ] To avoid versioning
- [ ] To make prompts longer
- [x] A fix for one case can break others
- [ ] It is required by the tokenizer

*Answer:* A fix for one case can break others. Regression testing catches unintended side effects.
