# An Evaluation Harness — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/e-harness

> Run any version of the system over the same cases and report results by category.

## The same cases, every time

An **evaluation harness** is a small program that takes a set of **test cases** (input, plus the expected answer or a checking rule), runs a candidate version of your system on each, **scores** the outputs, and prints a report by category with the failing cases listed. Make it fast, deterministic where possible (fix random seeds, low temperature), and runnable by anyone with one command, locally and in CI. Store cases as data files, tag them with **categories** (languages, topics, difficulty, adversarial), include **edge cases and "should refuse" cases**, and keep a **held-out set** you do not tune against. Scoring methods, from cheap to costly: **exact match or regex**, **programmatic checks** (valid JSON, correct fields, numbers within range), **reference-based similarity**, **LLM-as-judge**, and **human review**. Prefer programmatic checks wherever they capture what you care about.

## Cases, scores, noise, comparison

Build a test set, score it automatically, understand statistical noise, and compare versions fairly.

![Five parts: set, scorer, interval, paired test, judge.](assets/figures/llm-engineering/section-2-map.svg) — Figure 2.1 — Set, scorer, interval, paired test and judge.

## Comparing two prompt versions, run

I ran this with plain Python 3 (standard library only). This is a simulation with invented numbers and seeded random draws, so it shows the shape of the effect, not measurements of any product. Two stand-in classifiers play "prompt v1" and "prompt v2" on six labelled cases. The report shows 4/6 for v1 and 6/6 for v2, the pass counts per category and exactly which cases v1 failed (ids 2 and 4). A real harness would call the model; the structure is the same.

```python
import re

CASES = [
    {"id": 1, "cat": "billing",   "input": "I was charged twice", "check": ("equals", "billing")},
    {"id": 2, "cat": "billing",   "input": "refund my invoice",   "check": ("equals", "billing")},
    {"id": 3, "cat": "technical", "input": "app crashes on start", "check": ("equals", "technical")},
    {"id": 4, "cat": "technical", "input": "cannot log in",        "check": ("equals", "technical")},
    {"id": 5, "cat": "other",     "input": "where is your office", "check": ("equals", "other")},
    {"id": 6, "cat": "other",     "input": "tell me a joke",       "check": ("equals", "other")},
]
def score(output, check):
    kind, want = check
    return output.strip().lower() == want if kind == "equals" else bool(re.search(want, output))

def prompt_v1(text):          # stand-in for "model + prompt version 1": keyword rules only
    t = text.lower()
    return "billing" if "charged" in t else "technical" if "crash" in t else "other"
def prompt_v2(text):          # version 2 adds a few more cues
    t = text.lower()
    if any(w in t for w in ("charged", "refund", "invoice")): return "billing"
    if any(w in t for w in ("crash", "log in", "error")): return "technical"
    return "other"

def run(model):
    by_cat, results = {}, []
    for c in CASES:
        ok = score(model(c["input"]), c["check"]); results.append((c["id"], ok))
        by_cat.setdefault(c["cat"], []).append(ok)
    return results, {k: f"{sum(v)}/{len(v)}" for k, v in by_cat.items()}

for name, model in (("v1", prompt_v1), ("v2", prompt_v2)):
    results, cats = run(model)
    print(name, "pass", sum(ok for _, ok in results), "/", len(results), "| by category:", cats, "| failed ids:", [i for i, ok in results if not ok])

```

Output:

```
v1 pass 4 / 6 | by category: {'billing': '1/2', 'technical': '1/2', 'other': '2/2'} | failed ids: [2, 4]
v2 pass 6 / 6 | by category: {'billing': '2/2', 'technical': '2/2', 'other': '2/2'} | failed ids: []
```

## Put cases in data files

Plain JSONL files can be reviewed, versioned and extended by anyone on the team.

**Quiz:** Why keep a held-out test set that you do not tune against?

- [ ] Held-out sets run faster
- [x] Tuning on the same cases overstates quality (overfitting to the test)
- [ ] They are required by APIs
- [ ] They remove the need for scoring

*Answer:* Tuning on the same cases overstates quality (overfitting to the test). You need an honest measurement on cases the system was not tuned to.
