Lesson 21 / 27

Evaluating LLM Output

Build a test set and choose metrics beyond vibes.

A fixed test set beats impressions

Create an evaluation set of realistic inputs with expected answers or grading rules, and re-run it whenever you change the prompt, model or retrieval. Use exact or rule-based checks where possible (valid JSON, correct label, contains required fields). For open-ended text use rubrics and LLM-as-judge scoring, but spot-check the judge against human ratings because judges have biases (for example toward longer answers). Track groundedness (is every claim supported by the sources?), refusal rate, latency and cost alongside quality. Public benchmark scores are a starting point; your own task data decides.

Measure it, then protect it

You cannot improve what you do not measure; and an LLM app has new attack surfaces.

Three checks: test set, metrics, guardrails.
Figure 6.1 — Test set, metrics and guardrails.

Exact versus lenient scoring, run

I ran this plain-Python (standard library only) example. The same three answers score 1/3 under exact match but 2/3 once case is ignored; "warm" is genuinely wrong. Choosing the metric changes the story, so decide it before looking at results.

cases = [
    ("2+2", "4"), ("capital of France", "Paris"), ("opposite of hot", "cold"),
]
fake_model = {"2+2": "4", "capital of France": "paris", "opposite of hot": "warm"}
exact = sum(fake_model[q] == a for q, a in cases)
loose = sum(fake_model[q].lower() == a.lower() for q, a in cases)
print("exact match:", exact, "/", len(cases))
print("case-insensitive:", loose, "/", len(cases))

Output:

exact match: 1 / 3
case-insensitive: 2 / 3

Include hard and adversarial cases

Add typos, other languages, empty inputs, trick questions and cases where the right answer is "I don't know".

Quick check: Why keep a fixed evaluation set?

  • To train the base model
  • To compare changes to prompts or models objectively
  • To reduce token cost
  • To avoid testing
Answer

To compare changes to prompts or models objectively — The same test inputs let you see whether a change really helped.