Lesson 21 / 27
Evaluating LLM Output
Build a test set and choose metrics beyond vibes.
A fixed test set beats impressions
Create an evaluation set of realistic inputs with expected answers or grading rules, and re-run it whenever you change the prompt, model or retrieval. Use exact or rule-based checks where possible (valid JSON, correct label, contains required fields). For open-ended text use rubrics and LLM-as-judge scoring, but spot-check the judge against human ratings because judges have biases (for example toward longer answers). Track groundedness (is every claim supported by the sources?), refusal rate, latency and cost alongside quality. Public benchmark scores are a starting point; your own task data decides.
Measure it, then protect it
You cannot improve what you do not measure; and an LLM app has new attack surfaces.
Exact versus lenient scoring, run
I ran this plain-Python (standard library only) example. The same three answers score 1/3 under exact match but 2/3 once case is ignored; "warm" is genuinely wrong. Choosing the metric changes the story, so decide it before looking at results.
cases = [
("2+2", "4"), ("capital of France", "Paris"), ("opposite of hot", "cold"),
]
fake_model = {"2+2": "4", "capital of France": "paris", "opposite of hot": "warm"}
exact = sum(fake_model[q] == a for q, a in cases)
loose = sum(fake_model[q].lower() == a.lower() for q, a in cases)
print("exact match:", exact, "/", len(cases))
print("case-insensitive:", loose, "/", len(cases))
Output:
exact match: 1 / 3 case-insensitive: 2 / 3
Include hard and adversarial cases
Add typos, other languages, empty inputs, trick questions and cases where the right answer is "I don't know".
Quick check: Why keep a fixed evaluation set?
- To train the base model
- To compare changes to prompts or models objectively
- To reduce token cost
- To avoid testing
Answer
To compare changes to prompts or models objectively — The same test inputs let you see whether a change really helped.