# Evaluating LLM Output — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/e-eval

> Build a test set and choose metrics beyond vibes.

## A fixed test set beats impressions

Create an **evaluation set** of realistic inputs with expected answers or grading rules, and re-run it whenever you change the prompt, model or retrieval. Use **exact or rule-based checks** where possible (valid JSON, correct label, contains required fields). For open-ended text use **rubrics** and **LLM-as-judge** scoring, but spot-check the judge against human ratings because judges have biases (for example toward longer answers). Track **groundedness** (is every claim supported by the sources?), **refusal rate**, **latency** and **cost** alongside quality. Public benchmark scores are a starting point; your own task data decides.

## Measure it, then protect it

You cannot improve what you do not measure; and an LLM app has new attack surfaces.

![Three checks: test set, metrics, guardrails.](assets/figures/llms/section-6-map.svg) — Figure 6.1 — Test set, metrics and guardrails.

## Exact versus lenient scoring, run

I ran this plain-Python (standard library only) example. The same three answers score 1/3 under exact match but 2/3 once case is ignored; "warm" is genuinely wrong. Choosing the metric changes the story, so decide it before looking at results.

```python
cases = [
    ("2+2", "4"), ("capital of France", "Paris"), ("opposite of hot", "cold"),
]
fake_model = {"2+2": "4", "capital of France": "paris", "opposite of hot": "warm"}
exact = sum(fake_model[q] == a for q, a in cases)
loose = sum(fake_model[q].lower() == a.lower() for q, a in cases)
print("exact match:", exact, "/", len(cases))
print("case-insensitive:", loose, "/", len(cases))

```

Output:

```
exact match: 1 / 3
case-insensitive: 2 / 3
```

## Include hard and adversarial cases

Add typos, other languages, empty inputs, trick questions and cases where the right answer is "I don't know".

**Quiz:** Why keep a fixed evaluation set?

- [ ] To train the base model
- [x] To compare changes to prompts or models objectively
- [ ] To reduce token cost
- [ ] To avoid testing

*Answer:* To compare changes to prompts or models objectively. The same test inputs let you see whether a change really helped.
