# Test-Set Hygiene: Leakage and Contamination — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/e-leak

> Keep test cases out of prompts, training data and examples.

## If the test leaks, the score lies

A score is only honest if the system has not **seen the answers**. Leakage is common: test cases copied into few-shot examples in the prompt; the same documents in the fine-tuning data and the evaluation set; test questions in the retrieval index with their answers; public benchmark items that a model saw during pretraining (**contamination**). Practices: split data into **train/dev/test** before work starts and tag each case; check for **near-duplicates** between sets (for example shared word n-grams, as in the example); never use test cases as prompt examples; keep a **private held-out set** that nobody tunes on; **refresh** the set with new real cases over time; and be sceptical of public benchmark scores for your own decisions. After fixing a failure from the test set, add a **new** similar case instead of re-using the old one as proof.

## Detecting leaked test cases with n-grams, run

I ran this with plain Python 3 (standard library only). The clean test case shares no 5-word sequences with the training text, while the leaked one (copied from training with one added word) shares 9 and is flagged.

```python
def ngrams(text, n=5):
    w = text.lower().split(); return {" ".join(w[i:i + n]) for i in range(len(w) - n + 1)}

train = ["Our refund policy allows returns within thirty days of delivery for unused items",
         "Passwords must be at least twelve characters and include a symbol"]
test_ok  = "How long do I have to return an item I bought last week"
test_bad = "Our refund policy allows returns within thirty days of delivery for unused items today"

train_grams = set().union(*(ngrams(t) for t in train))
for name, t in (("clean test case", test_ok), ("leaked test case", test_bad)):
    overlap = ngrams(t) & train_grams
    print(f"{name:17} shared 5-grams with training data: {len(overlap)}", "-> LEAK" if overlap else "")

```

Output:

```
clean test case   shared 5-grams with training data: 0 
leaked test case  shared 5-grams with training data: 9 -> LEAK
```

## Add new cases, do not reuse old proof

After fixing a failure, test with a fresh similar case.

**Quiz:** Why is it a problem if test questions appear in the few-shot examples of the prompt?

- [ ] Few-shot examples are forbidden
- [ ] It makes the prompt shorter
- [x] The system has seen the answers, so the score overstates real performance
- [ ] It has no effect

*Answer:* The system has seen the answers, so the score overstates real performance. Evaluation requires cases the system has not been shown.
