Lesson 9 / 29
Test-Set Hygiene: Leakage and Contamination
Keep test cases out of prompts, training data and examples.
If the test leaks, the score lies
A score is only honest if the system has not seen the answers. Leakage is common: test cases copied into few-shot examples in the prompt; the same documents in the fine-tuning data and the evaluation set; test questions in the retrieval index with their answers; public benchmark items that a model saw during pretraining (contamination). Practices: split data into train/dev/test before work starts and tag each case; check for near-duplicates between sets (for example shared word n-grams, as in the example); never use test cases as prompt examples; keep a private held-out set that nobody tunes on; refresh the set with new real cases over time; and be sceptical of public benchmark scores for your own decisions. After fixing a failure from the test set, add a new similar case instead of re-using the old one as proof.
Detecting leaked test cases with n-grams, run
I ran this with plain Python 3 (standard library only). The clean test case shares no 5-word sequences with the training text, while the leaked one (copied from training with one added word) shares 9 and is flagged.
def ngrams(text, n=5):
w = text.lower().split(); return {" ".join(w[i:i + n]) for i in range(len(w) - n + 1)}
train = ["Our refund policy allows returns within thirty days of delivery for unused items",
"Passwords must be at least twelve characters and include a symbol"]
test_ok = "How long do I have to return an item I bought last week"
test_bad = "Our refund policy allows returns within thirty days of delivery for unused items today"
train_grams = set().union(*(ngrams(t) for t in train))
for name, t in (("clean test case", test_ok), ("leaked test case", test_bad)):
overlap = ngrams(t) & train_grams
print(f"{name:17} shared 5-grams with training data: {len(overlap)}", "-> LEAK" if overlap else "")
Output:
clean test case shared 5-grams with training data: 0 leaked test case shared 5-grams with training data: 9 -> LEAK
Add new cases, do not reuse old proof
After fixing a failure, test with a fresh similar case.
Quick check: Why is it a problem if test questions appear in the few-shot examples of the prompt?
- Few-shot examples are forbidden
- It makes the prompt shorter
- The system has seen the answers, so the score overstates real performance
- It has no effect
Answer
The system has seen the answers, so the score overstates real performance — Evaluation requires cases the system has not been shown.