Lesson 19 / 27
Retrieval Metrics: Recall@k and MRR
Score whether the right chunks are found and how high they rank.
Build a golden set first
Create a golden set of 50 to 200 real questions, each with the chunk(s) or document(s) that contain the answer. Then compute Recall@k (the fraction of relevant chunks found in the top k), Precision@k, MRR (mean reciprocal rank: the average of 1 divided by the rank of the first relevant result) and nDCG for graded relevance. Recall@k is the key number for RAG, because the generator cannot use evidence it never received. Re-run the set after every change to chunking, embeddings, filters or rerankers.
Measure each stage
Retrieval and generation fail differently, so score them separately.
Recall@k and MRR, run
I ran this plain-Python (standard library only) example. Across three questions, the right chunk is first for none (recall@1 = 0), found within the top 3 for half of the relevant chunks (recall@3 = 0.5), and MRR is 0.278.
def recall_at_k(retrieved, relevant, k):
return len(set(retrieved[:k]) & set(relevant)) / len(relevant)
def mrr(retrieved, relevant):
for i, d in enumerate(retrieved, start=1):
if d in relevant:
return 1 / i
return 0.0
cases = [
(["d3", "d1", "d9"], ["d1"]),
(["d2", "d4", "d5"], ["d5", "d7"]),
(["d8", "d6", "d0"], ["d1"]),
]
for k in (1, 3):
r = sum(recall_at_k(ret, rel, k) for ret, rel in cases) / len(cases)
print("recall@%d" % k, round(r, 3))
print("MRR", round(sum(mrr(ret, rel) for ret, rel in cases) / len(cases), 3))
Output:
recall@1 0.0 recall@3 0.5 MRR 0.278
Quick check: Why is Recall@k so important for RAG?
- It is unrelated to answers
- It measures the font quality
- It replaces the LLM
- The generator cannot use evidence that was never retrieved
Answer
The generator cannot use evidence that was never retrieved — Missing evidence caps the quality of any answer.