Lesson 19 / 27

Retrieval Metrics: Recall@k and MRR

Score whether the right chunks are found and how high they rank.

Build a golden set first

Create a golden set of 50 to 200 real questions, each with the chunk(s) or document(s) that contain the answer. Then compute Recall@k (the fraction of relevant chunks found in the top k), Precision@k, MRR (mean reciprocal rank: the average of 1 divided by the rank of the first relevant result) and nDCG for graded relevance. Recall@k is the key number for RAG, because the generator cannot use evidence it never received. Re-run the set after every change to chunking, embeddings, filters or rerankers.

Measure each stage

Retrieval and generation fail differently, so score them separately.

Three views: retrieval, answers, failures.
Figure 6.1 — Retrieval, answers and failures.

Recall@k and MRR, run

I ran this plain-Python (standard library only) example. Across three questions, the right chunk is first for none (recall@1 = 0), found within the top 3 for half of the relevant chunks (recall@3 = 0.5), and MRR is 0.278.

def recall_at_k(retrieved, relevant, k):
    return len(set(retrieved[:k]) & set(relevant)) / len(relevant)

def mrr(retrieved, relevant):
    for i, d in enumerate(retrieved, start=1):
        if d in relevant:
            return 1 / i
    return 0.0

cases = [
    (["d3", "d1", "d9"], ["d1"]),
    (["d2", "d4", "d5"], ["d5", "d7"]),
    (["d8", "d6", "d0"], ["d1"]),
]
for k in (1, 3):
    r = sum(recall_at_k(ret, rel, k) for ret, rel in cases) / len(cases)
    print("recall@%d" % k, round(r, 3))
print("MRR", round(sum(mrr(ret, rel) for ret, rel in cases) / len(cases), 3))

Output:

recall@1 0.0
recall@3 0.5
MRR 0.278

Quick check: Why is Recall@k so important for RAG?

  • It is unrelated to answers
  • It measures the font quality
  • It replaces the LLM
  • The generator cannot use evidence that was never retrieved
Answer

The generator cannot use evidence that was never retrieved — Missing evidence caps the quality of any answer.