Lesson 17 / 28

Recall@k, MRR and nDCG on a Golden Set

Quantify retrieval quality with a labelled set of real questions.

You cannot improve what you do not measure

Build a golden set: 50 to 200 real questions, each labelled with the document(s) that answer it, including questions that have no answer. Then compute: Recall@k (the share of relevant documents found in the top k; the key number when a later stage, such as an LLM, can only use what you retrieved), Precision@k, MRR (mean of 1 / rank of the first relevant result; rewards putting the answer first) and nDCG (handles graded relevance and rank positions). Report them per query type and look at the worst queries, since averages hide failures. Run the set whenever you change the model, chunking, index settings, filters or reranker, and keep the results alongside the configuration.

Measure, then improve

A labelled question set, recall and MRR, and error analysis show what to fix.

Three steps: label, measure, diagnose.
Figure 5.1 — Label, measure and diagnose.

Recall and MRR for three questions, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. For "vacation days" the right document is second, so recall@1 is 0 but recall@3 is 1 and MRR is 0.5. "hotel limit" is perfect. "stolen laptop" misses the relevant document completely. Averaged: recall@3 = 0.667, MRR = 0.5. The retrieved lists here are hand-written examples to show the calculation.

def recall_at_k(retrieved, relevant, k): return len(set(retrieved[:k]) & set(relevant)) / len(relevant)
def mrr(retrieved, relevant):
    for i, d in enumerate(retrieved, 1):
        if d in relevant: return 1 / i
    return 0.0

cases = {
    "vacation days":   (["d2", "d1", "d9"], ["d1"]),
    "hotel limit":     (["d6", "d5", "d7"], ["d6"]),
    "stolen laptop":   (["d3", "d4", "d8"], ["d10"]),
}
for name, (got, rel) in cases.items():
    print(f"{name:14} recall@1={recall_at_k(got, rel, 1):.1f} recall@3={recall_at_k(got, rel, 3):.1f} mrr={mrr(got, rel):.2f}")
n = len(cases)
print("mean recall@3:", round(sum(recall_at_k(g, r, 3) for g, r in cases.values()) / n, 3),
      "| mean MRR:", round(sum(mrr(g, r) for g, r in cases.values()) / n, 3))

Output:

vacation days  recall@1=0.0 recall@3=1.0 mrr=0.50
hotel limit    recall@1=1.0 recall@3=1.0 mrr=1.00
stolen laptop  recall@1=0.0 recall@3=0.0 mrr=0.00
mean recall@3: 0.667 | mean MRR: 0.5

Include unanswerable questions

About 10 to 20% of your set should have no answer, so you can measure whether the system correctly returns nothing.

Quick check: What does MRR reward?

  • Using bigger vectors
  • Returning more results
  • Placing the first relevant result near the top
  • Lower latency
Answer

Placing the first relevant result near the top — A first relevant result at rank 1 scores 1; at rank 4 it scores 0.25.