# Recall@k, MRR and nDCG on a Golden Set — Embeddings & Vector Search

Source: https://www.geekswithgeeks.com/en/embeddings/e-metrics

> Quantify retrieval quality with a labelled set of real questions.

## You cannot improve what you do not measure

Build a **golden set**: 50 to 200 real questions, each labelled with the document(s) that answer it, including questions that have **no answer**. Then compute: **Recall@k** (the share of relevant documents found in the top `k`; the key number when a later stage, such as an LLM, can only use what you retrieved), **Precision@k**, **MRR** (mean of 1 / rank of the first relevant result; rewards putting the answer first) and **nDCG** (handles graded relevance and rank positions). Report them per query type and look at the **worst queries**, since averages hide failures. Run the set whenever you change the model, chunking, index settings, filters or reranker, and keep the results alongside the configuration.

## Measure, then improve

A labelled question set, recall and MRR, and error analysis show what to fix.

![Three steps: label, measure, diagnose.](assets/figures/embeddings/section-5-map.svg) — Figure 5.1 — Label, measure and diagnose.

## Recall and MRR for three questions, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. For "vacation days" the right document is second, so recall@1 is 0 but recall@3 is 1 and MRR is 0.5. "hotel limit" is perfect. "stolen laptop" misses the relevant document completely. Averaged: recall@3 = 0.667, MRR = 0.5. The retrieved lists here are hand-written examples to show the calculation.

```python
def recall_at_k(retrieved, relevant, k): return len(set(retrieved[:k]) & set(relevant)) / len(relevant)
def mrr(retrieved, relevant):
    for i, d in enumerate(retrieved, 1):
        if d in relevant: return 1 / i
    return 0.0

cases = {
    "vacation days":   (["d2", "d1", "d9"], ["d1"]),
    "hotel limit":     (["d6", "d5", "d7"], ["d6"]),
    "stolen laptop":   (["d3", "d4", "d8"], ["d10"]),
}
for name, (got, rel) in cases.items():
    print(f"{name:14} recall@1={recall_at_k(got, rel, 1):.1f} recall@3={recall_at_k(got, rel, 3):.1f} mrr={mrr(got, rel):.2f}")
n = len(cases)
print("mean recall@3:", round(sum(recall_at_k(g, r, 3) for g, r in cases.values()) / n, 3),
      "| mean MRR:", round(sum(mrr(g, r) for g, r in cases.values()) / n, 3))

```

Output:

```
vacation days  recall@1=0.0 recall@3=1.0 mrr=0.50
hotel limit    recall@1=1.0 recall@3=1.0 mrr=1.00
stolen laptop  recall@1=0.0 recall@3=0.0 mrr=0.00
mean recall@3: 0.667 | mean MRR: 0.5
```

## Include unanswerable questions

About 10 to 20% of your set should have no answer, so you can measure whether the system correctly returns nothing.

**Quiz:** What does MRR reward?

- [ ] Using bigger vectors
- [ ] Returning more results
- [x] Placing the first relevant result near the top
- [ ] Lower latency

*Answer:* Placing the first relevant result near the top. A first relevant result at rank 1 scores 1; at rank 4 it scores 0.25.
