# Retrieval Metrics: Recall@k and MRR — Retrieval-Augmented Generation (RAG)

Source: https://www.geekswithgeeks.com/en/rag/e-retrieval

> Score whether the right chunks are found and how high they rank.

## Build a golden set first

Create a **golden set** of 50 to 200 real questions, each with the chunk(s) or document(s) that contain the answer. Then compute **Recall@k** (the fraction of relevant chunks found in the top k), **Precision@k**, **MRR** (mean reciprocal rank: the average of 1 divided by the rank of the first relevant result) and **nDCG** for graded relevance. Recall@k is the key number for RAG, because the generator cannot use evidence it never received. Re-run the set after every change to chunking, embeddings, filters or rerankers.

## Measure each stage

Retrieval and generation fail differently, so score them separately.

![Three views: retrieval, answers, failures.](assets/figures/rag/section-6-map.svg) — Figure 6.1 — Retrieval, answers and failures.

## Recall@k and MRR, run

I ran this plain-Python (standard library only) example. Across three questions, the right chunk is first for none (recall@1 = 0), found within the top 3 for half of the relevant chunks (recall@3 = 0.5), and MRR is 0.278.

```python
def recall_at_k(retrieved, relevant, k):
    return len(set(retrieved[:k]) & set(relevant)) / len(relevant)

def mrr(retrieved, relevant):
    for i, d in enumerate(retrieved, start=1):
        if d in relevant:
            return 1 / i
    return 0.0

cases = [
    (["d3", "d1", "d9"], ["d1"]),
    (["d2", "d4", "d5"], ["d5", "d7"]),
    (["d8", "d6", "d0"], ["d1"]),
]
for k in (1, 3):
    r = sum(recall_at_k(ret, rel, k) for ret, rel in cases) / len(cases)
    print("recall@%d" % k, round(r, 3))
print("MRR", round(sum(mrr(ret, rel) for ret, rel in cases) / len(cases), 3))

```

Output:

```
recall@1 0.0
recall@3 0.5
MRR 0.278
```

**Quiz:** Why is Recall@k so important for RAG?

- [ ] It is unrelated to answers
- [ ] It measures the font quality
- [ ] It replaces the LLM
- [x] The generator cannot use evidence that was never retrieved

*Answer:* The generator cannot use evidence that was never retrieved. Missing evidence caps the quality of any answer.
