# Reranking and Diversity (MMR) — Retrieval-Augmented Generation (RAG)

Source: https://www.geekswithgeeks.com/en/rag/q-rerank

> Re-order candidates for relevance and remove redundancy.

## Retrieve wide, then rank carefully

Fast first-stage retrieval (BM25, vectors) is good at **recall** but imperfect at ordering. A **reranker**, usually a **cross-encoder** model that reads the question and a chunk together and outputs a relevance score, re-orders the top 20 to 100 candidates much more accurately but is too slow to run on the entire collection. The common pattern is: retrieve 50, rerank, send the best 5. **Maximal Marginal Relevance (MMR)** addresses another problem, redundancy: if five chunks say the same thing you waste prompt space, so MMR picks results that are relevant to the query but different from those already chosen.

## MMR picks diverse results, run

I ran this plain-Python (standard library only) example. Plain top-2 returns A and its near-identical duplicate A_dup. MMR returns A and C instead, giving the model two different pieces of evidence.

```python
import math

def cos(a, b): return sum(x*y for x, y in zip(a, b)) / (math.sqrt(sum(x*x for x in a)) * math.sqrt(sum(y*y for y in b)))

q = [1.0, 0.0]
cands = {"A": [1.0, 0.1], "A_dup": [1.0, 0.1], "B": [0.7, 0.7], "C": [0.5, 0.9]}

def mmr(q, cands, k, lam=0.4):
    chosen = []
    while len(chosen) < k:
        best, best_s = None, -9
        for name, v in cands.items():
            if name in chosen: continue
            rel = cos(q, v)
            red = max((cos(v, cands[c]) for c in chosen), default=0)
            s = lam * rel - (1 - lam) * red
            if s > best_s: best, best_s = name, s
        chosen.append(best)
    return chosen

print("plain top-2:", sorted(cands, key=lambda n: -cos(q, cands[n]))[:2])
print("MMR top-2  :", mmr(q, cands, 2))

```

Output:

```
plain top-2: ['A', 'A_dup']
MMR top-2  : ['A', 'C']
```

## Retrieve more than you send

A common setting is to fetch 30 to 50 candidates, rerank, and pass only the best 3 to 5 to the model.

**Quiz:** What is the usual role of a cross-encoder reranker?

- [ ] Replace the LLM
- [ ] Create the embeddings of the whole collection
- [ ] Clean PDFs
- [x] Re-order a small candidate set more accurately than first-stage search

*Answer:* Re-order a small candidate set more accurately than first-stage search. Rerankers are accurate but costly, so they are applied to a shortlist.
