Lesson 14 / 28
Hybrid Search and Reranking
Combine keyword and vector results and re-order the shortlist.
Two retrievers, then a judge
Hybrid search runs a keyword retriever (BM25) and a vector retriever, then merges the two lists. Because their scores are on different scales, a robust merge is Reciprocal Rank Fusion: each document gets 1 / (k + rank) from each list and the sums decide the order. After retrieval, a reranker (a cross-encoder that reads the query and a candidate together) re-orders the top 20 to 100 candidates much more accurately than vector similarity alone, at a higher cost per pair, so it is applied only to a shortlist. The usual pipeline is: filter → hybrid retrieve ~50 → rerank → keep 3 to 5. Add a score threshold so weak matches are dropped.
The usual pipeline
Each stage trades a little cost for quality.
query
-> hard filters (tenant, country, date, permissions)
-> keyword retrieve (BM25, top 50) + vector retrieve (ANN, top 50)
-> merge with Reciprocal Rank Fusion (score = sum of 1/(60 + rank) over the lists)
-> cross-encoder rerank (top 50 -> best 5)
-> drop anything below a minimum score -> answer / show resultsTune the merge on your data
RRF with k=60 is a good default; if one retriever is clearly stronger on your golden set, weight it higher and re-measure.
Quick check: Why is a cross-encoder applied only to a shortlist?
- It replaces the filter
- It cannot read text
- It only works on images
- It is accurate but too costly to run on the whole collection
Answer
It is accurate but too costly to run on the whole collection — Fast retrieval narrows candidates; the expensive model refines them.