Lesson 14 / 28

Hybrid Search and Reranking

Combine keyword and vector results and re-order the shortlist.

Two retrievers, then a judge

Hybrid search runs a keyword retriever (BM25) and a vector retriever, then merges the two lists. Because their scores are on different scales, a robust merge is Reciprocal Rank Fusion: each document gets 1 / (k + rank) from each list and the sums decide the order. After retrieval, a reranker (a cross-encoder that reads the query and a candidate together) re-orders the top 20 to 100 candidates much more accurately than vector similarity alone, at a higher cost per pair, so it is applied only to a shortlist. The usual pipeline is: filter → hybrid retrieve ~50 → rerank → keep 3 to 5. Add a score threshold so weak matches are dropped.

The usual pipeline

Each stage trades a little cost for quality.

query
  -> hard filters (tenant, country, date, permissions)
  -> keyword retrieve (BM25, top 50)  +  vector retrieve (ANN, top 50)
  -> merge with Reciprocal Rank Fusion  (score = sum of 1/(60 + rank) over the lists)
  -> cross-encoder rerank (top 50 -> best 5)
  -> drop anything below a minimum score  -> answer / show results

Tune the merge on your data

RRF with k=60 is a good default; if one retriever is clearly stronger on your golden set, weight it higher and re-measure.

Quick check: Why is a cross-encoder applied only to a shortlist?

  • It replaces the filter
  • It cannot read text
  • It only works on images
  • It is accurate but too costly to run on the whole collection
Answer

It is accurate but too costly to run on the whole collection — Fast retrieval narrows candidates; the expensive model refines them.