Lesson 11 / 27
Hybrid Search and Reciprocal Rank Fusion
Combine keyword and vector results into one ranking.
Two weak signals beat one
Keyword search is precise on exact terms; vector search understands meaning. Hybrid search runs both and merges the lists. The scores of the two systems are on different scales, so a robust way to merge is Reciprocal Rank Fusion (RRF): each document gets 1 / (k + rank) from every list it appears in (k is often 60), and the sums decide the final order. Documents that rank well in both lists rise to the top, and no score calibration is needed. Many vector databases and search engines offer hybrid search built in.
RRF, run
I ran this plain-Python (standard library only) example. Document A is first in the keyword list and second in the vector list, C the reverse. A and C come out nearly tied at the top, E (only in the vector list) and D (only in the keyword list) trail.
def rrf(rankings, k=60):
scores = {}
for ranking in rankings:
for rank, doc in enumerate(ranking, start=1):
scores[doc] = scores.get(doc, 0) + 1 / (k + rank)
return sorted(scores.items(), key=lambda x: -x[1])
keyword = ["A", "B", "C", "D"]
vector = ["C", "A", "E", "B"]
for doc, s in rrf([keyword, vector]):
print(doc, round(s, 4))
Output:
A 0.0325 C 0.0323 B 0.0318 E 0.0159 D 0.0156
Weight if you must
RRF treats both lists equally. If one retriever is clearly stronger on your golden set, weight its contribution higher and re-measure.
Quick check: Why is RRF popular for merging results?
- It deletes duplicates only
- It needs no ranking at all
- It trains a new model
- It uses ranks, so incompatible score scales do not matter
Answer
It uses ranks, so incompatible score scales do not matter — Ranks are comparable across systems even when scores are not.