# Hybrid Search: Full-Text Plus Vectors — Vector Databases

Source: https://www.geekswithgeeks.com/en/vector-databases/f-hybrid

> Combine keyword relevance with vector similarity in one query.

## Two signals in one ORDER BY

Vector similarity understands meaning; keyword search nails exact terms (names, codes, rare words). **Hybrid search** uses both. In PostgreSQL you can add a `tsvector` column (here a generated column) with a full-text index, rank with `ts_rank`, and combine with vector similarity in the same `ORDER BY` using a weighted sum, or run both queries and merge the lists with **Reciprocal Rank Fusion** in the application or in SQL. Dedicated engines usually offer hybrid queries with sparse vectors (BM25-style) or built-in fusion. Choose weights by **measuring** recall on your labelled questions; scores from the two systems are on different scales, which is why rank-based fusion is popular.

## Vector similarity plus a keyword boost, run

I ran this SQL on PostgreSQL 16 with the pgvector extension, version 0.8.6, in a Docker container. For a query vector between "leave" and "travel", the vector scores of documents 1 and 3 are close (0.781 and 0.776). Adding the full-text rank for the word "hotels", weighted by 5, lifts the travel-and-hotels document to first place.

```sql
ALTER TABLE docs ADD COLUMN IF NOT EXISTS tsv tsvector GENERATED ALWAYS AS (to_tsvector('english', body)) STORED;
SELECT id, body,
       round((1 - (embedding <=> '[0.5,0.5,0]'))::numeric, 3) AS vec_sim,
       round(ts_rank(tsv, plainto_tsquery('english', 'hotels'))::numeric, 3) AS keyword_rank
FROM docs
WHERE tenant = 'acme'
ORDER BY (1 - (embedding <=> '[0.5,0.5,0]')) + 5 * ts_rank(tsv, plainto_tsquery('english', 'hotels')) DESC
LIMIT 3;
```

Output:

```
 id |       body        | vec_sim | keyword_rank 
----+-------------------+---------+--------------
  3 | travel and hotels |   0.776 |        0.061
  2 | old leave policy  |   0.851 |        0.000
  1 | leave policy 2025 |   0.781 |        0.000
(3 rows)
```

## Reciprocal Rank Fusion, run

I ran this plain-Python (standard library only) example. Document `a` ranks first in the vector list and third in the keyword list; `c` is third and first. RRF puts `a` first and `c` second, then `b`, `e`, `d`. It uses only ranks, so the incompatible score scales do not matter.

```python
def rrf(*rankings, k=60):
    scores = {}
    for ranking in rankings:
        for rank, doc in enumerate(ranking, start=1):
            scores[doc] = scores.get(doc, 0) + 1 / (k + rank)
    return [d for d, _ in sorted(scores.items(), key=lambda x: -x[1])]

vector_hits = ["a", "b", "c", "d"]
keyword_hits = ["c", "e", "a"]
print("RRF:", rrf(vector_hits, keyword_hits))

```

Output:

```
RRF: ['a', 'c', 'b', 'e', 'd']
```

## Normalise or fuse by rank

Adding raw keyword and vector scores directly is fragile. Normalise them, or use rank-based fusion such as RRF.

**Quiz:** Why is rank-based fusion popular for hybrid search?

- [x] Keyword and vector scores are on different scales; ranks are comparable
- [ ] It needs no ranking
- [ ] It removes the need for vectors
- [ ] It makes queries slower on purpose

*Answer:* Keyword and vector scores are on different scales; ranks are comparable. RRF avoids calibrating incompatible scores.
