# Metadata Filtering — Embeddings & Vector Search

Source: https://www.geekswithgeeks.com/en/embeddings/a-filter

> Combine similarity with hard constraints such as tenant, date and permissions.

## Similarity ranks; filters decide eligibility

Similarity alone ignores facts such as **who may see a document**, which **country or product** it covers, and **how recent** it is. Store such fields as **metadata** next to each vector and apply them as **hard filters**: "only documents in the user's tenant, for India, effective 2025". Two designs exist: **pre-filtering** (restrict candidates first, then search) and **post-filtering** (search, then drop non-matching results). Post-filtering is simple but can return fewer than `k` results if most nearest neighbours are filtered out, so over-fetch (retrieve more than `k`) or use an engine with filter-aware search. Permission filters are a **security control**: enforce them in the query itself, never only in the prompt or UI.

## Filter, combine, rerank, group

Real search combines metadata filters, keyword scores and reranking, and embeddings also power clustering and de-duplication.

![Four patterns: filter, hybrid, rerank, group.](assets/figures/embeddings/section-4-map.svg) — Figure 4.1 — Filter, hybrid, rerank and group.

## Filter then rank, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. Without a filter the old 2022 India policy and the UK policy win on raw similarity score. Filtering to India and 2025 first leaves chunks 1 and 4, ranked by score.

```python
chunks = [
    {"id": 1, "text": "Leave policy for India", "country": "IN", "year": 2025, "score": 0.91},
    {"id": 2, "text": "Leave policy for UK",    "country": "UK", "year": 2025, "score": 0.93},
    {"id": 3, "text": "Old leave policy India", "country": "IN", "year": 2022, "score": 0.95},
    {"id": 4, "text": "Travel policy India",    "country": "IN", "year": 2025, "score": 0.70},
]
def top_k(items, k, **where):
    ok = [c for c in items if all(c[f] == v for f, v in where.items())]     # hard filter, then rank by similarity
    return [c["id"] for c in sorted(ok, key=lambda c: -c["score"])[:k]]

print("no filter     :", top_k(chunks, 2))
print("IN + 2025 only:", top_k(chunks, 2, country="IN", year=2025))

```

Output:

```
no filter     : [3, 2]
IN + 2025 only: [1, 4]
```

## Over-fetch before post-filtering

Ask for 5 to 10 times k, filter, then cut to k, or use pre-filtering so you never return too few results.

**Quiz:** Where should "this user may see only their own documents" be enforced?

- [ ] Only in the prompt
- [x] In the retrieval query as a filter
- [ ] Only in the UI
- [ ] Nowhere

*Answer:* In the retrieval query as a filter. Documents the user may not see must never be retrieved.
