Lesson 13 / 28

Metadata Filtering

Combine similarity with hard constraints such as tenant, date and permissions.

Similarity ranks; filters decide eligibility

Similarity alone ignores facts such as who may see a document, which country or product it covers, and how recent it is. Store such fields as metadata next to each vector and apply them as hard filters: "only documents in the user's tenant, for India, effective 2025". Two designs exist: pre-filtering (restrict candidates first, then search) and post-filtering (search, then drop non-matching results). Post-filtering is simple but can return fewer than k results if most nearest neighbours are filtered out, so over-fetch (retrieve more than k) or use an engine with filter-aware search. Permission filters are a security control: enforce them in the query itself, never only in the prompt or UI.

Filter, combine, rerank, group

Real search combines metadata filters, keyword scores and reranking, and embeddings also power clustering and de-duplication.

Four patterns: filter, hybrid, rerank, group.
Figure 4.1 — Filter, hybrid, rerank and group.

Filter then rank, run

I ran this in a Python virtual environment with numpy 2.5.3, scikit-learn 1.9.1 and faiss-cpu 1.15.1, with fixed random seeds so the numbers repeat. Without a filter the old 2022 India policy and the UK policy win on raw similarity score. Filtering to India and 2025 first leaves chunks 1 and 4, ranked by score.

chunks = [
    {"id": 1, "text": "Leave policy for India", "country": "IN", "year": 2025, "score": 0.91},
    {"id": 2, "text": "Leave policy for UK",    "country": "UK", "year": 2025, "score": 0.93},
    {"id": 3, "text": "Old leave policy India", "country": "IN", "year": 2022, "score": 0.95},
    {"id": 4, "text": "Travel policy India",    "country": "IN", "year": 2025, "score": 0.70},
]
def top_k(items, k, **where):
    ok = [c for c in items if all(c[f] == v for f, v in where.items())]     # hard filter, then rank by similarity
    return [c["id"] for c in sorted(ok, key=lambda c: -c["score"])[:k]]

print("no filter     :", top_k(chunks, 2))
print("IN + 2025 only:", top_k(chunks, 2, country="IN", year=2025))

Output:

no filter     : [3, 2]
IN + 2025 only: [1, 4]

Over-fetch before post-filtering

Ask for 5 to 10 times k, filter, then cut to k, or use pre-filtering so you never return too few results.

Quick check: Where should "this user may see only their own documents" be enforced?

  • Only in the prompt
  • In the retrieval query as a filter
  • Only in the UI
  • Nowhere
Answer

In the retrieval query as a filter — Documents the user may not see must never be retrieved.