Lesson 27 / 28
Case Study: Semantic Search for a Help Centre
Design the full pipeline for 50,000 help articles in English and Hindi.
The design
Goal: users search 50,000 help articles in English and Hindi, and an assistant answers from them. Data: split each article by headings into ~300-token chunks (about 400,000 chunks) with the article title prepended; metadata for product, language, audience (public/internal), version and URL. Model: shortlist two multilingual embedding models from a leaderboard, then pick by recall@5 on 150 labelled questions (including Hindi, English and mixed queries and 20 unanswerable ones); record the model name and version. Storage: 400,000 × 768 × 4 bytes is about 1.2 GB, so a flat or HNSW index fits in memory; start with exact search for ground truth, then HNSW tuned to recall@10 at least 0.98 against flat. Retrieval: filters first (audience, language, product), hybrid BM25 + vector with RRF, top 50, cross-encoder rerank to 5, a minimum-score threshold for "not found". Operations: nightly incremental updates keyed by article ID and content hash, delete-by-source, parallel index for model upgrades with golden-set comparison before switching, query-embedding cache, p95 latency and recall dashboards, access control enforced in the query, redacted logs.
Embed, index, retrieve, measure
A measured pipeline of good chunks, a suitable model, the right index and a golden set makes search dependable.
The design on one page
Each line maps to a section of this course.
Chunks heading-aware ~300 tok + title prefix + metadata (product, lang, audience, version) (Sec 2)
Model 2 multilingual candidates -> pick by recall@5 on 150 own questions (Hindi/English/mixed) (Sec 2, 5)
Storage ~400k x 768 x 4B = ~1.2 GB -> fits memory; exact first, then HNSW (>= 0.98 recall@10) (Sec 3)
Retrieval filters -> BM25 + vector (RRF, top 50) -> cross-encoder -> best 5 -> score threshold (Sec 4)
Operations incremental updates by id + hash, parallel index for model upgrades, query-embedding cache (Sec 5, 6)
Safety ACL filter inside the query, vectors = source-level sensitivity, redacted logs (Sec 6)Start with exact search and a golden set
Get a working baseline with exact search and measured recall before adding any index, reranker or compression.
Quick check: Why is exact (flat) search used first in this design?
- Because flat is always faster
- Because HNSW cannot work with 400,000 vectors
- To avoid embeddings
- To get ground truth for tuning the HNSW index and measuring recall
Answer
To get ground truth for tuning the HNSW index and measuring recall — Approximate recall is only meaningful relative to exact answers.