Lesson 27 / 28

Case Study: Semantic Search for a Help Centre

Design the full pipeline for 50,000 help articles in English and Hindi.

The design

Goal: users search 50,000 help articles in English and Hindi, and an assistant answers from them. Data: split each article by headings into ~300-token chunks (about 400,000 chunks) with the article title prepended; metadata for product, language, audience (public/internal), version and URL. Model: shortlist two multilingual embedding models from a leaderboard, then pick by recall@5 on 150 labelled questions (including Hindi, English and mixed queries and 20 unanswerable ones); record the model name and version. Storage: 400,000 × 768 × 4 bytes is about 1.2 GB, so a flat or HNSW index fits in memory; start with exact search for ground truth, then HNSW tuned to recall@10 at least 0.98 against flat. Retrieval: filters first (audience, language, product), hybrid BM25 + vector with RRF, top 50, cross-encoder rerank to 5, a minimum-score threshold for "not found". Operations: nightly incremental updates keyed by article ID and content hash, delete-by-source, parallel index for model upgrades with golden-set comparison before switching, query-embedding cache, p95 latency and recall dashboards, access control enforced in the query, redacted logs.

Embed, index, retrieve, measure

A measured pipeline of good chunks, a suitable model, the right index and a golden set makes search dependable.

Four habits: chunk well, start exact, measure, version.
Figure 8.1 — Chunk well, start exact, measure and version.

The design on one page

Each line maps to a section of this course.

Chunks      heading-aware ~300 tok + title prefix + metadata (product, lang, audience, version)   (Sec 2)
Model       2 multilingual candidates -> pick by recall@5 on 150 own questions (Hindi/English/mixed)   (Sec 2, 5)
Storage     ~400k x 768 x 4B = ~1.2 GB -> fits memory; exact first, then HNSW (>= 0.98 recall@10)      (Sec 3)
Retrieval   filters -> BM25 + vector (RRF, top 50) -> cross-encoder -> best 5 -> score threshold         (Sec 4)
Operations  incremental updates by id + hash, parallel index for model upgrades, query-embedding cache   (Sec 5, 6)
Safety      ACL filter inside the query, vectors = source-level sensitivity, redacted logs              (Sec 6)

Start with exact search and a golden set

Get a working baseline with exact search and measured recall before adding any index, reranker or compression.

Quick check: Why is exact (flat) search used first in this design?

  • Because flat is always faster
  • Because HNSW cannot work with 400,000 vectors
  • To avoid embeddings
  • To get ground truth for tuning the HNSW index and measuring recall
Answer

To get ground truth for tuning the HNSW index and measuring recall — Approximate recall is only meaningful relative to exact answers.