Lesson 26 / 27

Case Study: A Support Knowledge-Base Assistant

Design a RAG assistant for customer-support articles, with permissions and citations.

The design

Goal: support agents ask questions and get answers with links to help-centre articles. Ingestion: nightly crawl and a webhook on edits; extract clean text, keep headings, split by section into 300-token chunks with the article title prepended; metadata for product, language, audience (public or internal), article version and URL. Retrieval: hybrid BM25 + embeddings merged with RRF, filtered by product and audience, top 30 reranked to 5. Generation: instructions to answer only from numbered excerpts, cite as [n], reply "not found" otherwise, JSON output. Safeguards: code verifies citation IDs, a groundedness check flags weak answers, injection-resistant wording, and internal articles never reach external users. Evaluation: 150-question golden set including 20 unanswerable questions, tracking recall@5, faithfulness, refusal accuracy, latency and cost on every change. Operations: incremental re-indexing, deletion handling, feedback buttons and a weekly review of failed questions.

A measured, trustworthy assistant

Combine ingestion, hybrid retrieval, grounded generation and evaluation into one dependable system.

Four habits: clean data, hybrid search, grounded answers, constant measurement.
Figure 8.1 — Clean data, hybrid search, grounded answers and measurement.

The design on one page

Each line maps to a section of this course.

Ingest      clean text, section chunks (300 tok), title prefix, metadata   (Sec 2)
Retrieve    BM25 + embeddings -> RRF -> filter -> rerank 30->5              (Sec 3, 4)
Generate    numbered excerpts, cite [n], "not found" allowed, JSON          (Sec 5)
Guard       citation-ID check, groundedness flag, injection wording, ACLs   (Sec 5, 7)
Evaluate    150-question golden set incl. unanswerable; recall@5, faithful  (Sec 6)
Operate     incremental index, deletes, feedback, weekly failure review     (Sec 7)

Quick check: Why are internal articles filtered out in the retrieval query?

  • So external users can never receive internal content
  • To speed up the font
  • To enlarge chunks
  • Because embeddings dislike them
Answer

So external users can never receive internal content — Permissions must be enforced before text reaches the model.