Lesson 26 / 27

Case Study: A Policy Question-Answering Assistant

Design an internal assistant that answers HR policy questions with citations.

The design

Goal: employees ask HR questions and get accurate, cited answers. Data: policy PDFs split into overlapping chunks with metadata (document, section, date, country). Retrieval: hybrid keyword + embedding search filtered by the employee's country, then re-ranking. Prompt: system rules ("answer only from the provided excerpts, quote the section, say you don't know otherwise"), the excerpts, the question; temperature low; JSON output with answer and citations. Safeguards: code checks that each citation exists in the retrieved excerpts, a prompt-injection filter on documents, no personal data in logs, and escalation to a human for legal or sensitive topics. Evaluation: 100 real questions with reference answers, scored for correctness, groundedness and refusal behaviour on every change. Operations: monitor cost, latency, thumbs-up/down feedback and unanswered questions, and refresh the index when policies change.

From idea to a reliable feature

Combine prompting, retrieval, evaluation and safeguards into a dependable feature.

Four habits: ground, test, guard, monitor.
Figure 8.1 — Ground, test, guard and monitor.

The design on one page

Each line maps to a section of this course.

Tokens/cost     count tokens, cap output, cache the fixed prompt        (Sec 1, 5)
Retrieval       chunk + metadata filter + hybrid search + re-rank       (Sec 2, 5)
Prompt          answer only from excerpts, cite, abstain if unsure      (Sec 5)
Sampling        low temperature, JSON schema output                     (Sec 2, 5)
Guardrails      citation check in code, injection filter, no PII logs   (Sec 6)
Evaluation      100-question test set run on every change               (Sec 6)
Ops             monitor cost/latency/feedback, refresh index            (Sec 7)

Quick check: Why does the assistant verify citations in code?

  • Models can invent citations, so they must be checked against retrieved text
  • To make responses longer
  • To train the model
  • Citations are never wrong
Answer

Models can invent citations, so they must be checked against retrieved text — Programmatic verification catches unsupported claims before users see them.