# Case Study: A Policy Question-Answering Assistant — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/v-case

> Design an internal assistant that answers HR policy questions with citations.

## The design

Goal: employees ask HR questions and get accurate, cited answers. **Data**: policy PDFs split into overlapping chunks with metadata (document, section, date, country). **Retrieval**: hybrid keyword + embedding search filtered by the employee's country, then re-ranking. **Prompt**: system rules ("answer only from the provided excerpts, quote the section, say you don't know otherwise"), the excerpts, the question; temperature low; JSON output with `answer` and `citations`. **Safeguards**: code checks that each citation exists in the retrieved excerpts, a prompt-injection filter on documents, no personal data in logs, and escalation to a human for legal or sensitive topics. **Evaluation**: 100 real questions with reference answers, scored for correctness, groundedness and refusal behaviour on every change. **Operations**: monitor cost, latency, thumbs-up/down feedback and unanswered questions, and refresh the index when policies change.

## From idea to a reliable feature

Combine prompting, retrieval, evaluation and safeguards into a dependable feature.

![Four habits: ground, test, guard, monitor.](assets/figures/llms/section-8-map.svg) — Figure 8.1 — Ground, test, guard and monitor.

## The design on one page

Each line maps to a section of this course.

```text
Tokens/cost     count tokens, cap output, cache the fixed prompt        (Sec 1, 5)
Retrieval       chunk + metadata filter + hybrid search + re-rank       (Sec 2, 5)
Prompt          answer only from excerpts, cite, abstain if unsure      (Sec 5)
Sampling        low temperature, JSON schema output                     (Sec 2, 5)
Guardrails      citation check in code, injection filter, no PII logs   (Sec 6)
Evaluation      100-question test set run on every change               (Sec 6)
Ops             monitor cost/latency/feedback, refresh index            (Sec 7)
```

**Quiz:** Why does the assistant verify citations in code?

- [x] Models can invent citations, so they must be checked against retrieved text
- [ ] To make responses longer
- [ ] To train the model
- [ ] Citations are never wrong

*Answer:* Models can invent citations, so they must be checked against retrieved text. Programmatic verification catches unsupported claims before users see them.
