Lesson 30 / 31

Case Study: A Documentation Assistant

Design an assistant that answers questions about product docs with citations.

The design

Goal: developers ask questions about a product's docs and get answers with links. Ingestion: load Markdown docs, split by headings into chunks of about 300 tokens (with overlap), attach metadata (page URL, section, version), embed and store in a vector database; re-index changed pages nightly. Query path (LangChain chain, or a LlamaIndex query engine): condense follow-ups with chat history, retrieve with a version filter, rerank, format numbered passages, prompt "answer only from the passages and cite", parse to a Pydantic object {answer, citations}, and verify citations in code. Tools (optional): a read-only code-search tool, with a step limit. Reliability: retries with backoff on the model call, a fallback model, timeouts, streaming to the UI. Quality: 120-question dataset with unanswerable cases, recall@5 and faithfulness tracked in CI, traces for every request. Security: keys in a secrets manager, retrieved text treated as data, no write-capable tools, logs redacted.

A tested, traceable assistant

Combine retrieval, chains, tools, guardrails and tests into one dependable application.

Four habits: compose, ground, guard, test.
Figure 8.1 — Compose, ground, guard and test.

The design on one page

Each line maps to a section of this course.

Ingest     Markdown -> heading-aware chunks (~300 tok) -> metadata -> vector DB      (Sec 4, 5)
Query       condense -> filtered retrieve -> rerank -> numbered passages -> prompt      (Sec 2, 4)
Output      Pydantic {answer, citations} + code check that citations exist              (Sec 2)
Reliability retry+backoff, fallback model, timeouts, streaming                          (Sec 2, 3)
Agent/tools read-only code search, step limit, no write tools                           (Sec 3, 6)
Quality     120-question dataset, recall@5 + faithfulness in CI, traces                 (Sec 7)
Security    secrets manager, data-not-instructions, redacted logs, pinned versions      (Sec 7)

Quick check: Why does the design give the agent only read-only tools?

  • To avoid citations
  • Read-only tools are faster
  • Write tools do not exist
  • To limit harm if the model is tricked by injected text
Answer

To limit harm if the model is tricked by injected text — Least privilege bounds what a compromised agent can do.