Lesson 27 / 28
Case Study: Securing a Customer-Support Assistant
Apply the whole course to one realistic assistant.
The design
The assistant answers customers about orders, reads their order history, searches a help centre and can propose refunds. Threat model: untrusted text enters from customer messages, uploaded screenshots' text, retrieved help articles (some user-contributed) and web pages it summarises; the lethal-trifecta risk is private order data plus untrusted content plus email/refund actions, so at least one leg is removed (no free-form outbound email: it drafts text and a human sends). Tools: get_order(order_id) runs as the authenticated customer with a read-only account; propose_refund(order, amount<=5000) is queued for human approval and signed to the exact amount; no shell, no generic SQL, no file tool; fetching URLs only from an allow-list through an isolated fetcher. Prompting: untrusted text fenced and labelled as data; a canary token in the system prompt; no secrets in prompts. Output: replies escaped when rendered, no auto-rendered remote images, structured fields validated, URLs filtered. Data: minimum customer data sent to the model, personal data redacted from logs, retrieval filtered by tenant and article visibility, ingestion strips hidden text. Limits: per-user rate limits, token and step caps, daily spend cap with alert. Supply chain: pinned dependencies, models from verified sources in a safe format, tools reviewed. Operations: attack suite in CI (injection, extraction, cross-user, tool abuse, output payloads, Hindi and Hinglish), audit logs, alerts on canary hits and denials, an incident runbook with kill switches.
Assume manipulation, contain the damage
A secure LLM app limits what the model can reach, validates what it produces, controls access to data and watches for abuse.
The design on one page
Each line maps to a section of this course.
Threat model untrusted text sources listed; lethal trifecta broken (human sends emails) (Sec 1)
Injection fenced + labelled data, canary token, no secrets in prompts, filters as monitoring (Sec 2)
Output escape on render, no auto-images, validate fields, URL allow-list (Sec 3)
Agency read-only per-customer tool, refund proposals signed + human approved, default-deny (Sec 4)
Data minimum data, redacted logs, retrieval filtered by tenant, ingestion strips hidden text (Sec 5)
Platform pinned deps, safe model formats, rate limits, token/step/spend caps (Sec 6)
Operations attack suite in CI, audit logs, alerts, kill switches, runbook (Sec 7)Re-run the threat model for each new tool
Every added capability changes the blast radius.
Quick check: Which design choice breaks the "lethal trifecta" in this assistant?
- Using more logging
- Using a larger model
- Using a longer prompt
- The assistant drafts emails and a human sends them, so it has no free outbound channel
Answer
The assistant drafts emails and a human sends them, so it has no free outbound channel — Removing one leg of the trifecta prevents exfiltration through that route.