Lesson 28 / 28
Revision: Cheat Sheet and Self-Check
Review the key ideas of the whole course.
Cheat sheet
Core: instructions and data share one text stream, so any text the model reads can steer it; assume injection works and limit the blast radius; the lethal trifecta = private data + untrusted content + external communication. Injection: direct vs indirect; filters are one thin layer (bypassed by paraphrase, spacing, Unicode, encodings); fence and label data; canary tokens detect prompt leaks; never keep secrets in prompts; capability limits are the strongest layer. Output handling: treat model output as untrusted input: parameterised SQL, HTML escaping and safe Markdown rendering (watch image/link exfiltration), no shell (argument lists + allow-lists), path confinement, URL allow-lists with parsed hostnames and SSRF protection, schema validation. Agency: narrow tools, minimum credentials, per-user authorisation (avoid the confused deputy), default-deny policy engine with argument limits, human approval bound to the exact action (signed), audit logs with redaction. Data: minimise and redact PII, access control inside retrieval, vector-store security, provenance and ingestion hygiene against poisoning. Platform: pin and verify models and packages, avoid code-executing formats such as pickle, vet plugins and MCP servers, rate limits, token/step caps, spend caps. Operate: attack suite and regression tests, monitoring and alerts (canary hits, denials, cost spikes), kill switches, incident runbook, fix controls not wording.
Quick check: A retrieved web page tells the assistant to email the customer list to an outsider. Which control most reliably prevents harm?
- A polite system prompt
- The assistant has no unrestricted email tool and cannot access the whole customer list
- A keyword blocklist alone
- A bigger context window
Answer
The assistant has no unrestricted email tool and cannot access the whole customer list — Capability limits bound damage even when manipulation succeeds.
Quick check: Model output is placed into an SQL string. What is the correct fix?
- Ask the model to avoid quotes
- Use parameterised queries (and least-privilege database access)
- Remove the database
- Use uppercase
Answer
Use parameterised queries (and least-privilege database access) — Parameters keep values from being interpreted as SQL.
Quick check: Why must RAG permission filtering happen in the retrieval query?
- It is only cosmetic
- It makes retrieval slower on purpose
- Models cannot read permissions
- Text that reaches the prompt can leak, whatever the model is told to hide
Answer
Text that reaches the prompt can leak, whatever the model is told to hide — Keep forbidden content out of the prompt entirely.