Lesson 28 / 28

Revision: Cheat Sheet and Self-Check

Review the key ideas of the whole course.

Cheat sheet

Core: instructions and data share one text stream, so any text the model reads can steer it; assume injection works and limit the blast radius; the lethal trifecta = private data + untrusted content + external communication. Injection: direct vs indirect; filters are one thin layer (bypassed by paraphrase, spacing, Unicode, encodings); fence and label data; canary tokens detect prompt leaks; never keep secrets in prompts; capability limits are the strongest layer. Output handling: treat model output as untrusted input: parameterised SQL, HTML escaping and safe Markdown rendering (watch image/link exfiltration), no shell (argument lists + allow-lists), path confinement, URL allow-lists with parsed hostnames and SSRF protection, schema validation. Agency: narrow tools, minimum credentials, per-user authorisation (avoid the confused deputy), default-deny policy engine with argument limits, human approval bound to the exact action (signed), audit logs with redaction. Data: minimise and redact PII, access control inside retrieval, vector-store security, provenance and ingestion hygiene against poisoning. Platform: pin and verify models and packages, avoid code-executing formats such as pickle, vet plugins and MCP servers, rate limits, token/step caps, spend caps. Operate: attack suite and regression tests, monitoring and alerts (canary hits, denials, cost spikes), kill switches, incident runbook, fix controls not wording.

Quick check: A retrieved web page tells the assistant to email the customer list to an outsider. Which control most reliably prevents harm?

  • A polite system prompt
  • The assistant has no unrestricted email tool and cannot access the whole customer list
  • A keyword blocklist alone
  • A bigger context window
Answer

The assistant has no unrestricted email tool and cannot access the whole customer list — Capability limits bound damage even when manipulation succeeds.

Quick check: Model output is placed into an SQL string. What is the correct fix?

  • Ask the model to avoid quotes
  • Use parameterised queries (and least-privilege database access)
  • Remove the database
  • Use uppercase
Answer

Use parameterised queries (and least-privilege database access) — Parameters keep values from being interpreted as SQL.

Quick check: Why must RAG permission filtering happen in the retrieval query?

  • It is only cosmetic
  • It makes retrieval slower on purpose
  • Models cannot read permissions
  • Text that reaches the prompt can leak, whatever the model is told to hide
Answer

Text that reaches the prompt can leak, whatever the model is told to hide — Keep forbidden content out of the prompt entirely.