Lesson 4 / 28
A Map of the Main Risk Families
Preview the risks this course covers and how they connect.
Ten-ish risks, four themes
The usual risk lists group into four themes. Manipulation: prompt injection (direct and indirect), jailbreaks, system-prompt leakage. Unsafe consequences: insecure handling of model output (XSS, SQL injection, command injection, path traversal, SSRF), excessive agency, over-reliance on wrong answers. Data and knowledge: sensitive-information disclosure, training-data or retrieval poisoning, weaknesses in embeddings and vector stores, RAG access-control failures. Platform and resources: supply chain (models, packages, plugins, serialised files), unbounded consumption (denial of service, denial of wallet), and monitoring gaps. The fixes come from the same short list of principles: least privilege, validate everything crossing a boundary, keep humans in the loop for risky actions, separate data from instructions where you can, limit resources, log and test. The rest of the course turns these into concrete practice.
Risk families and main controls
A cheat sheet for the sections ahead.
Theme Risk Main controls
Manipulation prompt injection, jailbreak, prompt leak data/instruction separation, least privilege, approvals, never keep secrets in prompts
Unsafe consequence insecure output handling escape / parameterise / allow-list / validate before use
excessive agency narrow tools, per-user auth, human approval, audit
Data & knowledge sensitive disclosure, RAG leaks, poisoning access control in retrieval, redaction, provenance, minimal retention
Platform supply chain, unbounded consumption pin & verify models/packages, safe formats, rate and spend limits, monitoringUse a maintained checklist
Start from the OWASP Top 10 for LLM Applications and adapt it to your system; the list changes over time.
Quick check: Which principle underlies most of the fixes?
- Using longer prompts
- Least privilege and validating everything that crosses a boundary
- Hiding the model name
- Using more GPUs
Answer
Least privilege and validating everything that crosses a boundary — Classic security principles still apply and do most of the work.