Lesson 4 / 28

A Map of the Main Risk Families

Preview the risks this course covers and how they connect.

Ten-ish risks, four themes

The usual risk lists group into four themes. Manipulation: prompt injection (direct and indirect), jailbreaks, system-prompt leakage. Unsafe consequences: insecure handling of model output (XSS, SQL injection, command injection, path traversal, SSRF), excessive agency, over-reliance on wrong answers. Data and knowledge: sensitive-information disclosure, training-data or retrieval poisoning, weaknesses in embeddings and vector stores, RAG access-control failures. Platform and resources: supply chain (models, packages, plugins, serialised files), unbounded consumption (denial of service, denial of wallet), and monitoring gaps. The fixes come from the same short list of principles: least privilege, validate everything crossing a boundary, keep humans in the loop for risky actions, separate data from instructions where you can, limit resources, log and test. The rest of the course turns these into concrete practice.

Risk families and main controls

A cheat sheet for the sections ahead.

Theme              Risk                                   Main controls
Manipulation       prompt injection, jailbreak, prompt leak   data/instruction separation, least privilege, approvals, never keep secrets in prompts
Unsafe consequence insecure output handling                   escape / parameterise / allow-list / validate before use
                   excessive agency                           narrow tools, per-user auth, human approval, audit
Data & knowledge   sensitive disclosure, RAG leaks, poisoning access control in retrieval, redaction, provenance, minimal retention
Platform           supply chain, unbounded consumption         pin & verify models/packages, safe formats, rate and spend limits, monitoring

Use a maintained checklist

Start from the OWASP Top 10 for LLM Applications and adapt it to your system; the list changes over time.

Quick check: Which principle underlies most of the fixes?

  • Using longer prompts
  • Least privilege and validating everything that crosses a boundary
  • Hiding the model name
  • Using more GPUs
Answer

Least privilege and validating everything that crosses a boundary — Classic security principles still apply and do most of the work.