Lesson 3 / 28

Blast Radius: What Can a Compromised Model Do

Measure risk by what the model is allowed to reach.

Assume injection works; limit the damage

Because prompt injection cannot be fully prevented today, the most reliable security question is not "can the model be tricked?" but "if it is tricked, what is the worst it can do?" That worst case is the blast radius, and it is set by the capabilities you hand over. A summariser that can only return text has a small radius (at worst, misleading text). Add per-user data access and a successful injection can leak that user's data; add the ability to read any customer's data and it can leak everyone's; add email and it can exfiltrate data to an attacker; add refunds and it can move money; add shell commands and it can compromise the host. A dangerous combination is the "lethal trifecta": access to private data, exposure to untrusted content, and a way to communicate externally. Reduce risk by removing at least one of the three.

Capabilities and blast radius, run

I ran this with plain Python 3 (standard library only). All attacks here are harmless demonstrations on local data, using no real systems. Each added capability raises the worst-case impact of a successful injection, from 0 (text only) to 6 (shell access). The scores are a teaching scale, not a standard.

# Estimating how much an attacker gains from a successful injection depends on what the app can reach.
capabilities = {
    "summarise text only": 0,
    "+ read the user's own orders": 1,
    "+ read any customer's data": 3,
    "+ send email": 4,
    "+ issue refunds": 5,
    "+ run shell commands": 6,
}
print("capability                         blast radius (0-6)")
for name, level in capabilities.items(): print(f"{name:34} {'#' * level}{'.' * (6 - level)} {level}")

Output:

capability                         blast radius (0-6)
summarise text only                ...... 0
+ read the user's own orders       #..... 1
+ read any customer's data         ###... 3
+ send email                       ####.. 4
+ issue refunds                    #####. 5
+ run shell commands               ###### 6

Remove one leg of the trifecta

If private data, untrusted content and outbound communication cannot all be present together, a whole class of attacks loses its payoff.

Quick check: Which combination is the "lethal trifecta"?

  • Logging + metrics + dashboards
  • Fast GPU + big RAM + cheap storage
  • Short prompt + low temperature + small model
  • Private data + untrusted content + external communication
Answer

Private data + untrusted content + external communication — With all three, an injection can read secrets and send them out.