Lesson 9 / 28

The Rule: Treat Model Output Like User Input

Apply the classic injection defences to everything the model produces.

Where output flows, attacks follow

If model output flows into another component, an attacker who can influence the model (by their own message or by planting text in content the model reads) may control what that component receives. Build the string of a SQL query, an HTML page, a shell command, a file path or a URL from model output and you recreate the classic injection bugs: SQL injection, cross-site scripting (XSS), command injection, path traversal and server-side request forgery (SSRF). The fix is the same as ever, applied at every sink: never build code from strings; use parameterised queries, output encoding/escaping, argument lists instead of shells, path canonicalisation with an allowed base folder, and allow-lists for hosts. Also validate model output against a schema and treat a failed validation as an error, not something to guess around.

Model output is untrusted input

Anything a model writes can be shaped by an attacker, so escape, parameterise and validate it before it reaches another system.

Five sinks: SQL, HTML, shell, files, URLs.
Figure 3.1 — SQL, HTML, shell, files and URLs.

Sinks and their defences

Each row is a classic bug applied to model output.

Sink (where output goes)     Attack if built from strings        Defence
SQL query                     SQL injection                       parameterised queries (?)
HTML page                     XSS (script in the reply)           escape / sanitise; strict CSP
shell command                 command injection                   no shell; argument list; allow-list of commands
file path                     path traversal (../../secret)       realpath + must stay inside allowed folder
URL fetched by a tool         SSRF (internal/metadata hosts)      https + host allow-list; block private ranges
code / eval / template        arbitrary code execution            never eval; sandbox if code must run

Validate against a schema

Ask for structured output and reject anything that does not match, instead of guessing what the model meant.

Quick check: How should model output be treated when it is used to build a database query?

  • As raw SQL to execute
  • As trusted because the model wrote it
  • As untrusted input passed through parameterised queries
  • As a password
Answer

As untrusted input passed through parameterised queries — Origin in a model does not make text safe.