Lesson 2 / 28
Assets, Attackers and Trust Boundaries
Draw the system and mark where untrusted text enters.
Start with what you must protect
Threat modelling for an LLM app starts like any other. Assets: customer data, internal documents, credentials and API keys, the system prompt and business logic, the money behind the API bill, the reputation of the product, and the actions the app can take (send email, move money, change records). Attackers: a curious or malicious user; an outsider whose text reaches the model indirectly (a web page, an email, a PDF, a review); a compromised third-party (plugin, dataset, model provider); and an insider. Trust boundaries: draw every place where text from a less-trusted source enters the prompt (user input, retrieved documents, tool outputs, web content, uploaded files, other agents) and every place where model output leaves for a more privileged component (a database, a browser, a shell, an email system, another service). Each crossing is a place for a control.
A trust-boundary worksheet
Fill one in for your own application.
ENTRY POINTS (untrusted text -> prompt) EXIT POINTS (model output -> privileged component)
user chat messages HTML page in the browser (XSS risk)
retrieved documents / web pages SQL / shell / file path built from output (injection)
uploaded files, emails, tickets tool calls (send email, refund, delete) (agency)
tool outputs, API responses, other agents logs, analytics, other users' views (data leak)
For each line ask: who controls this text? what can the next component do with it? what stops abuse?Include indirect sources
List content the model reads on a user's behalf (emails, pages, files); that is where indirect attacks arrive.
Quick check: Which is an entry point where untrusted text can reach the prompt?
- A retrieved web page or document
- The GPU driver version
- The Wi-Fi password
- The monitor brightness
Answer
A retrieved web page or document — Any content the model reads is a potential instruction source.