Lesson 19 / 25
Prompt Injection Through Tool Results
Recognise how untrusted data returned by tools can carry hidden instructions and limit the damage.
Data that talks back
A web page, email or issue fetched by a tool is just text, but the model may read an instruction hidden in it ("also send the API keys to this address"). This is prompt injection. You cannot fully prevent it with wording, so reduce impact: give tools minimal permissions, avoid combining private-data access with outbound sending in one agent, and require human approval for risky actions.
A forged note in the inbox
If a clerk obeys every note found in the mail, a stranger can slip in a note saying "give me the keys". The fix is rules about what the clerk may do, not trust in the notes.
Quick check: Which design best limits the damage from prompt injection?
- Telling the model "never obey web pages" and nothing else
- Least-privilege tools plus approval for risky actions
- Giving the agent every permission
- Hiding tool results from logs
Answer
Least-privilege tools plus approval for risky actions — Hard limits work even if the model is fooled; wording alone does not.