Lesson 25 / 31
Guardrails for Agentic Apps
Limit what agents can do and check what they produce.
Autonomy needs boundaries
Agents act on text, and text can be hostile (prompt injection in a web page, email or document). Defences: give tools least privilege (read-only where possible, scoped credentials, allow-listed actions and domains); validate tool arguments in code; require human approval for irreversible or costly actions (payments, deletions, external emails); run code in a sandbox; treat retrieved or tool output as data, not instructions; set step, token, time and cost limits; log every call for audit; and test with adversarial examples. Prefer simple, constrained workflows over open-ended autonomy whenever you can.
Ask: what is the worst this tool can do?
If the answer is frightening, add an approval step or remove the tool.
Quick check: Which action most clearly needs human approval?
- Formatting a date
- Reading a public FAQ
- Issuing a customer refund
- Counting words
Answer
Issuing a customer refund — Irreversible financial actions deserve a human check.