Lesson 25 / 31

Guardrails for Agentic Apps

Limit what agents can do and check what they produce.

Autonomy needs boundaries

Agents act on text, and text can be hostile (prompt injection in a web page, email or document). Defences: give tools least privilege (read-only where possible, scoped credentials, allow-listed actions and domains); validate tool arguments in code; require human approval for irreversible or costly actions (payments, deletions, external emails); run code in a sandbox; treat retrieved or tool output as data, not instructions; set step, token, time and cost limits; log every call for audit; and test with adversarial examples. Prefer simple, constrained workflows over open-ended autonomy whenever you can.

Ask: what is the worst this tool can do?

If the answer is frightening, add an approval step or remove the tool.

Quick check: Which action most clearly needs human approval?

  • Formatting a date
  • Reading a public FAQ
  • Issuing a customer refund
  • Counting words
Answer

Issuing a customer refund — Irreversible financial actions deserve a human check.