# What Can Go Wrong — AI Coding-Agent Guardrails

Source: https://www.geekswithgeeks.com/en/coding-agent-guardrails/tm-risk-taxonomy

> Classify the main risks: destructive actions, leaks, injection, supply chain and runaway cost.

## Six families of risk

(1) **Destructive actions**: `rm -rf`, dropping a database, force-pushing. (2) **Secret exposure**: reading `.env` or sending keys to a model or a server. (3) **Prompt injection**: instructions hidden in files, issues or web pages. (4) **Supply chain**: installing a wrong or malicious package. (5) **Scope creep**: editing files you never asked about, such as CI or infrastructure. (6) **Runaway cost or time**: loops that never end. Each needs its own control.

## Risk to control map

Use it as a checklist: every row should have at least one control you can name.

```text
Risk                    First control
Destructive actions     command policy + sandbox + branch
Secret exposure         deny secret paths, scrub env, redact
Prompt injection        least privilege, no network by default
Supply chain            approval for installs, lockfile review
Scope creep             protected paths, diff size gate
Runaway cost/time       step, time and token budgets
```

## Start with the worst-case

Ask: if this agent were fully tricked, what is the worst it could do with the access it has today? Reduce that worst case first, then work on likelihood.

**Quiz:** An agent edits the CI workflow when asked to fix a typo. Which risk family is this?

- [x] Scope creep
- [ ] Network latency
- [ ] Font rendering
- [ ] Cache misses

*Answer:* Scope creep. Changing files outside the task is scope creep, controlled by protected paths and diff review.
