Lesson 23 / 25

Incident Response and the Kill Switch

Contain, investigate and learn when an agent does something it should not have.

Stop, revoke, review, fix

When something goes wrong: stop the agent (a kill switch or process kill); revoke any credentials it could have used, and rotate anything that might have leaked; review the audit log and the Git history to learn exactly what changed; revert the damage using Git and backups; then fix the guardrail that failed and add a test. Run a short blameless review, because the lesson is about the system, not a person.

A response checklist

Write it down before you need it. Under stress people skip steps.

1. STOP     kill the agent process / disable its token
2. REVOKE   rotate keys it could read; revoke its Git and cloud access
3. REVIEW   audit log + git log --stat + git diff for the whole session
4. REVERT   git revert / restore from backup; check deployed state
5. FIX      patch the failed guardrail, add a red-team test
6. LEARN    short blameless write-up; update the project guide

Practise the kill switch

Know how to stop an agent in under a minute and test it occasionally. The first time you look for the off switch should not be during a real incident.

Quick check: An agent may have exposed a key. What comes right after stopping it?

  • Revoke or rotate the key
  • Write a blog post
  • Wait a week
  • Rename the repository
Answer

Revoke or rotate the key — Until the credential is revoked, anyone who captured it can keep using it.