Lesson 23 / 26

AI Incident Management

Define severity levels, response steps and learning reviews for AI-related incidents.

Prepare before it happens

AI incidents include harmful or false outputs reaching users, personal-data exposure, a biased outcome discovered in audit, or a vendor model change that breaks behaviour. Prepare: a way for staff and users to report; severity levels with response times; a runbook (contain, assess, notify, fix, review); named decision-makers; and notification rules for regulators and affected people where laws require it. After each incident run a blameless review, fix the system rather than the person, and add the case to your evaluation set.

Keep watching after launch

Governance continues after launch: incidents are handled, outcomes are measured and findings feed back.

Three loops: detect, respond, learn.
Figure 7.1 — Detect, respond and learn.

Severity classification, run

I ran this. Any data exposure is SEV1, a harmful output that reached over 100 users is SEV1, a smaller harmful output is SEV2, and a minor issue is SEV3.

def severity(users_affected, harmful, data_exposed):
    if data_exposed or (harmful and users_affected > 100):
        return "SEV1"
    if harmful or users_affected > 100:
        return "SEV2"
    return "SEV3"

print(severity(5, False, True), severity(500, True, False), severity(10, False, False))

Output:

SEV1 SEV1 SEV3

Rehearse once

Run a short tabletop exercise: "the assistant just told customers the wrong refund policy; what do we do in the first hour?" Gaps in ownership and contact lists show up immediately and cheaply.

Quick check: What is the point of a blameless post-incident review?

  • To fix the system and process so it does not recur
  • To punish the person involved
  • To keep the incident secret
  • To avoid learning anything
Answer

To fix the system and process so it does not recur — Blame hides information; focusing on the system surfaces causes and improves safeguards.