Even well-tested AI systems fail in production. A prepared response limits the harm.
What Counts as an AI Incident
- Harmful, offensive or dangerous outputs.
- Systematically unfair outcomes for a group.
- Leaks of personal or confidential data.
- An agent taking unintended actions.
- Significant accuracy degradation affecting decisions.
- Successful manipulation, such as prompt injection.
Detection
Monitoring dashboards, automated output checks, user reports, complaints and staff observations. Make it easy for anyone — users and employees — to report a problem.
Containment
Have options ready: disable a feature, fall back to a simpler system or human process, restrict an agent's permissions, raise a confidence threshold, or roll back to a previous model or prompt version.
Investigation
Reconstruct what happened from logs: inputs, outputs, model and prompt versions, retrieved content and tool calls. Identify affected people and the root cause — data, model behaviour, integration, or misuse.
Remediation and Communication
Fix the cause, correct affected decisions where possible, notify affected people and regulators when required (for example, for personal data breaches), and communicate honestly.
Learn From It
Hold a blameless review, add the failure to evaluation test sets, update documentation and model cards, and share lessons across teams.
Prepare in Advance
Keep logs that make investigation possible, assign owners, write down the playbook, and practise it before you need it.