
- 01AI incidents fall into four categories. Unauthorized tool invocation, tenant data crossing a boundary, runaway token consumption, and model output used to drive a downstream attack.
- 02The first responder action for any AI incident is to disable the affected agent or feature, not to investigate while it is still running. The blast radius grows with every additional model call.
- 03Reconstruction requires the audit log to contain the full model input, output, and tool call sequence. Missing any of these makes root cause analysis impossible.
- 04Post incident reviews must answer whether the launch gate would have caught the root cause. If not, the gate criteria must be updated before the next deployment.
Classic incident response playbooks assume a human attacker moving laterally through a network. AI incidents introduce a different pattern. An autonomous system making decisions faster than any human can observe, with authority over real resources, and with failure modes that can be triggered by input rather than by an attacker with credentials.
Security operations teams that have tried to apply existing playbooks to AI incidents report the same gap. The playbooks do not tell the analyst what to disable, what to collect, or how to determine whether the model was manipulated or malfunctioned. This article closes that gap.
The four AI incident categories and their first responder actions
/AI_INCIDENT_CATEGORIES
| Category | Symptoms | First Responder Action | Evidence to Preserve |
|---|---|---|---|
| Unauthorized tool invocation | Tool call log shows calls outside the expected scope, or to resources the feature should not access | Disable the agent feature flag. Do not restart. | Full tool call log with timestamps, model input at each step, tenant context |
| Tenant data crossing a boundary | Retrieval log shows content returned to an identity that does not own it | Revoke the retrieval service token. Notify affected tenants. | Retrieval query, returned chunks, calling identity, timestamp |
| Runaway token consumption | Per tenant token spend exceeds 3x the daily baseline | Trigger hard token budget kill switch for the affected tenant. | Conversation history, tool call loop sequence, entry point request |
| Model output used in downstream attack | Model response appears in a downstream SQL query, HTML injection, or command execution | Disable the integration that passes model output downstream. | Model response, downstream system logs, input that triggered the response |
Reconstruction workflow for AI incidents
Root cause analysis for an AI incident requires replaying the exact sequence of inputs the model received and the outputs it produced. This is only possible if the audit log captures the full context window, not just a summary. Many teams discover during an incident that their logging only preserved the final response. That is not sufficient.
The reconstruction workflow has five steps. First, identify the conversation ID or session ID for the affected interaction. Second, retrieve the complete input sequence from the audit log. Third, replay the sequence in a sandboxed environment using the same model version. Fourth, identify the specific input segment that caused the unexpected behavior. Fifth, determine whether that input could have been blocked at the launch gate.
Never replay suspicious inputs against a production model.
Replaying a suspected prompt injection against a live, model with tool access to confirm the finding will cause the same unauthorized actions again. Always replay in a sandboxed environment with all tools stubbed out.
Metrics that indicate IR readiness for AI incidents
- 01Mean time to disable. Time from first anomaly alert to the affected agent being disabled. Target is under 15 minutes.
- 02Audit log completeness rate. Percentage of AI incidents where a full input, output, and tool call sequence was recoverable from logs. Target is 100 percent.
- 03Post incident gate update rate. Percentage of AI incidents that produced an improvement to the launch gate criteria. Target is 100 percent.
- 04Tenant notification time. Time from confirmation of a tenant data boundary crossing to tenant notification. Target aligns with your breach notification SLA.
Building AI incident response capability before the first incident
The worst time to design an AI incident response capability is during an active incident. Security operations teams should run a tabletop exercise against each of the four incident categories before any AI feature reaches significant production traffic. The tabletop should test three things. Whether the team knows which feature flag or kill switch to activate, whether the audit log actually contains the data needed for reconstruction, and whether tenant notification procedures are defined.
The output of the tabletop is a one page runbook for each incident category. The runbook should be stored where on call analysts can find it in under 30 seconds. Teams that have done this work report that AI incidents, when they occur, resolve in the same time window as classic application security incidents. Teams that have not done this work report resolution times measured in days.
