The AI Incident Playbook: What to Do When Your AI Agent Goes Off-Script in Production

The AI Incident Playbook: What to Do When Your AI Agent Goes Off-Script in Production

Cyril Treacy

COO and Co-Founder

This post explains how to respond when an AI agent goes off-script in production, the four incident types, the warning signs that precede them, and what regulators expect in the AI Evidence trail.

This post explains how to respond when an AI agent goes off-script in production, the four incident types, the warning signs that precede them, and what regulators expect in the AI Evidence trail.

Key Takeaways

  • Drift is not a bug. It is a governance failure. The agent is doing what the policies allowed, not what you intended.

  • The four agent incident types are policy drift, prompt injection, cascading failure, and data exfiltration.

  • Only 1 in 5 enterprises has mature governance for autonomous agents, while around 80% have seen risky agent behaviour in production.

  • The playbook runs in five steps: detect, contain, assess, attribute, remediate and review.

  • Regulators read what your governance system did about the incident, not just what the incident was.

Key Takeaways

  • Drift is not a bug. It is a governance failure. The agent is doing what the policies allowed, not what you intended.

  • The four agent incident types are policy drift, prompt injection, cascading failure, and data exfiltration.

  • Only 1 in 5 enterprises has mature governance for autonomous agents, while around 80% have seen risky agent behaviour in production.

  • The playbook runs in five steps: detect, contain, assess, attribute, remediate and review.

  • Regulators read what your governance system did about the incident, not just what the incident was.

Why AI agent incidents are different

Traditional incident response assumes deterministic systems. Same input, same output. Agents do not behave that way. Same instruction, different output depending on context, model state, tool availability, and the policies in force.

Drift is not a bug. It is a governance failure. The agent is operating under the policies it was given. The behaviour is a statement about those policies, not the model.

This is what PowerPoint Governance looks like live. The policy register is current. The committee minutes are signed. The agent is still off-script. Deloitte's 2026 research puts only 1 in 5 enterprises at a mature governance model for autonomous agents, while around 80% have seen risky agent behaviour in production. The gap is policy versus runtime.

The four types of AI agent incident

Most agent incidents fall into one of four categories.

  1. Policy drift. Behaviour gradually moves outside the boundaries set at deployment. No single decision crossed the line. The aggregate did.

  1. Prompt injection. External inputs manipulate behaviour through malicious instructions hidden in retrieved content, tool responses, or user prompts. The agent follows because the governance layer never separated trusted from untrusted input.

  1. Cascading failure. One agent's output triggers downstream actions. A single bad decision propagates before anyone reads the first log line. This is the failure mode dismissed as Agentic Theatre in the demo and surfaces hard in live orchestration.

  1. Data exfiltration. The agent accesses or transmits data outside its permitted scope, through tool misuse or retrieval that pulled in records it should never have seen.

Category alone is enough to triage. Containment and attribution differ for each.

The warning signs that precede most incidents

Most incidents show up in telemetry before they show up in outputs. Four signals carry the weight.

  1. Unusual tool-call patterns. APIs called outside expected scope, or at frequencies outside the baseline.

  1. Output confidence degradation without input change. The same class of input produces less confident outputs over time.

  1. Unexpected inter-agent communication. Agents not designed to coordinate are coordinating. Volume spikes against the baseline.

  1. User escalation. End users escalating outputs they previously accepted.

No signal is conclusive alone. Together they form the early-warning layer. A structured monitoring approach reads all four.

The AI incident response playbook

Five steps. Run in order.

  1. Detect. Real-time monitoring triggers the alert. A threshold breach against the baseline, not a user complaint.

  1. Contain. Isolate the agent without destroying state. In software the instinct is to kill the process. In an agent incident the instinct is to preserve the decision chain. Cut the agent's ability to act while keeping the record of what it was about to do.

  1. Assess. Read the AI Evidence trail. What the agent saw, what policies were in force, what action it took, why it passed the controls. A summary log is not enough.

  1. Attribute. Decide whether the cause is model drift, prompt injection, a policy gap, or a data issue. The four types map to four remediation paths.

  1. Remediate and review. Update the policy, reconfigure the model, re-test before redeployment, update monitoring thresholds. A rollback alone is not remediation. The conditions that allowed the incident remain unless policy, tests, or triggers change.

What regulators expect after the incident

The supervisor question is not what happened. It is what the governance system did about it.

The EU AI Act, Article 26 requires deployers of high-risk AI systems to monitor operation, follow the provider's instructions, and report serious incidents to market surveillance authorities. Reporting is mandatory.

The FCA reads incident response under operational resilience and SMCR. Firms have to show they identified, contained, attributed, and remediated, with controls against recurrence. The evidence is structural: who owned the agent, what policy was in force, what the system did when the threshold breached.

The NIST AI Risk Management Framework sets the manage function as continuous documented risk treatment.

The common ground is the AI Evidence requirement. (A document repository with the policy register attached does not produce that record.)

How Disseqt covers incident response end-to-end

Disseqt is the only assurance layer built for the full enterprise AI lifecycle, unified in one platform. Test & Detect. Protect & Enforce. Prove & Comply. One spine, three pillars, every stage of the incident covered.

  1. Test & Detect. Continuous adversarial testing, prompt injection scenarios, and vulnerability detection before deployment. Find the failure modes before they land in production.

  1. Protect & Enforce. Run-time protection at the inference layer, inline policy enforcement on every agent decision, continuous monitoring while the AI is live. Containment without losing the decision chain.

  1. Prove & Comply. Automated compliance reporting with AI Evidence at the step level. Audit-ready records mapped to Article 26, SMCR, and NIST, with SOC2, SSO/SCIM, and RBAC. What supervisors actually read.

One Window for the Full AI Assurance Lifecycle. One data model. One platform. End-to-end.

Bottom Line

Agent incidents are not software incidents. The response model has to account for non-determinism, the difference between drift and defect, and the supervisor cycle. Firms running the full AI Assurance Lifecycle on one platform have a clean AI Evidence record when the incident lands. Firms running PowerPoint Governance reconstruct what was never captured.

When the supervisor calls, who reads the file. If the answer is the team that owns the policy register, the file is empty. See the AI Assurance Layer.

FAQs

01

What is AI agent drift?

AI agent drift is the gradual movement of an agent's behaviour outside the boundaries set at deployment. No single decision crosses the line, the aggregate does. Drift is a governance failure, not a model bug.

02

How do you detect when an AI agent goes off-script?

Four signals: unusual tool-call patterns, output confidence degradation without input change, unexpected inter-agent communication, and user escalation patterns. Read all four against the deployment baseline.

03

What should an AI incident response playbook include?

Five steps: detect, contain, assess, attribute, remediate and review. Detect uses real-time triggers. Contain preserves the decision chain. Assess reads the AI Evidence trail. Attribute maps to one of four incident types. Remediate and review update policy, tests, and thresholds.

04

What do regulators require after an AI agent incident?

EU AI Act Article 26 requires deployers to report serious incidents. The FCA reads incident response under operational resilience and SMCR. The NIST AI RMF sets continuous documented risk treatment. The common requirement is the AI Evidence trail.

AUTHOR

Cyril Treacy

COO and Co-Founder

Cyril is Co-Founder and COO at Disseqt, leading go-to-market, partnerships, and customer success. He brings 20+ years of enterprise sales, pre-sales leadership, and scaling expertise from Salesforce and the Irish startup ecosystem.

See Disseqt in action
Book a 30-minute walkthrough

Our team will walk you through a live workflow using your own AI environment. No slides. No generic demo. A real walkthrough of how Disseqt fits into your stack.

See Disseqt in action
Book a 30-minute walkthrough

Our team will walk you through a live workflow using your own AI environment. No slides. No generic demo. A real walkthrough of how Disseqt fits into your stack.

See Disseqt in action
Book a 30-minute walkthrough

Our team will walk you through a live workflow using your own AI environment. No slides. No generic demo. A real walkthrough of how Disseqt fits into your stack.