Skip to content

AI Security: Prompt Injection, Data Leakage and Red Teaming · Incident response and capstone · lesson 16 of 17 · 14 min

Incident response for AI systems

AI incidents are security incidents, with twists

When an LLM system misbehaves in a way that harms users, leaks data, costs money or damages trust, you need the same disciplined response as any security incident, plus some AI-specific steps. Existing frameworks (for example NIST's incident handling guidance, and the Manage function of the NIST AI RMF) apply; this lesson adapts them.

What counts as an AI security incident

  • Data exposure via model output, retrieval, logs or tools (including cross-tenant leakage).
  • Unauthorized actions by an agent (refunds, emails, deletions, purchases).
  • Successful prompt injection or jailbreak with material impact (harmful content published under your brand).
  • Poisoned knowledge base, memory or model artifact discovered.
  • Compromised MCP server, plugin, package or model supply chain.
  • Cost abuse (denial of wallet) or credential theft of AI API keys.
  • Hidden context exposure revealing sensitive business information.

Preparation (before anything happens)

  1. Logging and tracing sufficient to reconstruct what happened: prompts and outputs (with privacy controls), retrieved document IDs, tool calls with arguments, agent and user identities, model and prompt versions. Without this, investigations are guesswork.
  2. Kill switches per tool, route and agent; credential revocation procedures; read-only modes (previous module).
  3. Runbooks for the scenarios above, with owners and contact lists (security, engineering, legal, privacy, communications, vendor contacts).
  4. Severity matrix that includes AI-specific impacts (brand harm from generated content, regulated advice).
  5. Legal and regulatory map: breach-notification obligations differ by jurisdiction (for example UK GDPR generally requires notifying the regulator within 72 hours of becoming aware of a reportable personal-data breach; other regimes, including in the UAE and Saudi Arabia, set their own rules). Confirm with counsel in advance.

Response phases

Detect. Signals include canary-string hits, anomaly alerts (spikes in tool calls, refunds, outbound messages, token spend), PII detectors firing on outputs, user reports, social media posts, and red-team or bug-bounty reports.

Triage. Confirm it is real; classify type and severity; open an incident channel; assign an incident commander.

Contain. Flip kill switches; disable affected tools or routes; switch agents to read-only; revoke and rotate credentials (AI API keys, OAuth tokens, MCP server tokens); remove poisoned documents from indexes and memory; block malicious sources; pin or roll back model and prompt versions.

Investigate. Reconstruct timelines from traces: which inputs carried the attack, which tools were called, what data was exposed, to whom. Preserve evidence (logs, transcripts, index snapshots) with access controls. Determine scope: how many users, tenants, records.

Eradicate and recover. Fix root causes (architecture, permissions, validation), not only the specific payload. Re-index clean data; restore tools gradually behind feature flags; monitor closely.

Notify. Regulators, affected customers, partners and vendors as required and appropriate; coordinate wording with legal and communications. Be factual about what an AI system did and did not do.

Learn. Blameless post-incident review; add the attack to red-team and regression suites; update the threat model, runbooks and training. Consider sharing anonymized lessons with the community (for example via MITRE ATLAS case-study style write-ups or industry groups).

Hands-on: a runbook skeleton

# Runbook: Suspected data exfiltration via agent
Owner: Security on-call | Backup: AI platform lead
## Signals
- Canary token in outbound content; egress alert; PII detector on outputs; user report
## Immediate containment (first 30 minutes)
- [ ] Flip `agent.<name>.enabled=false` (or read-only)
- [ ] Revoke agent credentials: <list OAuth apps, API keys, MCP tokens>
- [ ] Block destination domains/addresses at egress
- [ ] Snapshot traces, index and memory for evidence (restricted bucket)
## Investigation
- [ ] Identify triggering inputs (trace IDs), affected users/tenants, data categories
- [ ] Check other agents/tools for the same pattern
## Notification decision
- [ ] Privacy + legal assess regulatory obligations and deadlines
## Recovery
- [ ] Root-cause fix merged + regression test added
- [ ] Staged re-enable with monitoring
## Post-incident
- [ ] Blameless review within 5 working days; update threat model and red-team suite

Worked example

A travel company in Dubai discovered through a customer tweet that its booking assistant was recommending an unfamiliar "partner" site for visa services. Traces showed a supplier's hotel description, recently updated, contained hidden instructions. Containment: the assistant's link output was restricted to the company domain within an hour; the supplier content was removed and the index re-built; affected conversations were identified from traces. Root-cause fixes: supplier content moved to a lower trust tier and sanitized at ingestion, link allowlisting made permanent, and a new red-team case added. Customers who had clicked were contacted with guidance.

Pitfalls

  • No traces, so scope cannot be determined.
  • Fixing only the payload (deleting one document) rather than the class of weakness.
  • Forgetting credentials held by agents and MCP servers during containment.
  • Improvised notifications without legal input.

How to measure success

Runbooks exist for each AI incident type, kill switches and credential revocation are drilled, mean time to contain is measured in drills, and every incident produces regression tests and threat-model updates.

Video lecture: Incident response for AI systems

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

  1. Incident response for AI systems
  2. Why a dedicated playbook
  3. Analogy: aviation safety
  4. AI incident types
  5. Preparation
  6. Detect → Triage → Contain
  7. Investigate → Recover → Notify → Learn
  8. Example: canary in an outbound draft
  9. Case: Dubai travel assistant
  10. Common mistakes
  11. Deeper: handling evidence
  12. Watch me do it: runbook + tabletop
  13. Recap

Lecture transcript

Incident response for AI systems

It is two in the morning, and a customer has just tweeted a screenshot: your AI assistant is recommending a website nobody at your company has heard of, for visa services. Is this a bug? A prank? An attack? Who do you call, and what do you switch off first? In this lecture you will learn how to prepare for and respond to AI security incidents, adapting proven incident response to the specific ways LLM systems fail.

Why a dedicated playbook

Why does this need its own lesson? Because AI incidents are security incidents, with twists. The attack may arrive through a hotel description, not a network port. The harmful action may be a sentence your brand published, not a stolen file. And the evidence lives in prompts, retrieved documents and tool calls, which many teams never log. Existing guidance, like NIST's incident handling work and the Manage function of the AI Risk Management Framework, still applies. You just have to adapt it.

Analogy: aviation safety

Here is an analogy. Think of an airline's response to a problem with an aircraft. They do not wait for an accident to decide who is in charge. There are checklists for each scenario, flight recorders that capture exactly what happened, the ability to ground a model of plane immediately, and a review afterwards that changes procedures for everyone. Your traces are the flight recorder. Your kill switches ground the plane. Your runbooks are the checklists.

AI incident types

First, what counts as an AI security incident? Data exposure through outputs, retrieval, logs or tools, including cross-tenant leaks. Unauthorized agent actions, like refunds, emails or deletions. A successful injection or jailbreak with real impact, such as harmful content published under your brand. A poisoned knowledge base, memory or model file. A compromised MCP server, plugin or package. Cost abuse or stolen AI API keys. And hidden context exposure that reveals sensitive business information.

Preparation

Preparation decides how well you respond. You need logging and tracing good enough to reconstruct what happened: prompts and outputs with privacy controls, retrieved document identifiers, tool calls and arguments, agent and user identities, and model and prompt versions. You need kill switches and credential revocation procedures. Runbooks with owners and contact lists across security, engineering, legal, privacy and communications. A severity matrix that includes AI harms. And a legal map of notification duties. For example, UK GDPR generally requires notifying the regulator within seventy-two hours of becoming aware of a reportable personal data breach, and other regimes, including in the UAE and Saudi Arabia, set their own rules. Confirm with counsel before you need it.

Detect → Triage → Contain

Now the response phases. Detect, from canary hits, anomaly alerts on tool calls or spend, personal data detectors, user reports and social media. Triage: confirm it is real, classify type and severity, open an incident channel and name an incident commander. Contain: flip kill switches, disable tools or routes, move agents to read-only, revoke and rotate API keys, OAuth tokens and MCP tokens, remove poisoned documents, block malicious sources, and pin or roll back model and prompt versions.

Investigate → Recover → Notify → Learn

Then investigate: reconstruct the timeline from traces, identify which inputs carried the attack, which tools were called, and what data went where. Preserve evidence with access controls, and determine scope: how many users, tenants and records. Eradicate and recover by fixing the root cause, not just the payload, then re-enable gradually behind feature flags with close monitoring. Notify regulators, customers and partners as required, with legal and communications. And learn, with a blameless review that adds the attack to your red-team suite and updates the threat model.

Example: canary in an outbound draft

Let's walk a simple example. A canary string planted in an agent's hidden context appears in an outbound email draft. That is your detection signal. Within thirty minutes, following the runbook in the lesson, you switch the agent to read-only, revoke its mail token, block the destination domain at egress, and snapshot traces and memory for evidence. Then you investigate which inbound message carried the instruction, check other agents for the same pattern, and ask privacy and legal whether notification is required.

Case: Dubai travel assistant

Now the realistic case from the opening. A travel company in Dubai learned from a customer's tweet that its booking assistant recommended an unfamiliar partner site for visas. Traces showed a supplier's recently updated hotel description contained hidden instructions. Within an hour, link output was restricted to the company's own domain. The supplier content was removed and the index rebuilt, and affected conversations were identified from traces. The root-cause fixes: supplier content moved to a lower trust tier and sanitized at ingestion, link allowlisting made permanent, and a new red-team case added. Customers who had clicked were contacted with guidance.

Common mistakes

Common mistakes. Having no traces, so you cannot tell who was affected. Fixing only the payload, deleting one document, instead of the class of weakness that let it in. Forgetting the credentials held by agents and MCP servers when you contain. And improvising customer or regulator notifications without legal input. Ask yourself: if this happened tonight, could you name every credential your agent holds?

Deeper: handling evidence

One level deeper on evidence handling. Snapshots go into a bucket that only the incident team can read, access is logged, and the retention period is agreed with legal. Transcripts with personal data are redacted in the working copy used for analysis, while the original is preserved untouched in case regulators or courts need it.

Watch me do it: runbook + tabletop

Watch me do it: filling in the runbook skeleton for our support agent, then running a tabletop. Owner: security on-call; backup: the AI platform lead. Signals: canary token in outbound content, an egress alert, a PII detector hit on outputs, or a customer report. Immediate containment, first thirty minutes. Flip the agent's enabled flag to off, or to read-only; I write the exact flag name. Revoke credentials; here is where the real work is. I list them: the provider API key for this agent, the OAuth app for the shared mailbox, the CRM service token, and the token our tracking MCP server uses. Four credentials, each with the console where it is revoked. Block destinations at egress. Snapshot traces, the vector index and memory into the restricted evidence bucket. Investigation: find triggering inputs by trace ID, list affected customers and data categories, and check other agents for the same pattern. Notification decision: privacy and legal assess obligations and deadlines, including the seventy-two-hour window under UK GDPR where it applies. Recovery: root-cause fix plus a regression test, then a staged re-enable. Now the tabletop. I read a scenario aloud: the canary appeared in a draft email at two in the morning. The team talks through each step. We discover nobody knew who could revoke the CRM token at night. That becomes action item one.

Recap

Recap. AI incidents need the same discipline as any security incident, plus AI-specific preparation: rich traces, kill switches, credential inventories, runbooks and a legal map. Respond through detect, triage, contain, investigate, recover, notify and learn, and always fix the class of weakness. Try this now: copy the runbook skeleton from the lesson for your highest-risk agent, fill in real credential names and owners, and schedule a thirty-minute tabletop exercise with your team.

Key takeaways

  • AI incidents include data exposure, unauthorized agent actions, impactful injections, poisoning, supply-chain compromise, cost abuse and hidden context exposure.
  • Prepare traces that reconstruct events, kill switches, credential inventories, runbooks, a severity matrix and a legal notification map.
  • Respond through detect, triage, contain, investigate, recover, notify and learn.
  • Fix the class of weakness, not just the payload, and turn every incident into regression tests and threat-model updates.

Try it

Fill in the runbook skeleton for your highest-risk agent with real credential names and owners, and run a 30-minute tabletop exercise.