AI Security: Prompt Injection, Data Leakage and Red TeamingDefense in depth and red teaming · Lesson 13 of 17

Defense-in-depth architecture for LLM apps and agents

Article · 15 min · 9 min lecture

Video lecture

Defense-in-depth architecture for LLM apps and agents

13 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 13

Defense in depth for LLM systems

  • Why layers
  • Reference architecture
  • Key patterns
  • Kill switches

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

No single control stops prompt injection

Every lesson so far points to the same conclusion: detection is probabilistic, models can be manipulated, and individual defenses get bypassed (EchoLeak chained several bypasses). The answer is defense in depth: independent layers such that an attacker must defeat all of them, and the architecture limits damage when they do.

A layered reference architecture

 Layer 0  Scope & design       -> minimal capabilities, trifecta check, threat model
 Layer 1  Identity & access    -> authenticated users, user-scoped credentials, tenant isolation
 Layer 2  Input handling       -> size limits, invisible-char scan, injection classifiers, provenance tags
 Layer 3  Context assembly     -> minimization, spotlighting untrusted content, trust tiers
 Layer 4  Model                -> current model, provider safety features, structured outputs
 Layer 5  Action control       -> narrow tools, authorization in tools, limits, human approval, sandboxes
 Layer 6  Output handling      -> validation, encoding, link/image allowlists, CSP, PII redaction
 Layer 7  Egress               -> network allowlists, recipient allowlists, no generic fetch on private-data agents
 Layer 8  Detection & response -> tracing, canaries, anomaly alerts, kill switches, incident runbooks

Layers 0, 1, 5 and 7 carry the most weight because they hold regardless of what the model does.

Architectural patterns worth knowing

Privilege separation. Split components by trust: a reader that processes untrusted content has no dangerous tools; an actor that has tools never sees raw untrusted content (dual LLM pattern). Research such as Design Patterns for Securing LLM Agents against Prompt Injections (2025) catalogs further patterns: action-selector (the model picks from fixed actions, and tool output never returns to the model), plan-then-execute, map-reduce over untrusted items in isolation, and context minimization.

Capability and data-flow control. Approaches like Google DeepMind researchers' CaMeL track where each value came from and enforce policies on where it may flow (for example, data from an untrusted email may not be used as an email recipient).

Allowlists over denylists. Allowlist tools per route, domains for fetch and links, recipients for email, file paths for reads, commands for shells. Denylists always miss something.

Human approval, done right. Approve high-impact actions with the actual parameters displayed, rate-limit approvals, and require re-authentication for the riskiest ones. Log who approved what.

Content provenance. Track source and trust tier for every context item; show sources in answers. For generated media, industry provenance standards such as C2PA Content Credentials can signal origin; they complement rather than replace access controls.

Guardrail services. Input and output classifiers and policy engines (provider-native services, open-source guardrail frameworks) add a useful probabilistic layer. Measure their false-positive and false-negative rates on your traffic.

Kill switches and safe degradation

Design the off switch before you need it:

  • Feature flags to disable a tool, a route or an entire agent instantly.
  • Circuit breakers that disable actions automatically when anomaly thresholds are crossed (for example, a spike in refunds or outbound emails).
  • Read-only mode for agents during investigations.
  • Credential revocation runbooks for every agent identity.
# circuit breaker for a high-impact tool (sketch)
def guarded_send_email(ctx, **kwargs):
    if flags.is_off("agent.send_email"):
        raise ToolDenied("email sending temporarily disabled")
    if metrics.count("agent.send_email", window_minutes=10) > limits["send_email_per_10m"]:
        flags.turn_off("agent.send_email", reason="rate anomaly")
        alerts.page("Agent email volume anomaly; tool disabled")
        raise ToolDenied("email sending paused pending review")
    return send_email(ctx, **kwargs)

Worked example: hardening an inbox agent

A consultancy in Dubai deployed an agent that triages a shared inbox, drafts replies and schedules meetings. Hardened design:

  • Layer 0: removed autonomous sending; drafts only. Calendar tool limited to proposing times.
  • Layer 1: OAuth tokens scoped to the shared mailbox; no access to personal mailboxes.
  • Layer 3: emails spotlighted as untrusted, with sender reputation as a provenance tag.
  • Layer 5: draft tool cannot add recipients beyond the original thread; attachments limited to files from the thread.
  • Layer 6: outputs rendered with link allowlisting and CSP in the review UI.
  • Layer 7: agent runtime has no general internet egress.
  • Layer 8: traces with canary tokens; alerts on unusual draft volume; one-click disable.

Red-team tests that previously produced exfiltration drafts now produced harmless drafts that humans rejected.

Pitfalls

  • Stacking probabilistic layers only (classifier + prompt + judge) with no deterministic boundary.
  • Approval fatigue that turns human review into a rubber stamp.
  • No kill switch, or one nobody has tested.
  • Assuming one pattern fits all routes. Apply the strictest pattern to the highest-risk flows.

How to measure success

For each high-risk flow you can name the deterministic control that stops exfiltration and damaging actions even if the model is fully compromised, and your kill switch has been tested in a drill.

Key takeaways

  • No single control stops prompt injection; build independent layers.
  • Deterministic layers (scope, identity, action control, egress) must hold even if the model is fully compromised.
  • Use privilege separation, action selectors, plan-then-execute, data-flow policies, allowlists and provenance.
  • Design and drill kill switches, circuit breakers and credential revocation before you need them.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which layer holds even if the model is fully manipulated?
  2. What does an 'action selector' pattern do?

Put it into practice

For your highest-risk flow, document the deterministic control that stops exfiltration and damaging actions, and run a kill-switch drill.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.