---
title: "Defense-in-depth architecture for LLM apps and agents"
description: "No single control stops prompt injection Every lesson so far points to the same conclusion: detection is probabilistic, models can be manipulated, and…"
url: https://optimizeall.com/learn/ai-security-and-red-teaming/defense-in-depth-architecture
updated: 2026-10-05
---

AI Security: Prompt Injection, Data Leakage and Red Teaming · Defense in depth and red teaming · lesson 13 of 17 · 15 min

# Defense-in-depth architecture for LLM apps and agents

## No single control stops prompt injection

Every lesson so far points to the same conclusion: detection is probabilistic, models can be manipulated, and individual defenses get bypassed (EchoLeak chained several bypasses). The answer is **defense in depth**: independent layers such that an attacker must defeat all of them, and the architecture limits damage when they do.

## A layered reference architecture

```text
 Layer 0  Scope & design       -> minimal capabilities, trifecta check, threat model
 Layer 1  Identity & access    -> authenticated users, user-scoped credentials, tenant isolation
 Layer 2  Input handling       -> size limits, invisible-char scan, injection classifiers, provenance tags
 Layer 3  Context assembly     -> minimization, spotlighting untrusted content, trust tiers
 Layer 4  Model                -> current model, provider safety features, structured outputs
 Layer 5  Action control       -> narrow tools, authorization in tools, limits, human approval, sandboxes
 Layer 6  Output handling      -> validation, encoding, link/image allowlists, CSP, PII redaction
 Layer 7  Egress               -> network allowlists, recipient allowlists, no generic fetch on private-data agents
 Layer 8  Detection & response -> tracing, canaries, anomaly alerts, kill switches, incident runbooks
```

Layers 0, 1, 5 and 7 carry the most weight because they hold regardless of what the model does.

## Architectural patterns worth knowing

**Privilege separation.** Split components by trust: a reader that processes untrusted content has no dangerous tools; an actor that has tools never sees raw untrusted content (dual LLM pattern). Research such as *Design Patterns for Securing LLM Agents against Prompt Injections* (2025) catalogs further patterns: action-selector (the model picks from fixed actions, and tool output never returns to the model), plan-then-execute, map-reduce over untrusted items in isolation, and context minimization.

**Capability and data-flow control.** Approaches like Google DeepMind researchers' CaMeL track where each value came from and enforce policies on where it may flow (for example, data from an untrusted email may not be used as an email recipient).

**Allowlists over denylists.** Allowlist tools per route, domains for fetch and links, recipients for email, file paths for reads, commands for shells. Denylists always miss something.

**Human approval, done right.** Approve high-impact actions with the actual parameters displayed, rate-limit approvals, and require re-authentication for the riskiest ones. Log who approved what.

**Content provenance.** Track source and trust tier for every context item; show sources in answers. For generated media, industry provenance standards such as C2PA Content Credentials can signal origin; they complement rather than replace access controls.

**Guardrail services.** Input and output classifiers and policy engines (provider-native services, open-source guardrail frameworks) add a useful probabilistic layer. Measure their false-positive and false-negative rates on your traffic.

## Kill switches and safe degradation

Design the off switch before you need it:

- **Feature flags** to disable a tool, a route or an entire agent instantly.
- **Circuit breakers** that disable actions automatically when anomaly thresholds are crossed (for example, a spike in refunds or outbound emails).
- **Read-only mode** for agents during investigations.
- **Credential revocation** runbooks for every agent identity.

```python
# circuit breaker for a high-impact tool (sketch)
def guarded_send_email(ctx, **kwargs):
    if flags.is_off("agent.send_email"):
        raise ToolDenied("email sending temporarily disabled")
    if metrics.count("agent.send_email", window_minutes=10) > limits["send_email_per_10m"]:
        flags.turn_off("agent.send_email", reason="rate anomaly")
        alerts.page("Agent email volume anomaly; tool disabled")
        raise ToolDenied("email sending paused pending review")
    return send_email(ctx, **kwargs)
```

## Worked example: hardening an inbox agent

A consultancy in Dubai deployed an agent that triages a shared inbox, drafts replies and schedules meetings. Hardened design:

- **Layer 0:** removed autonomous sending; drafts only. Calendar tool limited to proposing times.
- **Layer 1:** OAuth tokens scoped to the shared mailbox; no access to personal mailboxes.
- **Layer 3:** emails spotlighted as untrusted, with sender reputation as a provenance tag.
- **Layer 5:** draft tool cannot add recipients beyond the original thread; attachments limited to files from the thread.
- **Layer 6:** outputs rendered with link allowlisting and CSP in the review UI.
- **Layer 7:** agent runtime has no general internet egress.
- **Layer 8:** traces with canary tokens; alerts on unusual draft volume; one-click disable.

Red-team tests that previously produced exfiltration drafts now produced harmless drafts that humans rejected.

## Pitfalls

- **Stacking probabilistic layers only** (classifier + prompt + judge) with no deterministic boundary.
- **Approval fatigue** that turns human review into a rubber stamp.
- **No kill switch**, or one nobody has tested.
- **Assuming one pattern fits all routes.** Apply the strictest pattern to the highest-risk flows.

## How to measure success

For each high-risk flow you can name the deterministic control that stops exfiltration and damaging actions even if the model is fully compromised, and your kill switch has been tested in a drill.

## Video lecture: Defense-in-depth architecture for LLM apps and agents

Lecture coming soon · 13 chapters · about 9 minutes. Read the full transcript below.

1. Defense in depth for LLM systems
2. Analogy: airport security
3. Layers 0–4
4. Layers 5–8
5. Patterns
6. Principles
7. Kill switches
8. Case: Dubai inbox agent
9. Common mistakes
10. Walkthrough: the phishing review
11. Deeper: an action selector
12. Watch me do it: circuit-breaker drill
13. Recap

## Lecture transcript

### Defense in depth for LLM systems

If you remember one thing from this course, make it this: no single control stops prompt injection. Classifiers miss things. Models get manipulated. Filters get bypassed, sometimes all in the same attack, as EchoLeak showed. So professional teams build defense in depth, independent layers where an attacker must beat every one, and where the architecture limits the damage when they do. In this lecture you will get a layered reference architecture, the key patterns, and a way to switch things off safely when something goes wrong.

### Analogy: airport security

Think of it like airport security. There is not one check but many: identity at booking, bag scanning, a metal detector, a boarding pass check, locked cockpit doors. Each has gaps. Together they make it very hard to cause harm. And the most important one, the locked cockpit door, does not care how convincing the passenger is. In LLM systems, your locked cockpit doors are the deterministic controls: scoped permissions, tool authorization, egress blocks. They hold no matter what the model decides.

### Layers 0–4

Here is the reference architecture in the lesson, as nine layers. Layer zero, scope and design: minimal capabilities and the trifecta check. Layer one, identity and access: authenticated users, user-scoped credentials, tenant isolation. Layer two, input handling: size limits, invisible character scanning, injection classifiers and provenance tags. Layer three, context assembly: minimization, spotlighting and trust tiers. Layer four, the model itself, with provider safety features and structured outputs.

### Layers 5–8

Layer five, action control: narrow tools, authorization inside tools, hard limits, human approval and sandboxes. Layer six, output handling: validation, encoding, link and image allowlists, content security policy and redaction. Layer seven, egress: network and recipient allowlists, and no generic fetch tools on agents with private data. Layer eight, detection and response: tracing, canaries, anomaly alerts, kill switches and runbooks. Notice which layers carry the most weight: zero, one, five and seven, because they hold regardless of what the model does.

### Patterns

Now the patterns. Privilege separation: the component that reads untrusted content has no dangerous tools, and the component with tools never sees raw untrusted content. A twenty twenty-five research paper on design patterns for securing agents catalogs more: an action selector, where the model picks from fixed actions and tool output never comes back to it; plan-then-execute; map-reduce, processing untrusted items in isolation; and context minimization. CaMeL, from Google DeepMind researchers, tracks where every value came from and enforces policies on where it may flow.

### Principles

Three more principles. Prefer allowlists over denylists, for tools, domains, recipients, file paths and commands, because denylists always miss something. Do human approval properly: show the actual parameters, rate-limit approvals and require re-authentication for the riskiest actions. And track provenance for every piece of context, showing sources in answers. For generated media, standards like C2PA Content Credentials can signal origin, but they complement access controls rather than replacing them.

### Kill switches

Now the off switch, which you should design before you need it. Feature flags that disable a tool, a route or an entire agent instantly. Circuit breakers that disable actions automatically when anomaly thresholds are crossed, like a spike in outbound emails or refunds. A read-only mode for investigations. And credential revocation runbooks for every agent identity. The lesson includes a short circuit breaker around an email tool that pages on-call and pauses sending when volume looks abnormal.

### Case: Dubai inbox agent

A realistic example. A consultancy in Dubai deployed an agent that triages a shared inbox, drafts replies and schedules meetings. They removed autonomous sending entirely; drafts only. OAuth tokens were scoped to the shared mailbox. Emails were spotlighted as untrusted with sender reputation as a provenance tag. The draft tool could not add recipients beyond the original thread. Outputs were rendered with link allowlists. The runtime had no general internet access. And traces carried canary tokens, with alerts on unusual draft volume and a one-click disable. Red-team tests that once produced exfiltration drafts now produced harmless drafts that humans simply rejected.

### Common mistakes

Common mistakes. Stacking only probabilistic layers, a classifier plus a prompt plus a judge, with no deterministic boundary underneath. Approval fatigue, where human review becomes a rubber stamp. A kill switch that nobody has ever tested. And applying one pattern everywhere, instead of the strictest pattern to the highest-risk flows. Here is a question to test your own system: if the model were fully compromised right now, which deterministic control would stop data leaving and money moving?

### Walkthrough: the phishing review

Let's walk a simple example through the layers. A user asks a shopping assistant to summarize reviews, and one review contains an injection telling the model to show a phishing link with a fake discount. Layer two, the input scanner, may or may not flag it. Layer three spotlights reviews as untrusted, which lowers the odds of success. Suppose the model is fooled anyway, at layer four. Layer six, the output link allowlist, strips the link because the domain is not yours. And layer eight logs the attempt and alerts, so the review is removed. The attack needed every layer to fail. It only took one to stop it.

### Deeper: an action selector

One level deeper on the action-selector pattern. For an order assistant, the model chooses one of five fixed actions: show status, show delivery estimate, start a return, change address, or hand off to a human. The chosen action's parameters come from the verified session. Tool output is rendered by templates straight to the user and never returns to the model, so injected text in tool data cannot steer a next step.

### Watch me do it: circuit-breaker drill

Watch me do it: testing the circuit breaker from the lesson in a staging drill. The guarded send email function does three things before sending. First, it checks a feature flag for the email tool; if the flag is off, it refuses. Second, it counts emails sent by the agent in the last ten minutes and compares with the configured limit, say thirty. Third, if the count is over, it turns the flag off itself, pages on-call with the reason, and refuses. Now the drill. I write a script that simulates an injected agent trying to send a report to many addresses in a loop. Emails one to thirty go out to our test inbox. Email thirty-one trips the breaker. The flag flips off, the pager fires with rate anomaly, and every later attempt fails with email sending paused. I time it: the breaker tripped within seconds of crossing the limit. Then the second half of the drill, the human part. On-call opens the runbook, checks the traces for the agent's recent inputs, finds the planted instruction, and confirms no real customer emails were involved. They switch the agent to read-only, and only after review turn the email flag back on. Total drill time: twenty-two minutes. We write down two improvements, including a pre-filled trace query, and schedule the next drill.

### Recap

Recap. Assume every single control can fail. Build nine layers, and make sure the deterministic ones, scope, identity, action control and egress, hold even if the model is fully compromised. Use privilege separation and allowlists, do approvals properly, track provenance, and design your kill switches in advance. Try this now: for your highest-risk flow, write down the deterministic control that stops exfiltration and damaging actions, and schedule a drill to test your kill switch.

## Key takeaways

- No single control stops prompt injection; build independent layers.
- Deterministic layers (scope, identity, action control, egress) must hold even if the model is fully compromised.
- Use privilege separation, action selectors, plan-then-execute, data-flow policies, allowlists and provenance.
- Design and drill kill switches, circuit breakers and credential revocation before you need them.

## Try it

For your highest-risk flow, document the deterministic control that stops exfiltration and damaging actions, and run a kill-switch drill.

- [Previous: Secrets, PII, hidden context and unbounded consumption](https://optimizeall.com/learn/ai-security-and-red-teaming/secrets-pii-hidden-context-and-cost)
- [Next: Planning and running AI red-team engagements](https://optimizeall.com/learn/ai-security-and-red-teaming/red-teaming-methods)
- [All lessons of AI Security: Prompt Injection, Data Leakage and Red Teaming](https://optimizeall.com/learn/ai-security-and-red-teaming)
