AI Security: Prompt Injection, Data Leakage and Red Teaming · Threat modeling LLM applications and agents · lesson 2 of 17 · 14 min
A practical threat-modeling process for LLM systems
From sketch to threat model
A threat model answers four questions (Adam Shostack's widely used framing): What are we working on? What can go wrong? What are we going to do about it? Did we do a good job? For LLM systems, the process is the same as for any software, with AI-specific threat categories added.
Step 1: model the system
Draw a data flow diagram (DFD) with processes (orchestrator, model, tools), data stores (vector DB, memory, logs), external entities (users, third-party APIs, content sources) and trust boundaries. Include the prompt assembly step explicitly: what goes into the context window, from where, in what order.
Step 2: enumerate threats
Use two lenses together:
Classic STRIDE per element (Spoofing, Tampering, Repudiation, Information disclosure, Denial of service, Elevation of privilege) still applies: API keys stolen, logs tampered, cost-exhaustion DoS, privilege escalation via tools.
AI-specific categories, mapped to OWASP and MITRE ATLAS (next module):
| AI threat | Example question for your system | |---|---| | Direct prompt injection / jailbreak | Can a user make it ignore policy or reveal hidden context? | | Indirect prompt injection | Which retrieved or tool content could carry instructions? | | Exfiltration | Which channels can carry data out (links, images, tools, email)? | | Excessive agency | Which tools have more permission than the task needs? | | Data and model poisoning | Who can write to training data, fine-tuning sets, the RAG index or memory? | | Supply chain | Which models, datasets, packages, MCP servers do we trust, and why? | | Improper output handling | Where is model output rendered, executed or passed to another system? | | Sensitive information / hidden context exposure | What secrets or private data sit in prompts, context or logs? | | Unbounded consumption | Can someone drive unlimited tokens, tool calls or loops? | | Misinformation | Where do people act on outputs without verification? |
Step 3: write abuse cases
Turn threats into concrete, testable stories:
AC-07 Indirect injection via product reviews
As an attacker, I post a product review containing hidden instructions so that when a
shopper asks the assistant to "summarize reviews", it tells them to visit my phishing site
with a discount code.
Preconditions: reviews are retrieved into context unfiltered; assistant can output links.
Impact: customer phishing, brand damage. Likelihood: medium. Severity: high.
Mitigations: link allowlist on output, strip/neutralize instructions in retrieved text,
label retrieved content as untrusted, monitor for off-domain URLs.
Test: red-team case RT-07 (Module 6).
Abuse cases become red-team test cases and regression tests.
Step 4: rate and prioritize
Use a simple likelihood × impact matrix, or a scoring scheme your organization already uses. Rate impact by worst realistic outcome given current tool permissions, not by what the model "usually" does. A tool that can send money has high impact even if injection success is rare.
Step 5: decide mitigations and owners
For each high-priority threat, choose controls (Module 6 covers defense in depth): remove capability, restrict scope, require approval, isolate content, filter output, monitor. Assign an owner and a test.
Step 6: validate and revisit
Red-team the mitigations. Revisit the model when you add a tool, a data source, memory, a new model, or a new user group. Agents change fast; so should their threat models.
Worked example: a WhatsApp commerce agent
A Lahore-based D2C brand runs an agent on WhatsApp Business that answers product questions, checks order status and can create discount codes up to 10% for unhappy customers.
- Untrusted inputs: customer messages (any language, voice notes transcribed), product catalog text (supplier-provided), order notes.
- Actions:
get_order(order_id),create_discount(percent). - Top abuse cases: (1) customer talks the agent into creating many discount codes; (2) customer requests another person's order status by guessing IDs; (3) supplier catalog text contains instructions; (4) cost exhaustion via long voice notes.
- Mitigations chosen: discount tool enforces max percent and one code per phone number per 30 days in code;
get_orderrequires the phone number on the order to match the WhatsApp sender; catalog text sanitized on ingestion and marked as data; voice-note length limits and per-user rate limits.
Note that each mitigation lives in code or configuration, not in the prompt.
Hands-on: a threat-model worksheet
| ID | Element / flow | Threat (STRIDE or AI category) | Abuse case | Likelihood | Impact | Mitigation | Owner | Test ID |
|----|----------------|--------------------------------|------------|------------|--------|------------|-------|---------|
Aim for 10–20 rows for a typical feature in a first pass, then deepen the top five.
Pitfalls
- Threat modeling the model instead of the system. The risk is in the tools, data and integrations.
- Rating by typical behavior. Attackers are not typical users.
- One-off documents. A threat model that is not revisited when tools change is fiction.
How to measure success
Every high-impact threat has a mitigation implemented in code or configuration, an owner, and a red-team test that passes.
Video lecture: A practical threat-modeling process for LLM systems
Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.
- Threat modeling LLM apps
- Analogy: the fire engineer's walk-through
- Four questions
- 1. Model the system
- 2. Enumerate threats
- 3. Abuse cases
- 4–6. Prioritize, mitigate, revisit
- Case: WhatsApp commerce agent
- Mitigations in code, not prompts
- Drill: meeting-notes assistant
- Common mistakes
- Deeper: ranking two threats
- Watch me do it: one abuse case, end to end
- Recap
Lecture transcript
Threat modeling LLM apps
A threat model is a structured way of asking: what could go wrong, and what will we do about it, before an attacker asks for us. In this lecture you will learn a six-step process for threat modeling LLM applications and agents: model the system, enumerate threats with classic and AI-specific lenses, write abuse cases, prioritize, mitigate and revisit. You will leave with a worksheet you can use this week.
Analogy: the fire engineer's walk-through
An analogy for threat modeling. Before an architect finalizes a building, a fire engineer walks through the drawings and asks: where could a fire start, how would it spread, how do people get out, and what stops it? They do this on paper, when changes are cheap, not after the walls are up. Threat modeling is that walk-through for your AI system. You ask where an attack could start, how it would spread through tools and data, and what stops it, while changing the design is still cheap.
Four questions
Adam Shostack's widely used framing boils threat modeling down to four questions. What are we working on? What can go wrong? What are we going to do about it? Did we do a good job? For LLM systems the process is the same as for any software. You simply add AI-specific threat categories to the list of what can go wrong.
1. Model the system
Step one: model the system as a data flow diagram. Processes like the orchestrator, the model and tools. Data stores like the vector database, memory and logs. External entities like users, third-party APIs and content sources. And trust boundaries. Crucially, include prompt assembly as an explicit step: what goes into the context window, from where, and in what order. That is where untrusted content mixes with your instructions.
2. Enumerate threats
Step two: enumerate threats with two lenses. The classic STRIDE lens still applies: spoofing, tampering, repudiation, information disclosure, denial of service and elevation of privilege. Stolen keys, tampered logs, cost exhaustion and privilege escalation through tools. Then the AI-specific lens: direct and indirect prompt injection, exfiltration channels, excessive agency, poisoning, supply chain, improper output handling, hidden context exposure, unbounded consumption and misinformation.
3. Abuse cases
Step three: turn threats into abuse cases, concrete stories you can test. For example: as an attacker, I post a product review containing hidden instructions, so that when a shopper asks the assistant to summarize reviews, it sends them to my phishing site. Write down preconditions, impact, likelihood, mitigations and a test identifier. Every abuse case becomes a red-team case and later a regression test.
4–6. Prioritize, mitigate, revisit
Step four: prioritize with likelihood and impact. The key rule is to rate impact by the worst realistic outcome given the tool permissions you actually granted, not by what the model usually does. A tool that can send money has high impact even if injection rarely succeeds. Step five: choose mitigations and owners. Remove the capability, restrict its scope, require approval, isolate content, filter output, or monitor. Step six: validate with red teaming, and revisit whenever you add a tool, data source, memory or model.
Case: WhatsApp commerce agent
A worked example. A direct-to-consumer brand in Lahore runs a WhatsApp agent that answers product questions, checks order status and can create discount codes up to ten percent. Abuse cases: customers talking the agent into many discount codes, guessing other people's order numbers, supplier catalog text containing instructions, and cost exhaustion through long voice notes.
Mitigations in code, not prompts
Look at the mitigations they chose. The discount tool enforces the maximum percentage and one code per phone number per month, in code. Order lookup requires the phone number on the order to match the WhatsApp sender. Catalog text is sanitized on ingestion and marked as data. Voice notes have length limits and each user has rate limits. Notice that every single mitigation lives in code or configuration, not in the prompt.
Drill: meeting-notes assistant
Let's do a quick simple example together. Feature: a meeting-notes assistant that reads calendar invites and writes summaries into a shared drive. Ask the four questions. What are we working on: invites in, summaries out. What can go wrong: an invite from an outsider could carry an instruction; summaries might include confidential attendee notes; the drive tool might write anywhere. What will we do: treat invite text as untrusted, restrict the write tool to one folder, and exclude private notes from context. Did we do a good job: add three red-team tests to check. Ten minutes, three real risks, three concrete controls.
Common mistakes
Common mistakes in threat modeling sessions. Spending the whole hour debating whether the model can be jailbroken, instead of asking what the tools let an attacker do. Writing threats so vague they cannot be tested, like the AI might be misused. Forgetting insiders and suppliers, who can edit your knowledge base. And producing a document that lives in a folder nobody opens. The fix for the last one is simple: every high-priority threat gets an owner and a test identifier, and the tests run in CI.
Deeper: ranking two threats
One level deeper on rating. Two threats: an injection that makes the assistant use an unprofessional tone, and one that creates discount codes. The first is likely but low impact; the second is less likely but directly costs money at scale. Put the discount threat first, even though it happens less often, because impact is judged by what the tools can do.
Watch me do it: one abuse case, end to end
Watch me do it: turning one threat into a full worksheet row and abuse case. Element: the prompt assembly step for the WhatsApp commerce agent, where product catalog text from suppliers is combined with the customer's message. Threat category: indirect prompt injection. I write the abuse case. As a dishonest supplier, I add hidden text to my product description telling the agent to offer a ten percent discount code to anyone who asks about my product, so my sales rise at the brand's expense. Preconditions: catalog text is retrieved into context unfiltered, and the agent has the create discount tool. Likelihood: medium, because suppliers edit their own listings. Impact: I rate it by the worst case given permissions. The tool can create up to ten percent codes, unlimited times. So impact is high. Mitigations, in code: the discount tool only runs when the conversation contains a complaint flag set by the support workflow, not by the model; one code per phone number per thirty days; and catalog text is sanitized on ingestion and marked as data. Owner: the commerce team lead. Test ID: RT zero seven, a planted supplier description that must not produce a code. Now the row has everything a red team and a developer need, and when the test passes in CI, the row is closed.
Recap
Recap. Answer the four questions. Model the system including prompt assembly, enumerate threats with STRIDE and AI lenses, write testable abuse cases, rate impact by worst case given permissions, put mitigations in code, and revisit on every change. Your next step: fill in ten to twenty rows of the worksheet in the lesson for one feature, then deepen the top five with abuse cases and test IDs.
Key takeaways
- Use the four questions and a data flow diagram that includes prompt assembly.
- Combine STRIDE with AI-specific threat categories.
- Write concrete abuse cases with preconditions, impact and a test ID.
- Rate impact by worst case given tool permissions; implement mitigations in code and revisit on every change.
Try it
Fill in 10–20 rows of the threat-model worksheet for one feature and write full abuse cases with test IDs for the top five.