AI Security: Prompt Injection, Data Leakage and Red TeamingFrameworks: OWASP, MITRE ATLAS and NIST · Lesson 3 of 17
OWASP Top 10 for LLM and Agentic Applications
Video lecture
OWASP Top 10 for LLM and Agentic Applications
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 OWASP Top 10 for LLM and Agentic Applications
When a client's security team asks, have you tested for LLM zero one, they are speaking the language of the OWASP Top 10 for LLM Applications. In this lecture you will learn that list, what changed in the twenty twenty-six edition released in August, the separate Top 10 for Agentic Applications, and how to turn both into a coverage matrix that proves you have actually tested what matters.
0:30 Analogy: the pre-flight checklist
Why learn a list at all? Here is an analogy. A pilot's pre-flight checklist does not teach anyone to fly. But it guarantees that nobody forgets the fuel, the flaps or the doors, however experienced or tired they are. The OWASP lists are pre-flight checklists for LLM applications. They will not design your system for you, but they make sure nobody on your team forgets an entire category of risk, and they give you a shared language with clients, auditors and vendors.
1:06 2025 edition
The OWASP GenAI Security Project publishes the Top 10 for LLM Applications, the most widely referenced list of LLM app risks. The twenty twenty-five edition runs: prompt injection; sensitive information disclosure; supply chain; data and model poisoning; improper output handling; excessive agency; system prompt leakage; vector and embedding weaknesses; misinformation; and unbounded consumption.
1:29 2026 edition (Aug 2026)
In early August twenty twenty-six, OWASP released the twenty twenty-six edition. No categories were added or removed, but most moved. The ranking method also changed: it now combines the community vote with analysis of a large corpus of real-world incidents. Prompt injection stays at number one, with its scope expanded to cross-modal attacks, like instructions hidden in images or audio, and to injections that persist in memory or a retrieval corpus. Sensitive information disclosure stays at two.
2:02 Key 2026 moves
Three moves stand out. Excessive agency rises to number three, reflecting how many agents now hold real permissions. System prompt leakage is renamed hidden context exposure, at number eight, widening it to everything the app shows the model but not the user: retrieved documents, tool definitions, memory and configuration. And improper output handling drops to number ten. The full reported order is on screen and in your lesson text. Verify against the official document before quoting it in contracts.
2:36 Always cite the edition
Because numbering changed, always write the edition with the identifier, for example LLM zero eight colon twenty twenty-six, hidden context exposure. Many tools, reports and contracts still use twenty twenty-five numbers, so keep a mapping table in your security documentation. Otherwise, two teams can discuss LLM zero three and mean completely different risks.
2:59 Agentic Top 10 (ASI01–ASI10, Dec 2025)
For agents, there is a separate list. In December twenty twenty-five, the same project published the Top 10 for Agentic Applications, identified as A S I zero one to ten. It covers agent goal hijack, tool misuse and exploitation, agent identity and privilege abuse, agentic supply chain compromise, unexpected code execution, memory and context poisoning, insecure inter-agent communication, cascading agent failures, human-agent trust exploitation, and rogue agents. Use it alongside the LLM list for any system that plans, remembers and acts.
3:34 Use them well
How should you use these lists? As a coverage checklist, not a complete threat model, because your system can have risks outside both, like business logic abuse of a discount tool. Map each item to concrete red-team tests and to controls in your architecture. And link items to your own threat-model identifiers so findings reports use shared language. Beware checkbox compliance: saying we considered prompt injection without tests is not coverage.
4:05 Case: UK fintech coverage matrix
A worked example. A UK fintech preparing for a client security review built a coverage matrix from the twenty twenty-six items plus relevant agentic items. Two gaps appeared at once. Their retrieval index was shared across tenants with filtering applied only after retrieval, a vector and embedding weakness. And their agent stored user preferences that any conversation could overwrite, a memory poisoning risk. Both were fixed before the review.
4:35 Drill: a RAG chatbot vs the 2026 list
A simple example of using the list for coverage. Take a customer-support chatbot with retrieval but no tools. Walk down the twenty twenty-six list. Prompt injection: applies, through retrieved documents. Sensitive information disclosure: applies, order data. Excessive agency: barely, no tools yet. Supply chain: applies, your model provider and packages. Poisoning: applies to the knowledge base. Unbounded consumption: applies. Misinformation: applies to policy answers. Hidden context exposure: applies to the system prompt. Vector and embedding weaknesses: applies to the index. Improper output handling: applies to the chat UI. Ten minutes, and you know where your tests should go.
5:18 Scenario: the bank's questionnaire
And a realistic procurement moment. A bank in Abu Dhabi sends your agency a vendor questionnaire asking how you address the OWASP Top 10 for LLM Applications. A weak answer says, we follow OWASP best practices. A strong answer attaches your coverage matrix: each item with the edition year, how it manifests in the delivered system, the controls in code, the red-team test identifiers and the date last tested. That second answer wins contracts, and it also forces you to actually do the work.
5:55 Deeper: answering a 2025-based form
One level deeper on edition mapping. A client's 2025-based questionnaire asks about LLM zero seven, system prompt leakage. In your 2026-based matrix, that is LLM zero eight, hidden context exposure, which is broader. Answer with both IDs and note the broader 2026 scope. That shows you are current without making the client redo their form.
6:19 Watch me do it: coverage matrix
Watch me do it: filling the coverage matrix for a RAG support agent that can also create refunds. Row LLM zero one, twenty twenty-six, prompt injection. Applies: yes. How it manifests: customer messages, plus retrieved reviews and emails. Controls: spotlighting of retrieved content, a tool allowlist per route, no general internet egress. Tests: red-team cases RT zero one to RT zero nine. Last tested: this month. Owner: the security lead. Row LLM zero three, excessive agency. Applies: yes, because of the refund tool. Controls: a cap in code, a per-customer limit and human approval above a threshold. Tests: RT twelve. Row LLM zero eight, hidden context exposure. Applies: yes. How it manifests: the system prompt and tool definitions. Controls: no secrets in prompts, a canary string. Test: a canary check in the direct-injection suite. Row LLM zero nine, vector and embedding weaknesses. Applies: yes. Here I stop, because the controls column is empty. Our index is shared by all business units, and filtering happens after retrieval. That is a real gap, so I create a ticket with an owner. Row ASI zero six, memory and context poisoning. Our agent has no long-term memory, so I write not applicable, with the reason. Every row now has either controls and tests, a ticket, or a written reason.
7:52 Recap
Recap. Know the twenty twenty-five list and the August twenty twenty-six re-ranking, including hidden context exposure and the rise of excessive agency. Add the Agentic Top 10 for agents. Always cite edition years, and turn the lists into a coverage matrix with controls, tests and owners. Your next step: copy the matrix template from the lesson and fill it in for one of your systems.
Why use a shared taxonomy
The OWASP GenAI Security Project's Top 10 for LLM Applications is the most widely referenced list of LLM application risks. Security teams, auditors, vendors and clients use it as a common vocabulary: "Have you tested for LLM01?" is a question you will hear in procurement and pen-test scoping. Knowing it well lets you communicate risk quickly and check coverage systematically.
The 2025 edition
Published in late 2024 as the 2025 edition, it reflected the rise of RAG and agents:
| ID (2025) | Risk |
|---|---|
| LLM01:2025 | Prompt Injection |
| LLM02:2025 | Sensitive Information Disclosure |
| LLM03:2025 | Supply Chain |
| LLM04:2025 | Data and Model Poisoning |
| LLM05:2025 | Improper Output Handling |
| LLM06:2025 | Excessive Agency |
| LLM07:2025 | System Prompt Leakage |
| LLM08:2025 | Vector and Embedding Weaknesses |
| LLM09:2025 | Misinformation (absorbing the older "overreliance" theme) |
| LLM10:2025 | Unbounded Consumption (broadening the older "model denial of service") |
The 2026 update (August 2026)
OWASP released the 2026 edition in early August 2026. Based on OWASP's release and independent summaries (verify against the official document before quoting item text in contracts or reports):
- No categories were added or removed, but most entries moved.
- The ranking method changed: it now combines the community vote with analysis of a large corpus of real-world LLM security incidents, rather than relying on the vote alone.
- Prompt Injection stays at LLM01, with scope expanded to cross-modal attacks (instructions hidden in images, audio or video) and persistence via memory or a RAG corpus.
- Sensitive Information Disclosure stays at LLM02.
- Excessive Agency rises to LLM03, reflecting the growth of agents with real permissions.
- System Prompt Leakage is renamed Hidden Context Exposure (LLM08), widening it from the system prompt to everything the app puts in front of the model that the user cannot see: retrieved documents, tool definitions, memory and configuration.
- Improper Output Handling moves down to LLM10.
The resulting 2026 order, as reported: LLM01 Prompt Injection; LLM02 Sensitive Information Disclosure; LLM03 Excessive Agency; LLM04 Supply Chain; LLM05 Data and Model Poisoning; LLM06 Unbounded Consumption; LLM07 Misinformation; LLM08 Hidden Context Exposure; LLM09 Vector and Embedding Weaknesses; LLM10 Improper Output Handling.
Practical advice: many existing tools, reports and contracts still reference the 2025 IDs. Always write the edition with the ID ("LLM08:2026 Hidden Context Exposure") and keep a mapping table in your security documentation.
The OWASP Top 10 for Agentic Applications (2026)
In December 2025, the same project published a separate Top 10 for Agentic Applications (IDs ASI01–ASI10) for systems that plan, hold memory, use tools and act with delegated authority. Its categories include agent goal hijack, tool misuse and exploitation, agent identity and privilege abuse, agentic supply chain compromise, unexpected code execution, memory and context poisoning, insecure inter-agent communication, cascading agent failures, human-agent trust exploitation, and rogue agents. Use it alongside the LLM list for any agentic system.
Using the lists well
- As a coverage checklist, not a complete threat model. Your system may have risks outside both lists (for example business-logic abuse of a discount tool).
- Map each item to concrete tests in your red-team plan and to controls in your architecture.
- Map to your threat model IDs, so findings reports use shared language.
Worked example: a coverage matrix
A UK fintech preparing for a client security review built a matrix: rows were OWASP 2026 items plus relevant ASI items; columns were "applies?", "controls", "tests", "last tested", "owner". Two gaps emerged immediately: no tests for LLM09:2026 Vector and Embedding Weaknesses (their RAG index was shared across tenants with filtering done after retrieval), and no controls for ASI06 memory poisoning (their agent stored "user preferences" that any conversation could overwrite). Both were fixed before the review.
Hands-on: coverage matrix template
| Item | Applies? | How it manifests here | Controls (code/config) | Red-team tests | Last tested | Owner |
|------|----------|-----------------------|------------------------|----------------|-------------|-------|
| LLM01:2026 Prompt Injection | yes | reviews + emails in context | spotlighting, tool allowlist, egress block | RT-01..RT-09 | 2026-09-10 | @sec |
| LLM03:2026 Excessive Agency | yes | refund tool | cap in code, approval > limit | RT-12 | | |
| LLM08:2026 Hidden Context Exposure | | | | | | |
| ASI06 Memory & Context Poisoning | | | | | | |Pitfalls
- Quoting IDs without the edition year, causing confusion as numbering changed in 2026.
- Treating the list as exhaustive.
- Checkbox compliance: "we considered LLM01" without tests is not coverage.
How to measure success
A maintained coverage matrix that references both lists with edition years, where every applicable item has controls and passing tests.
Key takeaways
- The OWASP LLM Top 10 is the shared vocabulary for LLM app risks; the 2026 edition (Aug 2026) re-ranked the same ten categories.
- 2026 highlights: Prompt Injection stays LLM01 with cross-modal scope; Excessive Agency rises to LLM03; System Prompt Leakage becomes Hidden Context Exposure (LLM08).
- Always cite IDs with the edition year and keep a 2025↔2026 mapping.
- Use the Agentic Top 10 (ASI01–ASI10) for agents, and turn both lists into a coverage matrix with controls and tests.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Build a coverage matrix for one system using the 2026 LLM Top 10 plus relevant ASI items, listing controls, tests and owners.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.