AI Security: Prompt Injection, Data Leakage and Red TeamingAgents, tools, supply chain and RAG poisoning · Lesson 8 of 17
Excessive agency and secure tool design
Video lecture
Excessive agency and secure tool design
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Excessive agency and tool security
When a chatbot is manipulated, you get a bad paragraph. When an agent is manipulated, you might get a refund issued, a file deleted or an email sent to a stranger. That is why OWASP's twenty twenty-six update moved excessive agency up to number three. In this lecture you will learn the three root causes of excessive agency and seven design principles for tools that stay safe even when the model is fooled.
0:32 Analogy: the corporate card
Here is an analogy. When a company gives a new employee a corporate card, it does not give them unlimited credit on every merchant in the world. It sets a limit, restricts categories, requires receipts, and asks for manager approval above a threshold. The employee can be brilliant and trustworthy, and the controls still exist, because mistakes and fraud happen. Your agent's tools are its corporate card. Scope, limits, receipts and approvals belong in the card, not in a note asking the employee to be careful.
1:09 Root causes
OWASP describes three root causes. Excessive functionality: the agent has tools it does not need, like a summarizer that can also delete files. Excessive permissions: tools have broader access than needed, like a support bot whose database account can read every customer. And excessive autonomy: high-impact actions happen with no human verification. Fix all three.
1:33 Principles 1–2
Principle one: a minimal tool set per task. Principle two: narrow, purpose-built tools instead of generic ones. Instead of run SQL, offer get order status. Instead of a free-form send email, offer send order update, where the recipient is derived server-side from the order. Instead of an arbitrary HTTP request, a fixed endpoint. Narrow tools move security decisions out of the model and into code, where they belong.
2:03 Principle 3: authorize in the tool
Principle three: authorization lives in the tool, not the prompt. Every call must check that the end user is allowed to perform this action on this specific resource. Use the user's identity, for example a user-scoped token, not a powerful service account. That also defuses the classic confused deputy problem, where a privileged component is tricked into acting on behalf of someone who lacks the privilege.
2:32 Principles 4–5
Principle four: hard limits in code. Maximum refund amounts, per-user rate limits, allowed recipients, allowed file paths and time windows. The model can ask for anything; the tool enforces policy. Principle five: human approval for high-impact actions like payments, deletions, external messages and permission changes. Show the human the actual parameters, the recipient, the amount, the content, not the model's summary, because a manipulated agent can present persuasive but misleading justifications.
3:03 Principles 6–7
Principle six: prefer dry runs and reversible actions. Draft, stage, then commit, with undo where possible. Principle seven: give agents their own identities with short-lived, scoped credentials. Never hand an agent a human's long-lived keys. And log every action with both the agent identity and the user on whose behalf it acted, so incidents can be investigated.
3:28 Hands-on: create_refund
The lesson includes a hardened refund tool in Python. Look at what the model cannot influence. It cannot choose who the customer is, because the user identity comes from the authenticated session. It cannot exceed the order amount or the thirty-day refund limit. Anything above the automatic threshold goes to pending approval. And every call writes an audit record. The model can request. The tool decides.
3:57 Code execution + case
Agents that run code need special care, because generated code is effectively attacker-influenced input whenever injection is possible. The Agentic Top 10 calls this unexpected code execution. Use containers or micro virtual machines with no credentials, restricted network egress, resource and time limits, and disposable filesystems. Then a story: an accounting firm in Abu Dhabi piloted a reconciliation agent with a generic SQL tool and a service account that could read every client's ledger. An injected invoice note made it query another client's data. The redesign used four narrow, client-scoped tools and a draft-and-approve flow for writes.
4:39 Example: narrowing send_email
A simple example of narrowing a tool. Generic tool: send email with any recipient, any subject, any body. Narrow tool: send order update, taking an order number and a template name. The server looks up the customer's email from the order, fills the template, and sends. Now try to exfiltrate through it. You cannot choose the recipient. You cannot write free text. The worst a manipulated model can do is send a legitimate template to the legitimate customer. The attack surface went from unlimited to almost nothing.
5:17 Common mistakes
Common mistakes in tool design. A service account with god-mode access behind a friendly agent. Tools that accept a customer identifier as a free argument from the model. Approval screens that show the model's summary instead of the real parameters. And no audit trail linking an action to both the agent and the user it acted for. Here is a quick exercise: open your agent's tool list and, for each tool, write down the worst thing it could do if the model were fully manipulated.
5:54 Deeper: four narrow tools
One level deeper on the Abu Dhabi accounting agent. The four narrow tools were: list unmatched invoices, get invoice, propose match and flag for review. Each took the engagement ID from a per-session token, not from the model. Writes went through propose match, which only created a draft for an accountant to approve.
6:17 Watch me do it: create_refund × 3
Watch me do it: stepping through the hardened refund tool from the lesson with three requests. Request one, legitimate. The authenticated customer asks for a refund of three thousand rupees on their own order. The tool loads the order from the repository and checks the customer ID on the order matches the user ID in the context, which came from the session, not from the model. It matches. The amount is positive and below the paid amount. The customer has had no refunds in the last thirty days. The amount is below the automatic threshold, so status is approved, a refund row is created, and the audit log records the agent, the customer and the request ID. Request two, an injected attempt. Hidden text in a product review convinces the model to request a refund on a different customer's order. The model passes that order ID, but the ownership check fails and the tool raises tool denied: order not found for this customer. Notice the message doesn't confirm the order exists. Request three, persuasion. A customer talks the agent into asking for twelve thousand rupees. The tool accepts the request but sets status to pending approval because it exceeds the threshold, and a human sees the real amount and order. Three requests, and in every case the code, not the model, made the decision.
7:54 Recap
Recap. Cut functionality, permissions and autonomy to what the task needs. Build narrow tools, authorize every call against the real user and resource, enforce limits in code, require factual human approval for high-impact actions, prefer reversible operations, and give agents their own scoped identities with audit logs. Your next step: audit every tool in one agent against the seven principles and replace your most generic tool with a narrow one.
Excessive agency: the risk that grew fastest
OWASP's 2026 update moved Excessive Agency up to LLM03, and the Agentic Top 10 adds tool misuse and identity and privilege abuse. The reason is simple: agents now hold real permissions. When an LLM-driven system can act, every manipulation (injection, jailbreak, hallucination) becomes a potential real-world action.
OWASP describes three root causes:
- Excessive functionality: the agent has tools it does not need for its task (a summarizer that can also delete files).
- Excessive permissions: tools have broader access than needed (a support bot's database account can read all customers).
- Excessive autonomy: high-impact actions happen without human verification.
Design principles for tools
1. Minimal tool set per task. Give each agent or route only the tools it needs. A "read my calendar" assistant does not need send_email.
2. Narrow, purpose-built tools over generic ones. Compare:
| Generic (risky) | Narrow (safer) |
|---|---|
run_sql(query) | get_order_status(order_id) |
send_email(to, subject, body) | send_order_update(order_id, template_id) (recipient derived server-side from the order) |
http_request(url, method, body) | get_exchange_rate(currency) (fixed endpoint) |
run_shell(cmd) | run_tests(package) (fixed command, sandboxed) |
Narrow tools move security decisions from the model into code.
3. Authorization in the tool, not the prompt. Every tool call must check that the end user is allowed to perform it on that resource. Use the user's identity (for example a user-scoped OAuth token), not a powerful service account. This also addresses the classic confused deputy problem, where a privileged component is tricked into acting for someone who lacks the privilege.
4. Hard limits in code. Maximum refund amounts, rate limits per user, allowed recipients, allowed file paths, time windows. The model can ask for anything; the tool enforces policy.
5. Human approval for high-impact actions. Payments, deletions, external communications, permission changes. Show the human exactly what will happen (recipient, amount, content), not the model's summary of it. Beware approval fatigue and human-agent trust exploitation (ASI09): a manipulated agent can present misleading justifications. Keep approval screens factual and derived from the actual tool arguments.
6. Dry-run and reversible actions. Prefer tools that stage changes (draft, pending) and support undo.
7. Agent identity and credentials. Give agents their own identities with short-lived, scoped credentials; never share a human's long-lived keys with an agent. Log every action with the agent identity and the user on whose behalf it acted.
Hands-on: a hardened tool
# tools/refunds.py: the model can request; the tool decides
from decimal import Decimal
from dataclasses import dataclass
MAX_AUTO_REFUND = Decimal("5000") # in PKR minor-unit-safe Decimal; set per market in config
class ToolDenied(Exception): ...
@dataclass
class Ctx:
user_id: str # authenticated end user (from session, never from the model)
agent_id: str
request_id: str
def create_refund(ctx: Ctx, order_id: str, amount: str, reason: str) -> dict:
order = orders_repo.get(order_id)
if order is None or order.customer_id != ctx.user_id: # authorization on the resource
raise ToolDenied("order not found for this customer")
amt = Decimal(amount)
if amt <= 0 or amt > order.paid_amount:
raise ToolDenied("invalid amount")
if refunds_repo.count_for_customer(ctx.user_id, days=30) >= 2: # abuse limit
raise ToolDenied("refund limit reached; escalate to human")
status = "pending_approval" if amt > MAX_AUTO_REFUND else "approved"
refund = refunds_repo.create(order_id=order_id, amount=amt, reason=reason[:500],
status=status, requested_by=ctx.agent_id, on_behalf_of=ctx.user_id,
request_id=ctx.request_id)
audit_log.write("create_refund", ctx=ctx, order_id=order_id, amount=str(amt), status=status)
return {"refund_id": refund.id, "status": status}Note what the model cannot influence: who the customer is, the per-customer limit, the approval threshold, and the audit trail.
Code execution tools
Agents that run code (data analysis, coding agents) need strong sandboxes: containers or microVMs with no credentials, restricted network egress, CPU/memory/time limits, and disposable filesystems. The Agentic Top 10 lists unexpected code execution (ASI05) because generated code is effectively attacker-influenced input when injection is possible.
Worked example
An accounting firm in Abu Dhabi piloted an agent to reconcile invoices. The prototype used a generic run_sql tool with a service account that could read every client's ledger. Red-teaming showed an injected invoice note could make it query another client's data. The redesign replaced run_sql with four narrow tools scoped to the current client engagement via a per-session token, and moved write operations to a draft-and-approve flow.
Pitfalls
- Service accounts with god-mode behind the agent.
- Approval screens that show the model's description instead of actual parameters.
- Tools that accept identities from the model ("customer_id" as a free argument).
- No audit trail linking actions to users and agents.
How to measure success
Every tool has a documented scope, authorization check, hard limits and audit log; high-impact actions require factual human approval; red-team attempts to misuse tools are denied by code.
Key takeaways
- Excessive agency has three roots: too much functionality, too many permissions, too much autonomy.
- Prefer minimal, narrow, purpose-built tools; authorize every call against the real user and resource.
- Enforce limits in code and require factual human approval for high-impact actions.
- Give agents scoped, short-lived identities, prefer reversible actions, audit everything, and sandbox code execution.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Audit every tool in one agent against the seven principles and replace your most generic tool with a narrow, authorized one with limits and audit logging.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.