AI Security: Prompt Injection, Data Leakage and Red TeamingAgents, tools, supply chain and RAG poisoning · Lesson 8 of 17

Excessive agency and secure tool design

Article · 15 min · 8 min lecture

Video lecture

Excessive agency and secure tool design

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Excessive agency and tool security

  • Three root causes
  • Seven tool design principles
  • A hardened tool in code

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Excessive agency: the risk that grew fastest

OWASP's 2026 update moved Excessive Agency up to LLM03, and the Agentic Top 10 adds tool misuse and identity and privilege abuse. The reason is simple: agents now hold real permissions. When an LLM-driven system can act, every manipulation (injection, jailbreak, hallucination) becomes a potential real-world action.

OWASP describes three root causes:

  • Excessive functionality: the agent has tools it does not need for its task (a summarizer that can also delete files).
  • Excessive permissions: tools have broader access than needed (a support bot's database account can read all customers).
  • Excessive autonomy: high-impact actions happen without human verification.

Design principles for tools

1. Minimal tool set per task. Give each agent or route only the tools it needs. A "read my calendar" assistant does not need send_email.

2. Narrow, purpose-built tools over generic ones. Compare:

Generic (risky)Narrow (safer)
run_sql(query)get_order_status(order_id)
send_email(to, subject, body)send_order_update(order_id, template_id) (recipient derived server-side from the order)
http_request(url, method, body)get_exchange_rate(currency) (fixed endpoint)
run_shell(cmd)run_tests(package) (fixed command, sandboxed)

Narrow tools move security decisions from the model into code.

3. Authorization in the tool, not the prompt. Every tool call must check that the end user is allowed to perform it on that resource. Use the user's identity (for example a user-scoped OAuth token), not a powerful service account. This also addresses the classic confused deputy problem, where a privileged component is tricked into acting for someone who lacks the privilege.

4. Hard limits in code. Maximum refund amounts, rate limits per user, allowed recipients, allowed file paths, time windows. The model can ask for anything; the tool enforces policy.

5. Human approval for high-impact actions. Payments, deletions, external communications, permission changes. Show the human exactly what will happen (recipient, amount, content), not the model's summary of it. Beware approval fatigue and human-agent trust exploitation (ASI09): a manipulated agent can present misleading justifications. Keep approval screens factual and derived from the actual tool arguments.

6. Dry-run and reversible actions. Prefer tools that stage changes (draft, pending) and support undo.

7. Agent identity and credentials. Give agents their own identities with short-lived, scoped credentials; never share a human's long-lived keys with an agent. Log every action with the agent identity and the user on whose behalf it acted.

Hands-on: a hardened tool

# tools/refunds.py: the model can request; the tool decides
from decimal import Decimal
from dataclasses import dataclass

MAX_AUTO_REFUND = Decimal("5000")  # in PKR minor-unit-safe Decimal; set per market in config

class ToolDenied(Exception): ...

@dataclass
class Ctx:
    user_id: str           # authenticated end user (from session, never from the model)
    agent_id: str
    request_id: str

def create_refund(ctx: Ctx, order_id: str, amount: str, reason: str) -> dict:
    order = orders_repo.get(order_id)
    if order is None or order.customer_id != ctx.user_id:           # authorization on the resource
        raise ToolDenied("order not found for this customer")
    amt = Decimal(amount)
    if amt <= 0 or amt > order.paid_amount:
        raise ToolDenied("invalid amount")
    if refunds_repo.count_for_customer(ctx.user_id, days=30) >= 2:  # abuse limit
        raise ToolDenied("refund limit reached; escalate to human")
    status = "pending_approval" if amt > MAX_AUTO_REFUND else "approved"
    refund = refunds_repo.create(order_id=order_id, amount=amt, reason=reason[:500],
                                 status=status, requested_by=ctx.agent_id, on_behalf_of=ctx.user_id,
                                 request_id=ctx.request_id)
    audit_log.write("create_refund", ctx=ctx, order_id=order_id, amount=str(amt), status=status)
    return {"refund_id": refund.id, "status": status}

Note what the model cannot influence: who the customer is, the per-customer limit, the approval threshold, and the audit trail.

Code execution tools

Agents that run code (data analysis, coding agents) need strong sandboxes: containers or microVMs with no credentials, restricted network egress, CPU/memory/time limits, and disposable filesystems. The Agentic Top 10 lists unexpected code execution (ASI05) because generated code is effectively attacker-influenced input when injection is possible.

Worked example

An accounting firm in Abu Dhabi piloted an agent to reconcile invoices. The prototype used a generic run_sql tool with a service account that could read every client's ledger. Red-teaming showed an injected invoice note could make it query another client's data. The redesign replaced run_sql with four narrow tools scoped to the current client engagement via a per-session token, and moved write operations to a draft-and-approve flow.

Pitfalls

  • Service accounts with god-mode behind the agent.
  • Approval screens that show the model's description instead of actual parameters.
  • Tools that accept identities from the model ("customer_id" as a free argument).
  • No audit trail linking actions to users and agents.

How to measure success

Every tool has a documented scope, authorization check, hard limits and audit log; high-impact actions require factual human approval; red-team attempts to misuse tools are denied by code.

Key takeaways

  • Excessive agency has three roots: too much functionality, too many permissions, too much autonomy.
  • Prefer minimal, narrow, purpose-built tools; authorize every call against the real user and resource.
  • Enforce limits in code and require factual human approval for high-impact actions.
  • Give agents scoped, short-lived identities, prefer reversible actions, audit everything, and sandbox code execution.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which tool design is safest for a support agent?
  2. Why should approval dialogs show the actual tool arguments rather than the agent's summary?

Put it into practice

Audit every tool in one agent against the seven principles and replace your most generic tool with a narrow, authorized one with limits and audit logging.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.