Latest AI Techniques: RAG, Tool Use, Agents & MCPTool and function calling · Lesson 7 of 20

Designing good tools for models

Article · 11 min · 8 min lecture

Video lecture

Designing good tools for models

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Designing tools for models

  • Six principles
  • Before and after
  • Test tool selection like code

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Tools are an interface for a new kind of user

Designing tools for a model is like designing an API for a capable colleague who reads the documentation literally and cannot ask you questions. Many agent failures that look like "the model is not smart enough" are actually tool design problems.

Principles

1. Fewer, higher-level tools. Instead of list_customers, get_customer, list_orders and get_order, consider find_customer_orders(customer_email, status) if that is what the workflow needs. Each tool call costs a round trip and a chance to go wrong. Match tools to tasks, not to your database tables.

2. Clear, distinct names and boundaries. If two tools overlap (search_docs and search_knowledge), the model will choose inconsistently. Make each tool's purpose unmistakable, or merge them.

3. Inputs that are hard to get wrong. Use enums for fixed choices, explicit formats for dates (ISO 8601) and IDs, sensible defaults, and a minimal set of required fields.

4. Outputs designed for reading by a model. Return meaningful fields, human-readable names alongside IDs, and a sensible amount of data. Offer pagination or a detail_level parameter (summary vs full) rather than dumping everything.

5. Errors that teach. "Date must be in YYYY-MM-DD format; you sent 29/09/2026" lets the model fix its call on the next try.

6. Safe by construction. Separate read tools from write tools. Make destructive operations explicit (cancel_order rather than a generic update_order with a status field), so they can be gated with confirmation.

A before-and-after

Before:

{"name": "query", "description": "Run a query",
 "input_schema": {"type": "object", "properties": {"q": {"type": "string"}}}}

After:

{
  "name": "search_help_articles",
  "description": "Search published help-centre articles by meaning and keywords. Use for how-to and policy questions. Does NOT include order or account data; use get_order_status for those. Returns up to 5 results with title, url, snippet and last_updated.",
  "input_schema": {
    "type": "object",
    "properties": {
      "query": {"type": "string", "description": "The user's question rephrased as a standalone search query"},
      "product": {"type": "string", "enum": ["web", "ios", "android"], "description": "Filter by product if the user mentioned one"}
    },
    "required": ["query"]
  }
}

The second version tells the model what the tool covers, what it does not, what comes back, and how to fill the inputs.

Testing tools with the model

Write tests the same way you would for prompts:

  • A set of user requests, each with the expected tool (or no tool) and key arguments.
  • Measure tool selection accuracy (right tool chosen), argument accuracy (valid, correct values) and unnecessary calls.
  • Include near-miss cases: questions that sound like they need a tool but do not, and vice versa.

Read transcripts. When the model misuses a tool, ask first: "Would a new human colleague have misunderstood this description too?" Usually, the fix is in the description or the tool's shape.

How many tools?

There is no fixed limit, but selection quality tends to fall as the number of similar tools grows, and every definition consumes context. Strategies:

  • Group tools by task and expose only the relevant group (for example via a router step).
  • Use tool search or dynamic loading if your platform supports it, so the model discovers tools as needed.
  • Consolidate rarely used tools.

Worked example: a sales assistant

A small agency connects an assistant to its CRM. Version 1 exposes twelve thin CRM endpoints. The model often calls get_contact repeatedly and forgets to check deal stage. Version 2 has three tools: find_account(name_or_domain) returning account, main contacts and open deals in one summary; log_activity(account_id, type, notes); and draft_follow_up(account_id, goal), which drafts but never sends. Selection errors drop, runs are shorter, and nothing is sent without a human.

Failure modes

  • Tools that mirror internal APIs one-to-one.
  • Tool results with cryptic codes and no labels.
  • Write tools without confirmation or audit logs.
  • Changing a tool's behaviour without updating its description.

Hands-on: a tool-selection test harness

Treat tool definitions like code: test them. The harness below sends each test request to the model with your tool set, captures the first tool call (or none), and scores tool selection and key arguments. It uses the Claude API; the same idea works with any provider.

import json, os
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
TOOLS = json.load(open("tools.json"))      # your tool definitions, versioned with your code

CASES = [
    {"ask": "How do I reset my password on iOS?", "tool": "search_help_articles", "args": {"product": "ios"}},
    {"ask": "Where is order A55120?",             "tool": "get_order_status",     "args": {"order_id": "A55120"}},
    {"ask": "Thanks, that's all!",                "tool": None},  # near-miss: no tool needed
    {"ask": "Where is my order?",                 "tool": None},  # should ask for the ID, not guess
]

def first_call(ask):
    resp = client.messages.create(model=MODEL, max_tokens=1000, tools=TOOLS,
                                  messages=[{"role": "user", "content": ask}])
    for block in resp.content:
        if block.type == "tool_use":
            return block.name, block.input
    return None, {}

score = {"selection": 0, "args": 0, "n": len(CASES)}
for c in CASES:
    name, args = first_call(c["ask"])
    ok_sel = name == c["tool"]
    ok_args = all(args.get(k) == v for k, v in c.get("args", {}).items())
    score["selection"] += ok_sel
    score["args"] += ok_sel and ok_args
    if not (ok_sel and ok_args):
        print("MISS:", c["ask"], "->", name, args)
print(score)

Run it several times per case (model behaviour varies), keep a changelog of description edits, and re-run the suite whenever you change a description, add a tool or switch model. Near-miss cases such as "Where is my order?" (no ID) are the most valuable: they catch hallucinated arguments.

Designing for large tool sets

When an agent needs access to dozens or hundreds of tools (common once you connect several MCP servers), loading every definition into every request wastes context and hurts selection. Current options include:

  • Tool search / deferred loading: some APIs let you mark tools as deferred so the model searches for them only when needed; Anthropic's API, for example, offers regex and BM25 tool-search tools for this.
  • Routing: a cheap first step picks a toolset (billing, orders, marketing) before the main call.
  • Code-mode tools: instead of 30 thin tools, expose a small, sandboxed code-execution environment with a well-documented client library, so the model writes a short script that calls several operations and returns only the summary. This can cut round trips and context dramatically, but it raises the security stakes: sandbox it properly.

Whichever you use, the design principles in this lesson still decide success.

Going further

Version tool definitions alongside prompts, and include them in your evaluation runs. A tool description change is a behaviour change. For tools exposed to many AI applications at once, a standard interface such as the Model Context Protocol (module 5) lets you define them once and reuse them across compatible clients.

Key takeaways

  • Tool design, not model intelligence, causes many agent failures.
  • Prefer fewer, task-level tools with distinct boundaries, hard-to-misuse inputs and model-readable outputs.
  • Return errors that teach, separate read and write tools, and make destructive actions explicit.
  • Test tool selection and argument accuracy with near-miss cases and read transcripts.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Two tools, search_docs and search_knowledge, overlap. What is the likely result?
  2. Which error message best helps a model recover?
  3. Why prefer cancel_order over update_order(status='cancelled')?

Put it into practice

Take three API endpoints or spreadsheet actions you would expose to an AI assistant and redesign them as task-level tools with improved descriptions, enums and error messages.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.