Latest AI Techniques: RAG, Tool Use, Agents & MCPTool and function calling · Lesson 7 of 20
Designing good tools for models
Video lecture
Designing good tools for models
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Designing tools for models
When an AI agent misbehaves, the easy conclusion is that the model isn't smart enough. Surprisingly often, the real culprit is the tools you gave it: vague names, overlapping purposes, cryptic outputs and errors that say nothing. In this lesson you'll learn six principles for designing tools that models use reliably, see a before and after, and set up a small test harness so tool quality becomes something you measure rather than hope for.
0:32 Why it matters
Why focus on tool design? Because it's usually the cheapest fix available. Swapping to a bigger model costs money on every call forever. Rewriting a confusing tool description costs ten minutes once. Here's an analogy. Give a new employee a remote control with thirty unlabelled buttons and they'll press the wrong ones. Give them three clearly labelled buttons for the three jobs they actually do, and they'll get it right first time.
1:03 Principles 1–2
Think of tool design as writing an API for a capable colleague who reads documentation literally and can't ask you questions. Principle one: fewer, higher-level tools. Instead of list customers, get customer, list orders and get order, offer find customer orders if that's what the workflow needs. Every call is a round trip and a chance to go wrong. Principle two: clear, distinct names and boundaries. If search docs and search knowledge overlap, the model will choose inconsistently. Merge them, or state exactly what each one covers.
1:41 Principles 3–4
Principle three: inputs that are hard to get wrong. Use enums for fixed choices, explicit formats like ISO dates, sensible defaults, and as few required fields as possible. Principle four: outputs designed for a model to read. Return meaningful field names, human-readable names alongside IDs, and a sensible amount of data, with pagination or a detail level option instead of dumping everything.
2:08 Principles 5–6
Principle five: errors that teach. Date must be in year-month-day format; you sent twenty-nine slash nine lets the model fix its next call. Error four hundred teaches nothing. Principle six: safe by construction. Separate read tools from write tools, and make destructive operations explicit. A dedicated cancel order tool can be gated with confirmation and audit. A generic update order tool with a status field hides the danger inside a parameter.
2:39 Simple example: rename and describe
A simple example first. Your tool is called get data, with one input called q. The model has to guess what it does and what q means. Rename it search help articles, describe what it covers and doesn't, say it returns up to five results with titles and dates, and rename q to query with a description. Same code underneath, but now the model calls it at the right time with sensible input. Nothing about the model changed.
3:13 Worked example: agency CRM
Here's a real-world shaped example. A small agency connects an assistant to its CRM with twelve thin endpoints. The model keeps calling get contact repeatedly and forgets to check the deal stage. Version two has three tools. Find account returns the account, main contacts and open deals in one summary. Log activity records a note. Draft follow up drafts an email but never sends. Selection errors drop, runs get shorter, and nothing reaches a client without a human.
3:47 Business example (illustrative)
More detail on the agency CRM case, illustrative. With twelve thin tools, a typical account lookup took about six tool calls and the model forgot the deal stage in roughly one run in five. With three task-level tools, the same job took two calls, and the missed-stage problem disappeared from their test set. Runs got faster and cheaper, and because drafting never sends, the account managers kept full control of client communication.
4:18 Test your tools
Now make it measurable. Build a small set of user requests, each with the expected tool, or no tool, and the key arguments. Include near-misses: a thank-you message that needs no tool, and where is my order with no ID, where the right behaviour is to ask. Score selection accuracy and argument accuracy, run each case several times, and read the transcripts of misses. Ask yourself: would a new human colleague have misunderstood this description too? Usually the fix is in the description.
4:54 Scaling tool sets
As tool counts grow, especially after connecting several MCP servers, loading every definition into every request wastes context and hurts selection. Current answers include tool search, where tools are deferred until the model looks for them; a routing step that picks a toolset first; and code-mode, where the model writes a short sandboxed script against a documented library instead of calling thirty tiny tools. The lesson has a runnable harness that scores tool selection with the Claude API, plus notes on each of these patterns.
5:31 Common mistakes
Common mistakes. Mirroring your internal API one-to-one, so the model makes ten calls where one would do. Overlapping tools with near-identical names. Returning cryptic codes without labels. Changing a tool's behaviour without updating its description. Hiding dangerous actions inside generic update tools. And never testing tool selection, so you only discover problems when a customer does.
5:55 How you'll know it's working
How will you know your tool design improved? Selection accuracy goes up on your test set, including near-miss cases. Average tool calls per task goes down, because task-level tools do more per call. Invalid-argument errors fall. And, most tellingly, when you read transcripts, the model's choices look like what a sensible new colleague would do with the same descriptions.
6:21 Watch me do it: selection harness
Watch me do it. I open the selection harness. First, I load my tool definitions from tools dot json, so they're versioned with the code. Next, the cases list: each has a request, the expected tool or none, and key arguments. Notice the two near-misses: a thank-you message and where is my order without an ID. First call sends one request with the tools and returns the first tool use block's name and input, or none. Then the loop compares the tool name and arguments with expectations, counts passes, and prints every miss. I run it. The first run shows one miss: where is my order triggers get order status with an invented ID. I update the description to say, if no order ID was given, ask the customer for it. I run it three more times, because behaviour varies, and all four cases pass each time. I note the description change in the changelog.
7:28 Recap
To recap: many agent failures are tool design failures. Prefer fewer, task-level tools with distinct boundaries, hard-to-misuse inputs, readable outputs, teaching errors and explicit destructive actions. Test selection and arguments with near-miss cases, and version tool descriptions with your code, because a description change is a behaviour change. Your next step: redesign three endpoints or spreadsheet actions you'd expose to an AI as task-level tools. Next up, agents.
7:58 Try this now (30 minutes)
Try this now. Take three endpoints, spreadsheet actions or buttons you'd expose to an AI assistant. Redesign them as task-level tools: better names, descriptions with when to use and when not to, enums for fixed choices, and error messages that explain how to fix a bad call. Then write four test requests, including one where no tool should be called, and run them through the selection harness from the lesson.
Tools are an interface for a new kind of user
Designing tools for a model is like designing an API for a capable colleague who reads the documentation literally and cannot ask you questions. Many agent failures that look like "the model is not smart enough" are actually tool design problems.
Principles
1. Fewer, higher-level tools. Instead of list_customers, get_customer, list_orders and get_order, consider find_customer_orders(customer_email, status) if that is what the workflow needs. Each tool call costs a round trip and a chance to go wrong. Match tools to tasks, not to your database tables.
2. Clear, distinct names and boundaries. If two tools overlap (search_docs and search_knowledge), the model will choose inconsistently. Make each tool's purpose unmistakable, or merge them.
3. Inputs that are hard to get wrong. Use enums for fixed choices, explicit formats for dates (ISO 8601) and IDs, sensible defaults, and a minimal set of required fields.
4. Outputs designed for reading by a model. Return meaningful fields, human-readable names alongside IDs, and a sensible amount of data. Offer pagination or a detail_level parameter (summary vs full) rather than dumping everything.
5. Errors that teach. "Date must be in YYYY-MM-DD format; you sent 29/09/2026" lets the model fix its call on the next try.
6. Safe by construction. Separate read tools from write tools. Make destructive operations explicit (cancel_order rather than a generic update_order with a status field), so they can be gated with confirmation.
A before-and-after
Before:
{"name": "query", "description": "Run a query",
"input_schema": {"type": "object", "properties": {"q": {"type": "string"}}}}After:
{
"name": "search_help_articles",
"description": "Search published help-centre articles by meaning and keywords. Use for how-to and policy questions. Does NOT include order or account data; use get_order_status for those. Returns up to 5 results with title, url, snippet and last_updated.",
"input_schema": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "The user's question rephrased as a standalone search query"},
"product": {"type": "string", "enum": ["web", "ios", "android"], "description": "Filter by product if the user mentioned one"}
},
"required": ["query"]
}
}The second version tells the model what the tool covers, what it does not, what comes back, and how to fill the inputs.
Testing tools with the model
Write tests the same way you would for prompts:
- A set of user requests, each with the expected tool (or no tool) and key arguments.
- Measure tool selection accuracy (right tool chosen), argument accuracy (valid, correct values) and unnecessary calls.
- Include near-miss cases: questions that sound like they need a tool but do not, and vice versa.
Read transcripts. When the model misuses a tool, ask first: "Would a new human colleague have misunderstood this description too?" Usually, the fix is in the description or the tool's shape.
How many tools?
There is no fixed limit, but selection quality tends to fall as the number of similar tools grows, and every definition consumes context. Strategies:
- Group tools by task and expose only the relevant group (for example via a router step).
- Use tool search or dynamic loading if your platform supports it, so the model discovers tools as needed.
- Consolidate rarely used tools.
Worked example: a sales assistant
A small agency connects an assistant to its CRM. Version 1 exposes twelve thin CRM endpoints. The model often calls get_contact repeatedly and forgets to check deal stage. Version 2 has three tools: find_account(name_or_domain) returning account, main contacts and open deals in one summary; log_activity(account_id, type, notes); and draft_follow_up(account_id, goal), which drafts but never sends. Selection errors drop, runs are shorter, and nothing is sent without a human.
Failure modes
- Tools that mirror internal APIs one-to-one.
- Tool results with cryptic codes and no labels.
- Write tools without confirmation or audit logs.
- Changing a tool's behaviour without updating its description.
Hands-on: a tool-selection test harness
Treat tool definitions like code: test them. The harness below sends each test request to the model with your tool set, captures the first tool call (or none), and scores tool selection and key arguments. It uses the Claude API; the same idea works with any provider.
import json, os
import anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
TOOLS = json.load(open("tools.json")) # your tool definitions, versioned with your code
CASES = [
{"ask": "How do I reset my password on iOS?", "tool": "search_help_articles", "args": {"product": "ios"}},
{"ask": "Where is order A55120?", "tool": "get_order_status", "args": {"order_id": "A55120"}},
{"ask": "Thanks, that's all!", "tool": None}, # near-miss: no tool needed
{"ask": "Where is my order?", "tool": None}, # should ask for the ID, not guess
]
def first_call(ask):
resp = client.messages.create(model=MODEL, max_tokens=1000, tools=TOOLS,
messages=[{"role": "user", "content": ask}])
for block in resp.content:
if block.type == "tool_use":
return block.name, block.input
return None, {}
score = {"selection": 0, "args": 0, "n": len(CASES)}
for c in CASES:
name, args = first_call(c["ask"])
ok_sel = name == c["tool"]
ok_args = all(args.get(k) == v for k, v in c.get("args", {}).items())
score["selection"] += ok_sel
score["args"] += ok_sel and ok_args
if not (ok_sel and ok_args):
print("MISS:", c["ask"], "->", name, args)
print(score)Run it several times per case (model behaviour varies), keep a changelog of description edits, and re-run the suite whenever you change a description, add a tool or switch model. Near-miss cases such as "Where is my order?" (no ID) are the most valuable: they catch hallucinated arguments.
Designing for large tool sets
When an agent needs access to dozens or hundreds of tools (common once you connect several MCP servers), loading every definition into every request wastes context and hurts selection. Current options include:
- Tool search / deferred loading: some APIs let you mark tools as deferred so the model searches for them only when needed; Anthropic's API, for example, offers regex and BM25 tool-search tools for this.
- Routing: a cheap first step picks a toolset (billing, orders, marketing) before the main call.
- Code-mode tools: instead of 30 thin tools, expose a small, sandboxed code-execution environment with a well-documented client library, so the model writes a short script that calls several operations and returns only the summary. This can cut round trips and context dramatically, but it raises the security stakes: sandbox it properly.
Whichever you use, the design principles in this lesson still decide success.
Going further
Version tool definitions alongside prompts, and include them in your evaluation runs. A tool description change is a behaviour change. For tools exposed to many AI applications at once, a standard interface such as the Model Context Protocol (module 5) lets you define them once and reuse them across compatible clients.
Key takeaways
- Tool design, not model intelligence, causes many agent failures.
- Prefer fewer, task-level tools with distinct boundaries, hard-to-misuse inputs and model-readable outputs.
- Return errors that teach, separate read and write tools, and make destructive actions explicit.
- Test tool selection and argument accuracy with near-miss cases and read transcripts.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Take three API endpoints or spreadsheet actions you would expose to an AI assistant and redesign them as task-level tools with improved descriptions, enums and error messages.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.