Building Production AI Agents · Tools, reasoning and multi-agent design · lesson 4 of 18 · 16 min
Designing tools agents use correctly
Tools are the agent's user interface
Humans get buttons and forms; models get tool names, descriptions and schemas. Most agent failures that look like "the model is dumb" are really bad tool design: ambiguous names, missing descriptions, schemas that allow nonsense, or outputs that bury the answer in noise. Treat every tool as a product with a user (the model) and invest in it the way you would invest in a public API.
Seven principles
- Name tools for intent, not implementation.
find_customer_by_emailbeatsdb_query_v2. Use a consistent verb_noun convention and namespaces when you have many (crm_search_contacts,crm_update_deal). - Write descriptions like onboarding notes. Say what the tool does, when to use it, when not to, what it returns, and any limits. Three to six sentences is normal for important tools.
- Constrain inputs with schemas. Use
enumfor fixed choices, formats for dates,minimum/maximumfor numbers,requiredfields andadditionalProperties: false. Where the provider offers strict tool use (Claude and OpenAI both do), enable it so arguments always validate. - Consolidate chatty tools. Instead of
list_users,list_events,create_event, offerschedule_meeting(attendees, window)that does the lookup internally. Fewer, higher-level tools mean fewer steps and fewer errors. - Return concise, meaningful outputs. Return names and human-readable fields, not opaque UUID soup. Offer a
detail: "summary" | "full"parameter when both are needed. Truncate and paginate, and say so ("showing 10 of 240; pass page=2"). - Make errors actionable. "Invalid date" is weak; "date must be YYYY-MM-DD and in the future, got 2026-13-01" lets the model fix itself.
- Design for safety. Separate read and write tools; make write tools idempotent (accept an
idempotency_key); never expose a genericrun_sqlorhttp_requesttool to an agent that reads untrusted content.
Poka-yoke: make mistakes impossible
Borrow the manufacturing idea of mistake-proofing. If the model keeps passing relative file paths, require absolute ones. If it confuses currencies, make currency a required enum. If it sends a campaign to the wrong list, require list_id from a previous search_lists result rather than free text. Each change removes a class of errors permanently, which beats adding more instructions to the prompt.
Worked example: rewriting a weak tool
Before:
{"name": "email", "description": "Sends email",
"input_schema": {"type": "object", "properties": {"to": {"type": "string"}, "body": {"type": "string"}}}}
Problems: no subject, no guidance about when to use it, no limit on recipients, no draft step, irreversible.
After:
{
"name": "draft_client_email",
"description": "Create an email DRAFT to one existing client contact. Does not send. Use after you have confirmed the contact_id with crm_search_contacts. Returns draft_id and a preview. A human must approve with send_draft before anything is sent. Keep bodies under 200 words and in the client's language.",
"strict": true,
"input_schema": {
"type": "object",
"properties": {
"contact_id": {"type": "string", "description": "ID from crm_search_contacts, e.g. c_101"},
"subject": {"type": "string", "maxLength": 120},
"body": {"type": "string", "maxLength": 2000},
"language": {"type": "string", "enum": ["en", "ar", "ur"]}
},
"required": ["contact_id", "subject", "body", "language"],
"additionalProperties": false
}
}
The irreversible action is now split into draft (agent) and send (human-approved), and the enum covers English, Arabic and Urdu audiences for a PK/UAE client base.
Hands-on: a tool contract test
Treat tool definitions as code with tests. This small harness checks every tool against your house rules before deployment.
import re
RULES = [
("name is verb_noun snake_case", lambda t: re.fullmatch(r"[a-z]+(_[a-z0-9]+)+", t["name"])),
("description >= 25 words", lambda t: len(t.get("description", "").split()) >= 25),
("schema forbids extra props", lambda t: t["input_schema"].get("additionalProperties") is False),
("all props described or enum", lambda t: all("description" in p or "enum" in p
for p in t["input_schema"].get("properties", {}).values())),
("required declared", lambda t: bool(t["input_schema"].get("required"))),
]
def lint_tools(tools: list[dict]) -> list[str]:
problems = []
for t in tools:
for label, check in RULES:
if not check(t):
problems.append(f"{t['name']}: {label}")
return problems
# In CI: assert not lint_tools(TOOLS), lint_tools(TOOLS)
Then build a tool-use eval: 20–50 realistic requests with the expected tool and key arguments, and measure how often the model picks the right tool with valid arguments. Re-run it every time you edit a description.
Scaling to many tools
Past a few dozen tools, selection accuracy and cost suffer because every definition sits in context on every call. Options: group tools into specialist sub-agents; load tools on demand (Anthropic offers a tool-search capability with deferred loading, and MCP clients can filter servers); or replace many narrow tools with a small number of general ones such as a sandboxed code-execution tool.
Pitfalls
- Two tools with overlapping purposes ("search_docs" and "find_docs"); the model will pick inconsistently.
- Returning raw HTML or full API payloads.
- Hiding critical usage rules in the system prompt instead of the tool description.
- Letting write tools accept free-text identifiers.
Measuring success
Track tool-selection accuracy, argument-validity rate, tool error rate, average tool calls per task and tokens per tool result. A good redesign typically reduces calls per task and error rates at the same time.
Video lecture: Designing tools agents use correctly
Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.
- Designing tools agents use correctly
- Why tool design matters
- Principles 1–3
- Principles 4–6
- Principle 7: safety
- Simple example: get_data → get_sales_summary
- Mistake-proofing (poka-yoke)
- Worked example: email tool rewrite
- Test tools like software
- How long should a description be?
- Deeper: the wrong-language email
- Watch me do it: lint_tools()
- Try this now
- Recap
Lecture transcript
Designing tools agents use correctly
If an agent keeps making silly mistakes, the instinct is to blame the model. But very often the real culprit is the tools you gave it. For a model, tool names, descriptions and schemas are the user interface. In this lesson you'll learn seven principles for designing tools that agents use correctly, and how to test them like real software.
Why tool design matters
Why does tool design matter so much? Because the model can't ask you questions about a vague tool. It just guesses. Imagine giving a new assistant a remote control where every button is unlabeled. They'll press something, and sometimes it'll be the wrong thing. Labels, grouping and guards on the dangerous buttons fix that. Teams often spend weeks tuning prompts when ten minutes of renaming tools and tightening schemas would have fixed the problem. In practice, tool design is the highest leverage improvement you can make to an agent.
Principles 1–3
Principle one: name tools for intent. find customer by email tells the model exactly what happens. db query v2 tells it nothing. Principle two: write descriptions like onboarding notes for a new colleague. What does it do, when should you use it, when shouldn't you, what comes back, and what are the limits. Three to six sentences is normal for an important tool. Principle three: constrain the inputs. Use enums for fixed choices, date formats, number ranges, required fields, and no extra properties. If your provider offers strict tool use, switch it on so arguments always match the schema.
Principles 4–6
Principle four: consolidate chatty tools. Rather than list users, list events and create event, give the agent one schedule meeting tool that does the lookups inside. Fewer, higher level tools mean fewer steps and fewer chances to go wrong. Principle five: return concise, meaningful outputs. Names and readable fields, not a wall of IDs. Paginate, and tell the model: showing ten of two hundred and forty, ask for page two. Principle six: make errors actionable. Not invalid date, but date must be year month day and in the future, and here's what you sent.
Principle 7: safety
Principle seven is safety. Keep read tools and write tools separate. Make write tools idempotent, so if the same request is retried, nothing happens twice. And never hand a generic run any SQL or fetch any URL tool to an agent that reads untrusted content like emails or web pages. That's an open door for prompt injection, which we'll cover in module four.
Simple example: get_data → get_sales_summary
A simple example. Your agent has a tool called get data, with one text parameter called query. Users ask for last week's sales in Dubai, and the model sends query equals Dubai sales. Your code has no idea what to do with that. Rewrite it as get sales summary, with a required city from a list, a start date and an end date in year month day format, and a description saying it returns totals by day, maximum thirty one days. Now the model sends city Dubai, start and end dates, and your code knows exactly what to run. Same model, dramatically better behavior.
Mistake-proofing (poka-yoke)
There's a manufacturing idea called poka yoke, which means mistake proofing. Apply it to tools. If the model keeps using relative file paths, require absolute ones. If it mixes up pounds and dirhams, make currency a required choice. If it emails the wrong list, require a list id that came from a previous search, not free text. Each change removes a whole category of error forever, which is far better than adding another warning to the prompt and hoping.
Worked example: email tool rewrite
Here's a worked example. A weak tool called email, described as sends email, with a to and a body. No subject, no guidance, no limits, and it's irreversible. The rewrite is called draft client email. It only creates a draft for one existing contact, explains that the contact id must come from the CRM search, limits the length, offers English, Arabic or Urdu as the language, and says a human must approve sending. The risky action is now split into a safe draft and a human approved send.
Test tools like software
Finally, test your tools like software. The lesson includes a small linter that checks names, description length, required fields and strict schemas, which you can run in continuous integration. Then build a tool selection eval: twenty to fifty realistic requests, each with the tool and key arguments you expect. Re run it every time you change a description. And once you pass a few dozen tools, consider grouping them into specialist sub agents or loading them on demand, because every definition costs context on every call.
How long should a description be?
A question people often ask: how long should a tool description be? Longer than you think, but not a novel. For an important tool, three to six sentences is typical: what it does, when to use it, when not to, what it needs first, the limits, and what it returns. For a trivial tool, one or two sentences is fine. The test is simple: could a smart new colleague use this tool correctly after reading only the description? If they'd need to ask a question, the model needs the answer written down too.
Deeper: the wrong-language email
Let's deepen the email tool example with a PK and UAE agency context. The agency manages clients in Lahore and Dubai, writing to contacts in English, Arabic and Urdu. With the old email tool, the agent once replied to a Dubai client in Urdu because the conversation included an Urdu note from a Lahore colleague. The new draft tool requires a language field, and the agent reads each contact's preferred language from the CRM before drafting. It also can't reach anyone who isn't an existing client contact. In the first month, illustrative numbers again, the drafts needing language corrections fell from several a week to none, and the account team stopped worrying that the agent might message a stranger.
Watch me do it: lint_tools()
Watch me do it. I'll run the tool linter from the lesson against a weak tool and then a strong one. The rules list has five checks: the name must be verb and noun in snake case, the description needs at least twenty five words, the schema must forbid extra properties, every property needs a description or an enum, and required fields must be declared. First, the old email tool. The linter prints: name fails, description too short, extra properties allowed, properties undescribed, no required list. Five problems. Now the draft client email tool. It passes every check. In continuous integration, the assert line fails the build whenever a teammate adds a tool that breaks these rules, so quality doesn't depend on someone remembering. Next, I'd take the five test requests from the activity and check the model picks this tool with valid arguments, before and after my edit.
Try this now
Try this now. Pick the tool your agent uses most, or the one it gets wrong most often. Rewrite its name as verb and noun. Expand its description to cover what it does, when to use it, when not to, and what it returns. Add enums and formats to every parameter you can. Then write five test requests a real user might send, and check whether the model picks this tool with valid arguments before and after your change. Keep the before and after numbers. They make a very persuasive slide for your team.
Recap
Let's recap. Tools are the agent's user interface. Name them for intent, describe them richly, constrain inputs, consolidate, return concise outputs, make errors actionable, and keep writes safe and idempotent. Your next step: take one tool you plan to give an agent, rewrite it with these principles, run the linter, and write five test requests for your tool selection eval.
Key takeaways
- Most agent failures trace back to tool design, not model capability.
- Name for intent, describe like onboarding notes, and constrain inputs with strict schemas.
- Consolidate chatty tools and return concise, human-readable, paginated outputs.
- Split irreversible actions into draft and human-approved execute steps.
- Lint tool definitions in CI and run a tool-selection eval after every change.
Try it
Take one tool you plan to give an agent and rewrite it using the seven principles. Run the lint harness on it and write five test requests for a tool-selection eval.