---
title: "Designing good tools for models | Optimize All Academy"
description: "Tools are an interface for a new kind of user Designing tools for a model is like designing an API for a capable colleague who reads the documentation…"
url: https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/designing-good-tools
updated: 2026-10-05
---

Latest AI Techniques: RAG, Tool Use, Agents & MCP · Tool and function calling · lesson 7 of 20 · 11 min

# Designing good tools for models

## Tools are an interface for a new kind of user

Designing tools for a model is like designing an API for a capable colleague who reads the documentation literally and cannot ask you questions. Many agent failures that look like "the model is not smart enough" are actually tool design problems.

## Principles

**1. Fewer, higher-level tools.** Instead of `list_customers`, `get_customer`, `list_orders` and `get_order`, consider `find_customer_orders(customer_email, status)` if that is what the workflow needs. Each tool call costs a round trip and a chance to go wrong. Match tools to tasks, not to your database tables.

**2. Clear, distinct names and boundaries.** If two tools overlap (`search_docs` and `search_knowledge`), the model will choose inconsistently. Make each tool's purpose unmistakable, or merge them.

**3. Inputs that are hard to get wrong.** Use enums for fixed choices, explicit formats for dates (ISO 8601) and IDs, sensible defaults, and a minimal set of required fields.

**4. Outputs designed for reading by a model.** Return meaningful fields, human-readable names alongside IDs, and a sensible amount of data. Offer pagination or a `detail_level` parameter (summary vs full) rather than dumping everything.

**5. Errors that teach.** "Date must be in YYYY-MM-DD format; you sent 29/09/2026" lets the model fix its call on the next try.

**6. Safe by construction.** Separate read tools from write tools. Make destructive operations explicit (`cancel_order` rather than a generic `update_order` with a status field), so they can be gated with confirmation.

## A before-and-after

Before:

```json
{"name": "query", "description": "Run a query",
 "input_schema": {"type": "object", "properties": {"q": {"type": "string"}}}}
```

After:

```json
{
  "name": "search_help_articles",
  "description": "Search published help-centre articles by meaning and keywords. Use for how-to and policy questions. Does NOT include order or account data; use get_order_status for those. Returns up to 5 results with title, url, snippet and last_updated.",
  "input_schema": {
    "type": "object",
    "properties": {
      "query": {"type": "string", "description": "The user's question rephrased as a standalone search query"},
      "product": {"type": "string", "enum": ["web", "ios", "android"], "description": "Filter by product if the user mentioned one"}
    },
    "required": ["query"]
  }
}
```

The second version tells the model what the tool covers, what it does not, what comes back, and how to fill the inputs.

## Testing tools with the model

Write tests the same way you would for prompts:

- A set of user requests, each with the expected tool (or no tool) and key arguments.
- Measure **tool selection accuracy** (right tool chosen), **argument accuracy** (valid, correct values) and **unnecessary calls**.
- Include near-miss cases: questions that sound like they need a tool but do not, and vice versa.

Read transcripts. When the model misuses a tool, ask first: "Would a new human colleague have misunderstood this description too?" Usually, the fix is in the description or the tool's shape.

## How many tools?

There is no fixed limit, but selection quality tends to fall as the number of similar tools grows, and every definition consumes context. Strategies:

- Group tools by task and expose only the relevant group (for example via a router step).
- Use tool search or dynamic loading if your platform supports it, so the model discovers tools as needed.
- Consolidate rarely used tools.

## Worked example: a sales assistant

A small agency connects an assistant to its CRM. Version 1 exposes twelve thin CRM endpoints. The model often calls `get_contact` repeatedly and forgets to check deal stage. Version 2 has three tools: `find_account(name_or_domain)` returning account, main contacts and open deals in one summary; `log_activity(account_id, type, notes)`; and `draft_follow_up(account_id, goal)`, which drafts but never sends. Selection errors drop, runs are shorter, and nothing is sent without a human.

## Failure modes

- Tools that mirror internal APIs one-to-one.
- Tool results with cryptic codes and no labels.
- Write tools without confirmation or audit logs.
- Changing a tool's behaviour without updating its description.

## Hands-on: a tool-selection test harness

Treat tool definitions like code: test them. The harness below sends each test request to the model with your tool set, captures the first tool call (or none), and scores tool selection and key arguments. It uses the Claude API; the same idea works with any provider.

```python
import json, os
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
TOOLS = json.load(open("tools.json"))      # your tool definitions, versioned with your code

CASES = [
    {"ask": "How do I reset my password on iOS?", "tool": "search_help_articles", "args": {"product": "ios"}},
    {"ask": "Where is order A55120?",             "tool": "get_order_status",     "args": {"order_id": "A55120"}},
    {"ask": "Thanks, that's all!",                "tool": None},  # near-miss: no tool needed
    {"ask": "Where is my order?",                 "tool": None},  # should ask for the ID, not guess
]

def first_call(ask):
    resp = client.messages.create(model=MODEL, max_tokens=1000, tools=TOOLS,
                                  messages=[{"role": "user", "content": ask}])
    for block in resp.content:
        if block.type == "tool_use":
            return block.name, block.input
    return None, {}

score = {"selection": 0, "args": 0, "n": len(CASES)}
for c in CASES:
    name, args = first_call(c["ask"])
    ok_sel = name == c["tool"]
    ok_args = all(args.get(k) == v for k, v in c.get("args", {}).items())
    score["selection"] += ok_sel
    score["args"] += ok_sel and ok_args
    if not (ok_sel and ok_args):
        print("MISS:", c["ask"], "->", name, args)
print(score)
```

Run it several times per case (model behaviour varies), keep a changelog of description edits, and re-run the suite whenever you change a description, add a tool or switch model. Near-miss cases such as "Where is my order?" (no ID) are the most valuable: they catch hallucinated arguments.

## Designing for large tool sets

When an agent needs access to dozens or hundreds of tools (common once you connect several MCP servers), loading every definition into every request wastes context and hurts selection. Current options include:

- **Tool search / deferred loading:** some APIs let you mark tools as deferred so the model searches for them only when needed; Anthropic's API, for example, offers regex and BM25 tool-search tools for this.
- **Routing:** a cheap first step picks a toolset (billing, orders, marketing) before the main call.
- **Code-mode tools:** instead of 30 thin tools, expose a small, sandboxed code-execution environment with a well-documented client library, so the model writes a short script that calls several operations and returns only the summary. This can cut round trips and context dramatically, but it raises the security stakes: sandbox it properly.

Whichever you use, the design principles in this lesson still decide success.

## Going further

Version tool definitions alongside prompts, and include them in your evaluation runs. A tool description change is a behaviour change. For tools exposed to many AI applications at once, a standard interface such as the Model Context Protocol (module 5) lets you define them once and reuse them across compatible clients.

## Video lecture: Designing good tools for models

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Designing tools for models
2. Why it matters
3. Principles 1–2
4. Principles 3–4
5. Principles 5–6
6. Simple example: rename and describe
7. Worked example: agency CRM
8. Business example (illustrative)
9. Test your tools
10. Scaling tool sets
11. Common mistakes
12. How you'll know it's working
13. Watch me do it: selection harness
14. Recap
15. Try this now (30 minutes)

## Lecture transcript

### Designing tools for models

When an AI agent misbehaves, the easy conclusion is that the model isn't smart enough. Surprisingly often, the real culprit is the tools you gave it: vague names, overlapping purposes, cryptic outputs and errors that say nothing. In this lesson you'll learn six principles for designing tools that models use reliably, see a before and after, and set up a small test harness so tool quality becomes something you measure rather than hope for.

### Why it matters

Why focus on tool design? Because it's usually the cheapest fix available. Swapping to a bigger model costs money on every call forever. Rewriting a confusing tool description costs ten minutes once. Here's an analogy. Give a new employee a remote control with thirty unlabelled buttons and they'll press the wrong ones. Give them three clearly labelled buttons for the three jobs they actually do, and they'll get it right first time.

### Principles 1–2

Think of tool design as writing an API for a capable colleague who reads documentation literally and can't ask you questions. Principle one: fewer, higher-level tools. Instead of list customers, get customer, list orders and get order, offer find customer orders if that's what the workflow needs. Every call is a round trip and a chance to go wrong. Principle two: clear, distinct names and boundaries. If search docs and search knowledge overlap, the model will choose inconsistently. Merge them, or state exactly what each one covers.

### Principles 3–4

Principle three: inputs that are hard to get wrong. Use enums for fixed choices, explicit formats like ISO dates, sensible defaults, and as few required fields as possible. Principle four: outputs designed for a model to read. Return meaningful field names, human-readable names alongside IDs, and a sensible amount of data, with pagination or a detail level option instead of dumping everything.

### Principles 5–6

Principle five: errors that teach. Date must be in year-month-day format; you sent twenty-nine slash nine lets the model fix its next call. Error four hundred teaches nothing. Principle six: safe by construction. Separate read tools from write tools, and make destructive operations explicit. A dedicated cancel order tool can be gated with confirmation and audit. A generic update order tool with a status field hides the danger inside a parameter.

### Simple example: rename and describe

A simple example first. Your tool is called get data, with one input called q. The model has to guess what it does and what q means. Rename it search help articles, describe what it covers and doesn't, say it returns up to five results with titles and dates, and rename q to query with a description. Same code underneath, but now the model calls it at the right time with sensible input. Nothing about the model changed.

### Worked example: agency CRM

Here's a real-world shaped example. A small agency connects an assistant to its CRM with twelve thin endpoints. The model keeps calling get contact repeatedly and forgets to check the deal stage. Version two has three tools. Find account returns the account, main contacts and open deals in one summary. Log activity records a note. Draft follow up drafts an email but never sends. Selection errors drop, runs get shorter, and nothing reaches a client without a human.

### Business example (illustrative)

More detail on the agency CRM case, illustrative. With twelve thin tools, a typical account lookup took about six tool calls and the model forgot the deal stage in roughly one run in five. With three task-level tools, the same job took two calls, and the missed-stage problem disappeared from their test set. Runs got faster and cheaper, and because drafting never sends, the account managers kept full control of client communication.

### Test your tools

Now make it measurable. Build a small set of user requests, each with the expected tool, or no tool, and the key arguments. Include near-misses: a thank-you message that needs no tool, and where is my order with no ID, where the right behaviour is to ask. Score selection accuracy and argument accuracy, run each case several times, and read the transcripts of misses. Ask yourself: would a new human colleague have misunderstood this description too? Usually the fix is in the description.

### Scaling tool sets

As tool counts grow, especially after connecting several MCP servers, loading every definition into every request wastes context and hurts selection. Current answers include tool search, where tools are deferred until the model looks for them; a routing step that picks a toolset first; and code-mode, where the model writes a short sandboxed script against a documented library instead of calling thirty tiny tools. The lesson has a runnable harness that scores tool selection with the Claude API, plus notes on each of these patterns.

### Common mistakes

Common mistakes. Mirroring your internal API one-to-one, so the model makes ten calls where one would do. Overlapping tools with near-identical names. Returning cryptic codes without labels. Changing a tool's behaviour without updating its description. Hiding dangerous actions inside generic update tools. And never testing tool selection, so you only discover problems when a customer does.

### How you'll know it's working

How will you know your tool design improved? Selection accuracy goes up on your test set, including near-miss cases. Average tool calls per task goes down, because task-level tools do more per call. Invalid-argument errors fall. And, most tellingly, when you read transcripts, the model's choices look like what a sensible new colleague would do with the same descriptions.

### Watch me do it: selection harness

Watch me do it. I open the selection harness. First, I load my tool definitions from tools dot json, so they're versioned with the code. Next, the cases list: each has a request, the expected tool or none, and key arguments. Notice the two near-misses: a thank-you message and where is my order without an ID. First call sends one request with the tools and returns the first tool use block's name and input, or none. Then the loop compares the tool name and arguments with expectations, counts passes, and prints every miss. I run it. The first run shows one miss: where is my order triggers get order status with an invented ID. I update the description to say, if no order ID was given, ask the customer for it. I run it three more times, because behaviour varies, and all four cases pass each time. I note the description change in the changelog.

### Recap

To recap: many agent failures are tool design failures. Prefer fewer, task-level tools with distinct boundaries, hard-to-misuse inputs, readable outputs, teaching errors and explicit destructive actions. Test selection and arguments with near-miss cases, and version tool descriptions with your code, because a description change is a behaviour change. Your next step: redesign three endpoints or spreadsheet actions you'd expose to an AI as task-level tools. Next up, agents.

### Try this now (30 minutes)

Try this now. Take three endpoints, spreadsheet actions or buttons you'd expose to an AI assistant. Redesign them as task-level tools: better names, descriptions with when to use and when not to, enums for fixed choices, and error messages that explain how to fix a bad call. Then write four test requests, including one where no tool should be called, and run them through the selection harness from the lesson.

## Key takeaways

- Tool design, not model intelligence, causes many agent failures.
- Prefer fewer, task-level tools with distinct boundaries, hard-to-misuse inputs and model-readable outputs.
- Return errors that teach, separate read and write tools, and make destructive actions explicit.
- Test tool selection and argument accuracy with near-miss cases and read transcripts.

## Try it

Take three API endpoints or spreadsheet actions you would expose to an AI assistant and redesign them as task-level tools with improved descriptions, enums and error messages.

- [Previous: Function calling fundamentals](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/function-calling-fundamentals)
- [Next: Structured outputs and strict tool use](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/structured-outputs-and-strict-tools)
- [All lessons of Latest AI Techniques: RAG, Tool Use, Agents & MCP](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp)
