Advanced Prompt Engineering · Prompting tools and agents · lesson 8 of 17 · 14 min
Tool definitions and tool-use prompts
Tools are prompts too
When you give a model tools (functions it can ask your code to run), every tool's name, description and parameter schema becomes part of the prompt. The model decides whether to call a tool, which one, and with what arguments, based almost entirely on those descriptions. Weak tool definitions are one of the most common causes of agents that call the wrong tool, loop, or make up arguments.
The mental model: each tool definition is a mini job description for a function. Write it for a capable colleague who has never seen your codebase.
Anatomy of a good tool definition
- A specific, unambiguous name.
search_orders_by_emailbeatssearch. If two tools could be confused, rename them until they cannot. - A description that says what, when and when not. What the tool does, when to use it, what it returns, its limits, and which tool to use instead in nearby cases.
- Parameters with types, enums, formats and descriptions. Say "ISO 8601 date, e.g. 2026-09-30" rather than "date". Use enums for fixed choices.
- Useful, compact results. Return what the model needs to decide the next step, with clear errors. A 5,000-line JSON dump wastes context and hides the signal.
Worked example: weak versus strong
Weak:
{"name": "lookup", "description": "Looks up data",
"input_schema": {"type": "object", "properties": {"q": {"type": "string"}}}}
Strong:
{
"name": "get_order_status",
"description": "Returns the current status, carrier and estimated delivery date for ONE order. Use when the customer gives an order number (format NW-123456). Do not use for refunds (use create_refund_request) or for listing a customer's orders (use list_orders_by_email). Returns an error object if the order is not found.",
"strict": true,
"input_schema": {
"type": "object",
"properties": {
"order_id": {"type": "string", "description": "Order number in the format NW- followed by 6 digits, e.g. NW-004512"}
},
"required": ["order_id"],
"additionalProperties": false
}
}
strict: true (supported by several providers under similar names) makes the API guarantee that tool arguments validate against the schema. It does not guarantee the arguments are correct, so your code still validates business rules.
Hands-on: a complete tool loop with the Claude API
import json, os, anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")
TOOLS = [{
"name": "get_order_status",
"description": "Returns status, carrier and ETA for ONE order (format NW-123456). "
"Not for refunds or listing orders. Returns {'error': ...} if not found.",
"strict": True,
"input_schema": {"type": "object",
"properties": {"order_id": {"type": "string", "description": "e.g. NW-004512"}},
"required": ["order_id"], "additionalProperties": False},
}]
ORDERS = {"NW-004512": {"status": "in transit", "carrier": "Aramex", "eta": "2026-10-02"}}
def run_tool(name, args):
if name == "get_order_status":
return ORDERS.get(args["order_id"], {"error": "order not found"})
return {"error": f"unknown tool {name}"}
def chat(user_text, max_turns=6):
messages = [{"role": "user", "content": user_text}]
for _ in range(max_turns):
resp = client.messages.create(model=MODEL, max_tokens=4096, tools=TOOLS,
system="You are a delivery support assistant for Northwind.",
messages=messages)
messages.append({"role": "assistant", "content": resp.content}) # keep full blocks
if resp.stop_reason != "tool_use":
return "".join(b.text for b in resp.content if b.type == "text")
results = []
for block in resp.content:
if block.type == "tool_use":
out = run_tool(block.name, block.input)
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": json.dumps(out), "is_error": "error" in out})
messages.append({"role": "user", "content": results}) # all results in ONE message
return "Sorry, I couldn't complete that. A colleague will follow up."
print(chat("Where is my order NW-004512?"))
Key details: append the assistant's full content blocks (not just text), return every tool result in a single user message when the model calls several tools in parallel, mark failures with is_error, and cap the number of turns. The SDKs also provide tool-runner helpers that manage this loop for you. The OpenAI Responses API and the Gemini API follow the same pattern with their own field names: the model emits a function call, your code runs it, and you send the output back referencing the call's ID.
Prompting how tools are used
Your system prompt should say how tools fit the job, briefly:
Use get_order_status whenever the customer gives an order number; do not
guess delivery dates. If a tool returns an error, tell the customer what
you could not find and ask for the correct order number. You can call
independent tools in parallel.
Avoid shouting ("You MUST ALWAYS call the tool"). Current models follow tool guidance closely and aggressive wording tends to cause unnecessary calls.
Many tools: keep the menu short
Every tool definition costs tokens on every request and adds choices. With dozens of tools, accuracy can fall. Options include grouping related operations into one well-designed tool with an action enum, loading tools on demand (some APIs offer tool search so rarely used tools are only loaded when needed), or routing requests to specialised agents with small toolsets. Connecting tools via the Model Context Protocol (MCP) makes integration easier, but the same design rules apply to MCP tool descriptions.
Pitfalls
- Overlapping tools with vague descriptions, causing random choices.
- Returning raw database rows or full HTML pages as tool results.
- Silent failures: returning an empty string instead of a clear error.
- Letting the model call side-effecting tools (refunds, emails) without an approval step.
How to measure tool quality
Build a tool-use eval: 30 to 50 requests with the expected tool and key arguments. Measure tool-selection accuracy, argument accuracy, unnecessary calls and task completion. When you change a description, re-run it; small wording changes can shift selection behaviour noticeably.
Video lecture: Tool definitions and tool-use prompts
Lecture coming soon · 11 chapters · about 8 minutes. Read the full transcript below.
- Tool definitions and tool-use prompts
- Tools are prompts
- Anatomy of a good tool
- Weak vs strong
- The tool loop
- Prompting tool use
- Many tools: keep the menu short
- Example 1: two weather tools
- Example 2: a retail support agent (illustrative)
- Pitfalls and measurement
- Recap
Lecture transcript
Tool definitions and tool-use prompts
An agent keeps calling the wrong tool. The team rewrites the system prompt three times and adds ALWAYS in capital letters. Nothing changes. The problem was never the system prompt. It was the tool descriptions. In this lecture you will learn that tools are prompts too: how to write tool names, descriptions and schemas a model uses correctly, how the tool-use loop works in practice, how to keep large toolsets manageable, and how to measure tool quality.
Tools are prompts
When you give a model tools, functions it can ask your code to run, every tool's name, description and parameter schema becomes part of the prompt. The model decides whether to call a tool, which one, and with what arguments, based almost entirely on those descriptions. So here is the mental model: each tool definition is a mini job description for a function. Write it for a capable colleague who has never seen your codebase.
Anatomy of a good tool
A good tool has four parts. A specific, unambiguous name: search orders by email beats search. If two tools could be confused, rename them until they cannot. A description that says what, when and when not: what it does, when to use it, what it returns, its limits, and which tool to use instead in nearby cases. Parameters with types, enums, formats and descriptions, such as an order number in the format N W dash followed by six digits. And compact results with clear errors, rather than a five-thousand-line JSON dump.
Weak vs strong
Compare two definitions. Weak: a tool named lookup, described as looks up data, with a single parameter called q. The model has to guess. Strong: a tool named get order status, described as returning the status, carrier and estimated delivery for one order, to be used when the customer gives an order number, not for refunds, which use create refund request, and not for listing orders, which use list orders by email, and returning an error object if the order is not found. With strict mode on, which several providers offer, the API guarantees the arguments validate against the schema. It does not guarantee they are correct, so your code still checks business rules.
The tool loop
Now the loop. You send the request with your tools. If the model wants a tool, the response stops with a tool use reason and contains tool use blocks. Your code runs each tool and sends back tool result blocks that reference the call's ID, marking failures with an error flag. Four details matter. Append the assistant's full content blocks to the history, not just the text. When the model calls several tools in parallel, return all the results together in a single message. Cap the number of turns. And remember that SDK tool-runner helpers can manage this loop for you. OpenAI's Responses API and the Gemini API follow the same pattern with their own field names.
Prompting tool use
Your system prompt should say how tools fit the job, briefly and calmly. For example: use get order status whenever the customer gives an order number, and do not guess delivery dates. If a tool returns an error, tell the customer what you could not find and ask for the correct number. You can call independent tools in parallel. Avoid shouting, like you MUST ALWAYS call the tool. Current models follow tool guidance closely, and aggressive wording tends to cause unnecessary calls.
Many tools: keep the menu short
Every tool definition costs tokens on every request and adds another choice. With dozens of tools, accuracy can drop. You have options. Group related operations into one well-designed tool with an action enum. Load tools on demand; some APIs now offer tool search, so rarely used tools are only loaded when needed. Or route requests to specialised agents with small toolsets. Connecting tools through the Model Context Protocol, MCP, makes integration easier, but the same design rules apply to MCP tool descriptions.
Example 1: two weather tools
A simple worked example. A weather assistant has two tools: get forecast and get current conditions. Their descriptions were: gets weather data, and returns weather. The model picked randomly. The fix takes one minute. Get forecast: returns the forecast for the next seven days for one city; use when the user asks about future days, like tomorrow or the weekend; not for right now. Get current conditions: returns the current temperature and conditions for one city; use for right now or today's current weather. And both take a city parameter with an example, like Lahore. Now ask, will it rain on Saturday in Lahore? The model calls the forecast tool every time.
Example 2: a retail support agent (illustrative)
Now a business scenario, with illustrative numbers. An online electronics retailer in Riyadh runs a support agent with twelve tools and about eight thousand conversations a month. Their tool-selection eval, fifty requests with the expected tool, shows seventy-four percent correct selection and many unnecessary calls. The main culprits: three overlapping order tools, and a product search that returned full product pages, around four thousand tokens each. The team merges the order tools into one, with an action enum for status, cancel and return; rewrites every description with what, when and when not; and makes product search return just the name, price, stock and key specs. On the same eval, correct selection rises to around ninety-four percent, and average tokens per conversation fall by roughly a third. They also move the cancel action behind an approval step in code. Illustrative numbers.
Pitfalls and measurement
Four pitfalls cause most tool trouble. Overlapping tools with vague descriptions, which leads to random choices. Raw database rows or full web pages returned as results, which floods the context. Silent failures, such as returning an empty string instead of a clear error. And side-effecting tools, like refunds or emails, without an approval step enforced in code. Then measure. Build a tool-use eval of thirty to fifty requests, each with the expected tool and key arguments. Track tool-selection accuracy, argument accuracy, unnecessary calls and task completion, and re-run it whenever you change a description.
Recap
To recap. Tool names, descriptions and schemas are prompts. Write each as a mini job description with what, when and when not, typed parameters and compact results. Run the loop correctly, keep toolsets small or load tools on demand, and gate side effects in code. Try this now: rewrite two of your tool definitions using the what, when and when not pattern, add enums and formats, then build a thirty-request tool-selection eval and compare accuracy before and after. Next: system prompts for agents and long-running tasks.
Key takeaways
- Tool names, descriptions and schemas are prompts: write them as mini job descriptions with what, when and when not.
- Use specific names, enums, formats and strict schemas, and return compact results with clear errors.
- In the loop, append full assistant blocks, return all parallel results together and cap turns.
- Keep toolsets small or load tools on demand, gate side effects in code, and evaluate tool selection.
Try it
Rewrite two of your tool definitions using the what/when/when-not pattern, add enums and formats, then build a 30-request tool-selection eval and compare accuracy before and after.