---
title: "Designing MCP tools and servers that LLMs use well"
description: "An MCP server is a product whose user is a model A server can be protocol-perfect and still fail in practice: the model picks the wrong tool, passes bad…"
url: https://optimizeall.com/learn/model-context-protocol-mcp/designing-mcp-tools-for-llms
updated: 2026-10-05
---

Model Context Protocol (MCP): Connect AI to Your Tools and Data · Building MCP servers in Python and TypeScript · lesson 8 of 18 · 15 min

# Designing MCP tools and servers that LLMs use well

## An MCP server is a product whose user is a model

A server can be protocol-perfect and still fail in practice: the model picks the wrong tool, passes bad arguments, drowns in output, or burns through tokens. Everything from the agents course applies here, with some MCP-specific twists: your tools will be mixed with **other servers' tools** in hosts you don't control, and you cannot see the host's system prompt.

## Principles for MCP tool design

1. **Design for workflows, not endpoints.** Don't mirror your REST API one-to-one. Combine steps users actually perform: `schedule_campaign(segment, channel, date)` beats `create_campaign` + `set_segment` + `set_channel` + `set_date`.
2. **Unambiguous names, with a prefix when generic.** Hosts may namespace tools per server, but not all do and models still see names. Prefer `crm_search_contacts` over `search`.
3. **Descriptions carry the usage rules.** When to use, when not to, required order ("use `find_campaigns` first"), limits and what's returned. Use the server's `instructions` for cross-tool guidance.
4. **Tight input schemas.** Enums, formats, ranges, required fields; describe every property. 2026-07-28 allows full JSON Schema 2020-12, but keep it simple enough for models and clients.
5. **Structured, compact output.** Declare an `outputSchema`, return `structuredContent` plus a text version. Return identifiers the next tool needs. Offer a `detail` or `fields` parameter instead of always returning everything.
6. **Pagination and truncation, stated explicitly.** Cursor-based pagination for lists, "showing 20 of 312; pass cursor=..." in the text.
7. **Actionable errors.** Validation messages that tell the model how to fix the call.
8. **Idempotent writes.** Accept a client-supplied key; the stateless transport re-issues broken requests.
9. **Honest annotations.** Set `readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint` accurately; good hosts use them for confirmation UX.
10. **Deterministic ordering and stable definitions.** Changing descriptions constantly invalidates client caches and LLM prompt caches, and silently changing a tool's behavior after users approved it is exactly what "rug pull" attacks look like (module 5).

## How many tools?

Every tool definition occupies context in every model call in the host. Dozens of verbose tools from several servers can consume thousands of tokens and degrade selection. Strategies:

- Keep servers focused (one domain each) and tool counts modest.
- Use toolsets or feature flags so admins enable only what a team needs.
- Offer a small number of flexible tools (for example a `query` tool over a safe, read-only semantic layer) instead of dozens of narrow ones, when you can secure them.
- Some hosts support tool search or deferred loading; don't rely on it universally.

## Tokens are a design budget

Estimate the tokens your `tools/list` response adds (descriptions + schemas) and the typical size of each result. A 50 KB JSON result is roughly ten thousand or more tokens: expensive and distracting. Summaries, field selection and pagination fix it.

## Worked example: redesigning a HubSpot-style CRM server for an agency

Version 1 mirrored the API: 38 tools (`get_contact`, `get_contact_properties`, `list_associations`, ...). In testing, models called 6–9 tools per question and often picked `list_associations` wrongly.

Version 2:

- 9 tools organized around jobs: `crm_find_contacts`, `crm_contact_brief` (profile + recent activity + open deals in one call), `crm_find_deals`, `crm_pipeline_summary`, `crm_log_activity`, `crm_create_task`, `crm_draft_email`, `crm_update_deal_stage`, `crm_list_owners`.
- `crm_contact_brief` returns a 20-line summary plus IDs; `detail="full"` for more.
- Writes accept `idempotency_key`; `crm_update_deal_stage` carries `destructiveHint: true` so hosts confirm.

Result (measured on their 40-question eval, illustrative): average tool calls per question roughly halved and wrong-tool selections became rare.

## Hands-on: a tool design review checklist

Run this against every tool before release:

```text
[ ] Name is specific, verb_noun, prefixed if generic (e.g., crm_*)
[ ] Description: purpose, when to use, when NOT to use, prerequisites, limits, return shape
[ ] Every input property has a description or enum; required fields declared
[ ] outputSchema declared; structuredContent + text returned
[ ] Results capped; pagination cursor and "showing X of Y" message
[ ] Errors explain how to fix the call
[ ] Writes: idempotency key, accurate destructive/idempotent hints
[ ] No secrets or tokens accepted as tool arguments
[ ] Tool listed in deterministic order; description changes versioned in the changelog
[ ] Eval: 10+ realistic requests hit this tool correctly in at least two hosts
```

And a quick token estimate for your tool list (Python):

```python
import json
from mcp import Client
from server import mcp          # your MCPServer instance

async def tool_list_size():
    async with Client(mcp) as client:
        tools = (await client.list_tools()).tools
        payload = json.dumps([t.model_dump(exclude_none=True) for t in tools])
        print(f"{len(tools)} tools, ~{len(payload) // 4} tokens (rough 4-chars-per-token estimate)")
```

The four-characters-per-token rule is only a rough heuristic; use your model provider's token-counting endpoint for precise numbers.

## Pitfalls

- One tool per API endpoint.
- Generic names like `search` or `get` colliding across servers.
- Returning raw vendor payloads.
- Changing descriptions or behavior silently after release.

## Measuring success

Run the same eval in at least two hosts (for example Claude and an IDE agent). Track correct-tool rate, calls per task, tokens per task and error rate; watch how they change as you edit descriptions.

## Video lecture: Designing MCP tools and servers that LLMs use well

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Designing tools models use well
2. Why design matters
3. Workflows, names, descriptions
4. Inputs and outputs
5. Writes and stability
6. Simple example: two 'search' tools
7. How many tools?
8. Example: CRM server redesign
9. Hands-on: review checklist
10. Pitfalls and metrics
11. Breaking changes
12. Deeper: CRM redesign results (illustrative)
13. Watch me do it: checklist on crm_update_deal_stage
14. Try this now
15. Recap

## Lecture transcript

### Designing tools models use well

A server can follow the spec perfectly and still fail in the real world. The model picks the wrong tool, sends bad arguments, drowns in output, or burns through tokens. In this lesson you'll learn how to design MCP tools that models use well, even when they're mixed with other servers' tools in hosts you don't control.

### Why design matters

Why does design matter once the protocol works? Because in real hosts your server competes for the model's attention with dozens of other tools you've never seen. Think of a supermarket shelf. Products with clear labels, obvious purpose and sensible packaging get picked. Products with vague names get ignored, or worse, picked by mistake. Your tool names, descriptions and output sizes are your packaging. The model is a very fast shopper who reads every label but can't ask the staff.

### Workflows, names, descriptions

Start with the big idea: design for workflows, not endpoints. Don't mirror your REST API one to one. Combine the steps people actually do. Schedule campaign with a segment, channel and date beats four separate calls to create a campaign and set each field. Use specific names, prefixed when generic, like crm search contacts instead of search, because your tools will sit next to other servers' tools. And put usage rules in descriptions: when to use a tool, when not to, what to call first, the limits, and what comes back.

### Inputs and outputs

Next, tight inputs and compact outputs. Use enums, formats, ranges and required fields, and describe every property. Declare an output schema and return structured content plus a text version. Include the ids the next tool will need. Offer a detail or fields option instead of always returning everything. Paginate with cursors and say so plainly, like showing twenty of three hundred and twelve. And make errors actionable, telling the model how to fix the call.

### Writes and stability

For writes, accept an idempotency key, because the stateless transport re issues broken requests. Set annotations honestly, read only, destructive, idempotent and open world, so good hosts can ask for confirmation at the right moments. Return tools in a deterministic order, and keep definitions stable. Constantly changing descriptions breaks caches, and silently changing a tool's behavior after people approved it is exactly what a rug pull attack looks like.

### Simple example: two 'search' tools

A simple example. Two servers are connected to the same host, and both have a tool named search. One searches support tickets, the other searches product docs. A user asks, has anyone reported the login bug. The model picks the docs search half the time. The fix takes two minutes: rename them support search tickets and docs search articles, and give each a description saying when to use it and when not to. Selection errors drop, and nothing about your code changed.

### How many tools?

How many tools should you have? Remember, every tool definition takes up context in every model call. Dozens of verbose tools from several servers can eat thousands of tokens and make selection worse. Keep servers focused on one domain, use toolsets or flags so admins enable only what a team needs, and consider a few flexible, well secured tools instead of many narrow ones. Some hosts support tool search or deferred loading, but don't count on it everywhere.

### Example: CRM server redesign

A worked example. An agency's CRM server started as thirty eight tools mirroring the vendor API. In testing, models made six to nine calls per question and often chose the wrong one. Version two has nine tools built around jobs, like find contacts and contact brief, which returns the profile, recent activity and open deals in one call, as a twenty line summary with ids. Writes take idempotency keys, and updating a deal stage carries a destructive hint so hosts confirm. On their evaluation set, calls per question roughly halved.

### Hands-on: review checklist

The lesson includes a ten point review checklist. Run it for every tool: specific name, a description with purpose, when to use and when not to, prerequisites, limits and return shape, described inputs, an output schema, capped results with pagination, fixable errors, idempotent writes with honest hints, no secrets as arguments, stable ordering and versioned changes, and an evaluation in at least two hosts. There's also a small script to estimate your tool list's token cost. It's a rough rule of thumb, so use your provider's token counter for precise numbers.

### Pitfalls and metrics

Watch for the classics: one tool per API endpoint, generic names that collide, raw vendor payloads, and silent changes after release. Measure by running the same evaluation in at least two hosts, say Claude and an IDE agent, and track correct tool rate, calls per task, tokens per task and error rate as you edit descriptions.

### Breaking changes

How do you handle breaking changes to a tool? Treat tools like a public API. Add new tools or optional parameters rather than changing the meaning of existing ones. If you must break something, publish a new tool name, keep the old one working for a transition period with a description that points to the new one, announce it in your changelog, and remove it only after checking usage. Hosts and agents that depend on you will thank you.

### Deeper: CRM redesign results (illustrative)

Let's deepen the agency CRM redesign, with illustrative numbers from their forty question evaluation. Version one: thirty eight tools, about seven calls per question, and wrong tool selections in roughly one question in four. Version two: nine tools, about three calls per question, and wrong selections rare. The tool list shrank from several thousand tokens to well under half that, which also reduced cost on every call in every host. Account managers noticed the difference as speed: answers arrived noticeably faster. The team kept the evaluation running nightly, and when someone later shortened the contact brief description, accuracy dipped, and they restored it the next morning.

### Watch me do it: checklist on crm_update_deal_stage

Watch me do it. Let's run the design checklist on one real tool, crm update deal stage. Name: specific and prefixed, check. Description: I add, use after crm find deals, only for moving an existing deal between stages, not for creating deals; returns the updated stage. Inputs: deal id described as coming from crm find deals, stage as an enum of the six stages, and an idempotency key. Output schema: deal id, old stage, new stage. Results are tiny, so no pagination needed. Errors: unknown deal says to search first; invalid transition says which stages are allowed next. Annotations: not read only, destructive true, idempotent true. No secrets in arguments. Then I run the token estimate script: nine tools, about eighteen hundred tokens by the rough rule, then I confirm with the provider's token counter. Finally, ten realistic requests in two hosts. Checklist done.

### Try this now

Try this now. Take three tools you plan to ship and run the ten point checklist from the lesson on each. Estimate how many tokens your whole tool list adds to every model call. Then rewrite the weakest tool: consolidate it with a neighbor, sharpen its description, or cut its output down to what the next step needs. Test the before and after versions with five realistic requests in a real host, and keep the numbers.

### Recap

To recap: your server's user is a model working alongside other servers' tools. Design around workflows, name and describe precisely, return compact structured output, make writes idempotent, keep definitions stable, and budget tokens. Your next step: run the checklist on three tools you plan to ship, estimate your tool list's token cost, and rewrite the weakest tool.

## Key takeaways

- Design MCP tools around user workflows, not one-to-one API endpoints.
- Your tools share context with other servers' tools, so use specific names and rich descriptions.
- Return compact structured output with IDs, pagination and actionable errors.
- Every tool definition costs tokens in every call; keep servers focused and tool counts modest.
- Keep definitions stable and versioned; silent changes erode trust and caching.

## Try it

Apply the design checklist to three tools you plan to ship, estimate your tool list's token cost, and rewrite the weakest tool.

- [Previous: Build an MCP server in TypeScript with the official SDK](https://optimizeall.com/learn/model-context-protocol-mcp/typescript-server-with-official-sdk)
- [Next: Connecting MCP servers to Claude, ChatGPT, IDEs and agents](https://optimizeall.com/learn/model-context-protocol-mcp/connecting-servers-to-ai-hosts)
- [All lessons of Model Context Protocol (MCP): Connect AI to Your Tools and Data](https://optimizeall.com/learn/model-context-protocol-mcp)
