Model Context Protocol (MCP): Connect AI to Your Tools and DataBuilding MCP servers in Python and TypeScript · Lesson 8 of 18

Designing MCP tools and servers that LLMs use well

Article · 15 min · 9 min lecture

Video lecture

Designing MCP tools and servers that LLMs use well

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Designing tools models use well

  • Workflows over endpoints
  • Names, descriptions, schemas
  • Output and token budgets
  • A review checklist

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

An MCP server is a product whose user is a model

A server can be protocol-perfect and still fail in practice: the model picks the wrong tool, passes bad arguments, drowns in output, or burns through tokens. Everything from the agents course applies here, with some MCP-specific twists: your tools will be mixed with other servers' tools in hosts you don't control, and you cannot see the host's system prompt.

Principles for MCP tool design

  1. Design for workflows, not endpoints. Don't mirror your REST API one-to-one. Combine steps users actually perform: schedule_campaign(segment, channel, date) beats create_campaign + set_segment + set_channel + set_date.
  2. Unambiguous names, with a prefix when generic. Hosts may namespace tools per server, but not all do and models still see names. Prefer crm_search_contacts over search.
  3. Descriptions carry the usage rules. When to use, when not to, required order ("use find_campaigns first"), limits and what's returned. Use the server's instructions for cross-tool guidance.
  4. Tight input schemas. Enums, formats, ranges, required fields; describe every property. 2026-07-28 allows full JSON Schema 2020-12, but keep it simple enough for models and clients.
  5. Structured, compact output. Declare an outputSchema, return structuredContent plus a text version. Return identifiers the next tool needs. Offer a detail or fields parameter instead of always returning everything.
  6. Pagination and truncation, stated explicitly. Cursor-based pagination for lists, "showing 20 of 312; pass cursor=..." in the text.
  7. Actionable errors. Validation messages that tell the model how to fix the call.
  8. Idempotent writes. Accept a client-supplied key; the stateless transport re-issues broken requests.
  9. Honest annotations. Set readOnlyHint, destructiveHint, idempotentHint, openWorldHint accurately; good hosts use them for confirmation UX.
  10. Deterministic ordering and stable definitions. Changing descriptions constantly invalidates client caches and LLM prompt caches, and silently changing a tool's behavior after users approved it is exactly what "rug pull" attacks look like (module 5).

How many tools?

Every tool definition occupies context in every model call in the host. Dozens of verbose tools from several servers can consume thousands of tokens and degrade selection. Strategies:

  • Keep servers focused (one domain each) and tool counts modest.
  • Use toolsets or feature flags so admins enable only what a team needs.
  • Offer a small number of flexible tools (for example a query tool over a safe, read-only semantic layer) instead of dozens of narrow ones, when you can secure them.
  • Some hosts support tool search or deferred loading; don't rely on it universally.

Tokens are a design budget

Estimate the tokens your tools/list response adds (descriptions + schemas) and the typical size of each result. A 50 KB JSON result is roughly ten thousand or more tokens: expensive and distracting. Summaries, field selection and pagination fix it.

Worked example: redesigning a HubSpot-style CRM server for an agency

Version 1 mirrored the API: 38 tools (get_contact, get_contact_properties, list_associations, ...). In testing, models called 6–9 tools per question and often picked list_associations wrongly.

Version 2:

  • 9 tools organized around jobs: crm_find_contacts, crm_contact_brief (profile + recent activity + open deals in one call), crm_find_deals, crm_pipeline_summary, crm_log_activity, crm_create_task, crm_draft_email, crm_update_deal_stage, crm_list_owners.
  • crm_contact_brief returns a 20-line summary plus IDs; detail="full" for more.
  • Writes accept idempotency_key; crm_update_deal_stage carries destructiveHint: true so hosts confirm.

Result (measured on their 40-question eval, illustrative): average tool calls per question roughly halved and wrong-tool selections became rare.

Hands-on: a tool design review checklist

Run this against every tool before release:

[ ] Name is specific, verb_noun, prefixed if generic (e.g., crm_*)
[ ] Description: purpose, when to use, when NOT to use, prerequisites, limits, return shape
[ ] Every input property has a description or enum; required fields declared
[ ] outputSchema declared; structuredContent + text returned
[ ] Results capped; pagination cursor and "showing X of Y" message
[ ] Errors explain how to fix the call
[ ] Writes: idempotency key, accurate destructive/idempotent hints
[ ] No secrets or tokens accepted as tool arguments
[ ] Tool listed in deterministic order; description changes versioned in the changelog
[ ] Eval: 10+ realistic requests hit this tool correctly in at least two hosts

And a quick token estimate for your tool list (Python):

import json
from mcp import Client
from server import mcp          # your MCPServer instance

async def tool_list_size():
    async with Client(mcp) as client:
        tools = (await client.list_tools()).tools
        payload = json.dumps([t.model_dump(exclude_none=True) for t in tools])
        print(f"{len(tools)} tools, ~{len(payload) // 4} tokens (rough 4-chars-per-token estimate)")

The four-characters-per-token rule is only a rough heuristic; use your model provider's token-counting endpoint for precise numbers.

Pitfalls

  • One tool per API endpoint.
  • Generic names like search or get colliding across servers.
  • Returning raw vendor payloads.
  • Changing descriptions or behavior silently after release.

Measuring success

Run the same eval in at least two hosts (for example Claude and an IDE agent). Track correct-tool rate, calls per task, tokens per task and error rate; watch how they change as you edit descriptions.

Key takeaways

  • Design MCP tools around user workflows, not one-to-one API endpoints.
  • Your tools share context with other servers' tools, so use specific names and rich descriptions.
  • Return compact structured output with IDs, pagination and actionable errors.
  • Every tool definition costs tokens in every call; keep servers focused and tool counts modest.
  • Keep definitions stable and versioned; silent changes erode trust and caching.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Your MCP server exposes 38 tools mirroring every API endpoint and models make many calls per question. What is the best first change?
  2. Why prefix generic tool names such as search with a domain, like crm_search_contacts?
  3. A tool returns 50 KB of raw JSON per call. What is the main problem?

Put it into practice

Apply the design checklist to three tools you plan to ship, estimate your tool list's token cost, and rewrite the weakest tool.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.