Model Context Protocol (MCP): Connect AI to Your Tools and DataMCP in production · Lesson 13 of 18
Testing and debugging MCP servers with the Inspector and automated tests
Video lecture
Testing and debugging MCP servers with the Inspector and automated tests
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Testing and debugging
An MCP server has more ways to fail than a normal API. There's the protocol, the host's configuration, the model's choices and security. In this lesson you'll build a layered testing strategy, learn the MCP Inspector, write automated protocol and security tests, and debug the most common failures fast.
0:21 Why testing matters
Why so much emphasis on testing? Because an MCP server has an unusual user: a model that interprets your descriptions, in hosts you don't control, over protocol versions that are changing. Traditional tests check your code. They don't tell you whether a model will actually pick your tool, or whether a host will even launch your server. Think of a pilot's pre flight checklist: some checks are mechanical, some are instruments, some are a walk around the plane. You need all of them.
0:57 Six test layers
Think in six layers. Unit tests check the business logic behind each tool. Protocol tests call tools through an MCP client to check schemas, structured output and errors. The Inspector lets you explore interactively. Host tests confirm it works in the real apps your users run. Model behavior evals check whether models choose the right tools with the right arguments. And security tests check authentication, scopes and injection cases.
1:27 The MCP Inspector
The MCP Inspector is the official interactive tool for servers. It connects over standard I O or HTTP, lists your tools, resources and prompts, lets you call tools with any arguments, and shows the raw JSON RPC messages and server logs. With the Python SDK, mcp dev launches your server inside it. For any server, run the inspector package with npx and point it at your command or URL. Check that every schema field has a description, structured results match the output schema, errors come back as helpful error results, pagination works, and nothing leaks secrets.
2:09 Automated protocol tests
Then automate. In Python version two, a client connects directly to your server object in memory. The lesson's tests check that tools are listed in a stable order, that every input property has a description or an enum, and that an unknown customer returns an error result whose message tells the model to search first. In TypeScript, drive your HTTP handler in process with a transport that calls the handler's fetch, and assert on structured content and the error flag.
2:44 Simple debugging example
A simple debugging example. Your server shows as failed in Claude Desktop. In the terminal it runs fine. You open the host's log for your server and see, module not found. The host launched Python without your virtual environment, because it doesn't read your shell settings. The fix: use the absolute path to uv with the with flag, or the absolute path to your environment's Python. Restart the host fully. Five minutes, because you checked the log instead of guessing.
3:19 Debugging playbook
Now a debugging playbook. Server fails in the host but works in a terminal? Usually a relative path, a missing environment variable, or the host can't find uv or node. Disconnects right after start? Something printed to standard output. Four oh one loops on a remote server? Check the metadata URL, the token's audience and the server clock. Tool never chosen? The name or description is vague or overlaps another tool. Invalid arguments? The schema is too loose. And works on old clients but not new ones? A protocol version mismatch.
3:59 Evals and security tests
Model behavior evals catch problems no unit test can. Write twenty to fifty realistic requests per server with the tools and key arguments you expect. Run them through your client bridge against at least two models, and score tool selection accuracy, argument validity and calls per task. Re run after every description or schema change. Security tests should be automated too: no token gives four oh one, wrong audience gives four oh one, a read scope calling a write tool gives four oh three with a scope challenge, and traversal and SSRF inputs are rejected.
4:40 Example: marketplace orders server
A worked example. A Lahore marketplace ships an orders MCP server used by support staff in Claude and by an internal bot. Its CI runs unit tests, in memory protocol tests, security tests against staging with a test identity provider, and a nightly eval of forty support questions across two models. One pull request shortened a tool description to tidy it up, and the nightly eval showed tool selection accuracy dropping noticeably. They reverted it before release. That's the value of evals.
5:16 Pitfalls and metrics
Watch for four pitfalls: testing in only one host, no negative tests for auth and scopes, treating the Inspector as your test suite when it's for exploration, and not testing against both old and new protocol versions during the transition. Track protocol test coverage per tool, eval accuracy per model, security test pass rate and how long it takes to diagnose host connection issues.
5:44 Which hosts to test?
Should you test against every host? Aim for the ones your users actually run. Pick two or three priority hosts, run a short manual script in each before every release, and automate what you can with the Inspector's command line mode and your own client. Keep a small compatibility table in your docs, noting which features each host supports, so users know what to expect and your support team can answer quickly.
6:15 Deeper: the description regression (illustrative)
Let's deepen the Lahore marketplace's CI story. Their nightly model eval ran forty support questions against two models. One evening a well meaning pull request shortened get order status's description from four sentences to one, illustrative numbers again. The next morning, tool selection accuracy had dropped noticeably, mostly on questions like, where's my refund, which models now routed to a different tool. The developer restored the when to use sentence and accuracy recovered. After that, the team added a rule: any change to a tool name or description must include eval results in the pull request. Descriptions became reviewed like code, because they behave like code.
7:01 Watch me do it: protocol tests
Watch me do it. Let's run the automated protocol tests. The anyio backend fixture picks asyncio. The client fixture opens an in memory client on the server with raise exceptions on, and yields it. Test one lists tools twice and asserts the names match, so the order is stable, and that find campaigns is present. Test two loops over every tool's input schema and asserts each property has a description or an enum, which catches undocumented parameters before a model ever sees them. Test three calls get customer with a bad id, asserts the result is an error, and checks that the message mentions search customers, so the model gets a useful hint. I run pytest: three green tests in under a second, with no network, no subprocess and no API keys. Then I run the same server under the Inspector to eyeball schemas, and finally connect a host.
8:06 Try this now
Try this now. Add three protocol tests to your server: stable tool order, every property described, and one error path with helpful text. Add three security tests against a staging deployment: no token, wrong audience, and a read token calling a write tool. Then write ten realistic requests with the tool you expect for each, and run them through a model. Put all of it in continuous integration, with the eval running nightly.
8:38 Recap
To recap: test in six layers, explore with the Inspector, automate protocol and security tests, evaluate model behavior, and use the debugging playbook when hosts misbehave. Your next step: add three protocol tests and three security tests to your server's CI, and write ten realistic eval requests with the tools you expect.
A layered testing strategy
- Unit tests for the business logic behind each tool (pure functions, no MCP).
- Protocol tests that call tools through an MCP client in memory: schemas, structured output, error paths.
- Interactive exploration with the MCP Inspector.
- Host tests in the real clients your users run.
- Model-behavior evals: do models pick the right tools with the right arguments?
- Security tests: auth, scopes, injection and path/SSRF cases.
The MCP Inspector
The Inspector is the official interactive developer tool for MCP servers. It connects over stdio or Streamable HTTP, lists tools, resources and prompts, lets you call tools with custom arguments, and shows the raw JSON-RPC traffic and server logs.
- Python SDK:
uv run mcp dev server.pylaunches your server under the Inspector. - Any server:
npx @modelcontextprotocol/inspectorthen choose the transport and command/URL, or pass the launch command directly, e.g.npx @modelcontextprotocol/inspector node dist/server.js. - It also has a CLI mode useful for scripts and CI; check its README for current flags.
What to check in the Inspector: every tool's schema renders sensibly (descriptions on every field), structured results match the output schema, errors come back as isError with helpful text, pagination works, and nothing leaks secrets.
Protocol tests in code
Python (SDK v2) connects a Client directly to your server object:
import pytest
from mcp import Client
from server import mcp
@pytest.fixture
def anyio_backend():
return "asyncio"
@pytest.fixture
async def client():
async with Client(mcp, raise_exceptions=True) as c:
yield c
@pytest.mark.anyio
async def test_tools_listed_in_stable_order(client):
names1 = [t.name for t in (await client.list_tools()).tools]
names2 = [t.name for t in (await client.list_tools()).tools]
assert names1 == names2 and "find_campaigns" in names1
@pytest.mark.anyio
async def test_every_input_property_is_described(client):
for tool in (await client.list_tools()).tools:
for prop, spec in tool.input_schema.get("properties", {}).items():
assert "description" in spec or "enum" in spec, f"{tool.name}.{prop} lacks a description"
@pytest.mark.anyio
async def test_unknown_customer_is_tool_error(client):
result = await client.call_tool("get_customer", {"customer_id": "nope"})
assert result.is_error
assert "search_customers" in result.content[0].text # the error tells the model what to doTypeScript (SDK v2) can drive the HTTP handler in-process: a StreamableHTTPClientTransport whose fetch calls handler.fetch, then client.callTool(...) and assertions on structuredContent / isError. Close the client and handler.close() between tests.
Debugging common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Host shows server as failed; works in terminal | Relative path, missing env var, host can't find uv/node | Absolute paths; env in host config; absolute path to runtime |
| Connection drops right after start (stdio) | Something wrote to stdout | Log to stderr; remove prints |
| 401 loops on a remote server | Metadata URL wrong, audience mismatch, clock skew | Check Protected Resource Metadata, token aud, server time |
| Tool never chosen by the model | Vague name/description, overlaps with another tool | Rewrite description; eval in two hosts |
| Tool chosen but args invalid | Schema too loose or undocumented | Enums, formats, descriptions |
| Results truncated or huge | No pagination | Cursor pagination, summaries |
| Works on 2025 clients, fails on new ones (or vice versa) | Protocol version mismatch | Upgrade SDK; test both eras |
Server logs: Claude Desktop writes per-server logs (mcp-server-<name>.log) under its logs folder; Claude Code's /mcp shows status; the Inspector shows stderr. For remote servers, correlate by request ID and trace context.
Model-behavior evals
Write 20–50 realistic user requests per server with the expected tool(s) and key arguments. Run them through your own client bridge (lesson 10) against at least two models, and score tool-selection accuracy, argument validity and calls per task. Re-run after every description or schema change.
Security tests (automate them)
- Unauthenticated request → 401 with
WWW-Authenticate. - Token for another audience → 401.
- Read-scope token calling a write tool → 403 with scope challenge.
- Path traversal (
../../etc/passwd) and SSRF (http://169.254.169.254/) inputs → rejected. - Injection strings in data returned by tools do not trigger writes in your agent tests.
Worked example: CI for a Pakistani e-commerce MCP server
A Lahore marketplace ships an orders MCP server used by support agents in Claude and by an internal bot. Their CI runs unit tests, in-memory protocol tests (schemas described, stable order, error paths), a nightly model eval of 40 support questions across two models, and security tests against a staging deployment with a test identity provider. A pull request that shortened a tool description dropped tool-selection accuracy noticeably on the nightly eval; they caught and reverted it before release.
Pitfalls
- Testing only in one host.
- No negative tests for auth and scopes.
- Treating the Inspector as the test suite (it is for exploration).
- Not testing against both old and new protocol versions during the transition.
Measuring success
Protocol test coverage per tool, eval tool-selection accuracy per model, security test pass rate, and mean time to diagnose host connection issues.
Key takeaways
- Layer tests: unit, in-memory protocol, Inspector exploration, host, model-behavior evals and security.
- Use the MCP Inspector to explore schemas, results, errors and raw JSON-RPC traffic.
- Test schemas and error paths in code with an in-memory client in CI.
- Most host failures come from paths, env vars and stdout pollution.
- Automate negative security tests for audience, scopes, traversal and SSRF.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Add three protocol tests and three security tests to your server's CI, and write ten model-behavior eval requests with expected tools.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.