Model Context Protocol (MCP): Connect AI to Your Tools and DataMCP in production · Lesson 13 of 18

Testing and debugging MCP servers with the Inspector and automated tests

Article · 14 min · 9 min lecture

Video lecture

Testing and debugging MCP servers with the Inspector and automated tests

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Testing and debugging

  • Six test layers
  • The MCP Inspector
  • Automated tests
  • Debugging playbook

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

A layered testing strategy

  1. Unit tests for the business logic behind each tool (pure functions, no MCP).
  2. Protocol tests that call tools through an MCP client in memory: schemas, structured output, error paths.
  3. Interactive exploration with the MCP Inspector.
  4. Host tests in the real clients your users run.
  5. Model-behavior evals: do models pick the right tools with the right arguments?
  6. Security tests: auth, scopes, injection and path/SSRF cases.

The MCP Inspector

The Inspector is the official interactive developer tool for MCP servers. It connects over stdio or Streamable HTTP, lists tools, resources and prompts, lets you call tools with custom arguments, and shows the raw JSON-RPC traffic and server logs.

  • Python SDK: uv run mcp dev server.py launches your server under the Inspector.
  • Any server: npx @modelcontextprotocol/inspector then choose the transport and command/URL, or pass the launch command directly, e.g. npx @modelcontextprotocol/inspector node dist/server.js.
  • It also has a CLI mode useful for scripts and CI; check its README for current flags.

What to check in the Inspector: every tool's schema renders sensibly (descriptions on every field), structured results match the output schema, errors come back as isError with helpful text, pagination works, and nothing leaks secrets.

Protocol tests in code

Python (SDK v2) connects a Client directly to your server object:

import pytest
from mcp import Client
from server import mcp

@pytest.fixture
def anyio_backend():
    return "asyncio"

@pytest.fixture
async def client():
    async with Client(mcp, raise_exceptions=True) as c:
        yield c

@pytest.mark.anyio
async def test_tools_listed_in_stable_order(client):
    names1 = [t.name for t in (await client.list_tools()).tools]
    names2 = [t.name for t in (await client.list_tools()).tools]
    assert names1 == names2 and "find_campaigns" in names1

@pytest.mark.anyio
async def test_every_input_property_is_described(client):
    for tool in (await client.list_tools()).tools:
        for prop, spec in tool.input_schema.get("properties", {}).items():
            assert "description" in spec or "enum" in spec, f"{tool.name}.{prop} lacks a description"

@pytest.mark.anyio
async def test_unknown_customer_is_tool_error(client):
    result = await client.call_tool("get_customer", {"customer_id": "nope"})
    assert result.is_error
    assert "search_customers" in result.content[0].text   # the error tells the model what to do

TypeScript (SDK v2) can drive the HTTP handler in-process: a StreamableHTTPClientTransport whose fetch calls handler.fetch, then client.callTool(...) and assertions on structuredContent / isError. Close the client and handler.close() between tests.

Debugging common failures

SymptomLikely causeFix
Host shows server as failed; works in terminalRelative path, missing env var, host can't find uv/nodeAbsolute paths; env in host config; absolute path to runtime
Connection drops right after start (stdio)Something wrote to stdoutLog to stderr; remove prints
401 loops on a remote serverMetadata URL wrong, audience mismatch, clock skewCheck Protected Resource Metadata, token aud, server time
Tool never chosen by the modelVague name/description, overlaps with another toolRewrite description; eval in two hosts
Tool chosen but args invalidSchema too loose or undocumentedEnums, formats, descriptions
Results truncated or hugeNo paginationCursor pagination, summaries
Works on 2025 clients, fails on new ones (or vice versa)Protocol version mismatchUpgrade SDK; test both eras

Server logs: Claude Desktop writes per-server logs (mcp-server-<name>.log) under its logs folder; Claude Code's /mcp shows status; the Inspector shows stderr. For remote servers, correlate by request ID and trace context.

Model-behavior evals

Write 20–50 realistic user requests per server with the expected tool(s) and key arguments. Run them through your own client bridge (lesson 10) against at least two models, and score tool-selection accuracy, argument validity and calls per task. Re-run after every description or schema change.

Security tests (automate them)

  • Unauthenticated request → 401 with WWW-Authenticate.
  • Token for another audience → 401.
  • Read-scope token calling a write tool → 403 with scope challenge.
  • Path traversal (../../etc/passwd) and SSRF (http://169.254.169.254/) inputs → rejected.
  • Injection strings in data returned by tools do not trigger writes in your agent tests.

Worked example: CI for a Pakistani e-commerce MCP server

A Lahore marketplace ships an orders MCP server used by support agents in Claude and by an internal bot. Their CI runs unit tests, in-memory protocol tests (schemas described, stable order, error paths), a nightly model eval of 40 support questions across two models, and security tests against a staging deployment with a test identity provider. A pull request that shortened a tool description dropped tool-selection accuracy noticeably on the nightly eval; they caught and reverted it before release.

Pitfalls

  • Testing only in one host.
  • No negative tests for auth and scopes.
  • Treating the Inspector as the test suite (it is for exploration).
  • Not testing against both old and new protocol versions during the transition.

Measuring success

Protocol test coverage per tool, eval tool-selection accuracy per model, security test pass rate, and mean time to diagnose host connection issues.

Key takeaways

  • Layer tests: unit, in-memory protocol, Inspector exploration, host, model-behavior evals and security.
  • Use the MCP Inspector to explore schemas, results, errors and raw JSON-RPC traffic.
  • Test schemas and error paths in code with an in-memory client in CI.
  • Most host failures come from paths, env vars and stdout pollution.
  • Automate negative security tests for audience, scopes, traversal and SSRF.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Your stdio server works in the terminal but disconnects immediately in a host. What should you check first?
  2. Which test best catches a description change that makes models stop choosing a tool?
  3. What should a read-scope token calling a write tool receive?

Put it into practice

Add three protocol tests and three security tests to your server's CI, and write ten model-behavior eval requests with expected tools.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.