Model Context Protocol (MCP): Connect AI to Your Tools and DataCapstone: an MCP server for a marketing CRM · Lesson 18 of 18

Capstone part 3: secure, test and ship crescent-crm

Article · 20 min · 9 min lecture

Video lecture

Capstone part 3: secure, test and ship crescent-crm

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Capstone part 3

  • Auth and scopes
  • Deploy
  • Tests and evals
  • Security review
  • Rollout

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From working to shippable

crescent-crm works locally. Shipping it to the agency means: remote deployment over Streamable HTTP, OAuth with scopes per tool, real user identity in writes, automated tests and evals, a security review, and host rollout.

Step 1: add authentication and scopes

Reuse the JWT verifier from lesson 11 and require crm.read for every request. Then enforce crm.write inside write tools. In the Python SDK v2, get_access_token() returns the AccessToken your verifier built for the current HTTP request (and None over stdio or in-memory tests). A small helper keeps tools clean:

import os
from mcp.server.auth.middleware.auth_context import get_access_token
from mcp.server.mcpserver.exceptions import ToolError

ALLOW_LOCAL = os.environ.get("CRM_ALLOW_UNAUTHENTICATED") == "1"   # dev/tests over stdio only; never in prod

def require_scope(scope: str) -> str:
    """Return the caller's identity if the token has the scope, else raise a model-readable error."""
    token = get_access_token()
    if token is None:
        if ALLOW_LOCAL:
            return "local-dev-user"
        raise ToolError("Authentication required.")
    if scope not in token.scopes:
        raise ToolError(f"This action requires the {scope} scope. Ask your admin or re-authenticate.")
    return token.subject or token.client_id

In crm_log_activity, replace the placeholder with created_by = require_scope("crm.write"), and call require_scope("crm.write") at the top of crm_draft_client_email. In production, prefer returning an HTTP 403 with a scope challenge so clients can trigger step-up authorization; the SDK's auth middleware handles the required-scope check at the transport level, and per-tool checks add defense in depth.

Step 2: deploy

  • Containerize (lesson 14), run with transport="streamable-http" behind TLS on a container platform or behind the company API gateway.
  • Environment: CRM_DB (or a Postgres DSN), MCP_RESOURCE_URL, OAUTH_ISSUER, OAUTH_JWKS_URL from a secret manager.
  • Gateway: rate limit per user on Mcp-Name, log method/tool/user/duration, block request bodies over a size limit.
  • Two replicas from day one to prove statelessness.

Step 3: automated tests

import pytest
from mcp import Client
from server import mcp

@pytest.fixture
def anyio_backend():
    return "asyncio"

@pytest.mark.anyio
async def test_brief_hides_contact_details():
    async with Client(mcp) as c:
        companies = (await c.call_tool("crm_find_companies", {"query": "Noor"})).structured_content
        cid = companies["companies"][0]["company_id"]
        brief = (await c.call_tool("crm_company_brief", {"company_id": cid})).structured_content
        for contact in brief["key_contacts"]:
            assert set(contact) == {"contact_id", "full_name", "role"}

@pytest.mark.anyio
async def test_long_range_rejected():
    async with Client(mcp) as c:
        r = await c.call_tool("crm_campaign_performance", {"start": "2026-01-01", "end": "2026-06-01"})
        assert r.is_error and "92 days" in r.content[0].text

@pytest.mark.anyio
async def test_log_activity_is_idempotent():
    args = {"company_id": "co_noor", "type": "call", "note": "Wants proposal by Friday",
            "idempotency_key": "test-key-0001"}
    async with Client(mcp) as c:
        await c.call_tool("crm_log_activity", args)
        second = await c.call_tool("crm_log_activity", args)
        assert "already logged" in second.content[0].text

(Run these in-memory tests with CRM_ALLOW_UNAUTHENTICATED=1 and a test database; in-memory clients bypass the HTTP auth layer by design. Run auth tests over HTTP against staging.)

Security tests against staging: no token → 401; token with wrong audience → 401; crm.read token calling crm_log_activity → rejected; marketing draft to a non-consented contact → tool error; SQL metacharacters in query → treated as text.

Step 4: model evals in two hosts

Run your 25 eval requests through (a) the client bridge from lesson 10 with two models and (b) manually in Claude and VS Code for a subset. Score tool selection, argument validity, calls per task and answer correctness against the fictional data. Target: agreed thresholds with account-team representatives. Fix descriptions, not prompts, when selection is wrong.

Step 5: security review checklist

CheckStatus
Protected Resource Metadata served; audience validated
No token passthrough; server uses its own DB credentials (read-mostly DB user)
Scopes enforced per tool; write tools need crm.write
No generic SQL, export, send or delete tools
Parameterized queries; row caps; range limits
Personal data minimized in outputs and absent from logs
Tool definitions hashed and recorded in the catalog
Host/Origin validation for any local HTTP use
Dependency and container image scanning
Red-team: injection text in activity notes does not cause writes in the agent eval

Step 6: rollout

  1. Add a catalog entry (lesson 15) with owner, scopes, data classification and hosts.
  2. Pilot with five account managers and two analysts for two weeks; collect failures into the eval set.
  3. Add as an org-managed connector in Claude; share a VS Code mcp.json snippet; issue a client-credentials crm.read token to the reporting agent.
  4. Publish a changelog; schedule a review in six months.

Worked example: the pilot findings

In the pilot (illustrative), account managers loved crm_company_brief but asked for "last contacted" per contact; analysts hit the 92-day limit for quarter-over-quarter comparisons. The team added last_contacted to ContactSummary (no new personal data), and added compare_previous guidance to the description for quarters (two 92-day calls). Tool-selection accuracy on the eval stayed stable; the changelog noted both changes.

Final deliverable

A launch pack: design doc (part 1), repository with tests, security checklist with evidence, eval results in two hosts, catalog entry, rollout plan and changelog. That pack is exactly what a real platform team produces, and it is the evidence behind this course's badge.

Key takeaways

  • Shipping means remote deployment, OAuth with per-tool scopes, real identity in writes, tests, evals and review.
  • Enforce scopes at the transport and inside write tools for defense in depth.
  • Run at least two replicas from day one to prove statelessness.
  • Evaluate in two hosts and fix tool descriptions when selection goes wrong.
  • Ship with a launch pack: design, tests, security evidence, evals, catalog entry, rollout and changelog.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why run two replicas of crescent-crm from day one?
  2. The eval shows models choosing crm_find_deals when users ask for a company brief. What should you change first?
  3. Which item belongs on the security checklist for crescent-crm?

Put it into practice

Deploy crescent-crm to staging with OAuth, run the protocol, security and eval suites, complete the security checklist, and assemble the launch pack.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.