Model Context Protocol (MCP): Connect AI to Your Tools and DataCapstone: an MCP server for a marketing CRM · Lesson 18 of 18
Capstone part 3: secure, test and ship crescent-crm
Video lecture
Capstone part 3: secure, test and ship crescent-crm
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Capstone part 3
Crescent CRM works on your machine. Now let's make it something an agency can trust with real clients. In this final lesson you'll add authentication and scopes, deploy it remotely, automate tests and evals, run a security review, and roll it out to hosts.
0:19 Why shipping matters
Why does shipping need its own lesson? Because this is where good intentions become evidence. Anyone can say, it's secure and it works. A platform team shows the tests that prove it, the checklist with links, the eval results in real hosts, and the plan for when something goes wrong. Think of a building inspection. The builder's confidence doesn't matter; the inspector's signed checklist does. In this lesson, you're both the builder and the person preparing for inspection.
0:53 Step 1: auth and scopes
Step one is authentication. Reuse the JWT verifier from the authorization lesson and require the read scope on every request. Then enforce the write scope inside the write tools with a small helper that checks the verified token and returns the caller's identity. That identity replaces the placeholder in the created by field, so every logged activity records who really did it. In production, prefer a four oh three with a scope challenge so clients can step up, and keep the per tool check as defense in depth.
1:31 Step 2: deploy
Step two is deployment. Containerize the server, run it over Streamable HTTP behind TLS, and put it behind the company gateway. Load secrets like the database connection and OAuth settings from a secret manager. At the gateway, rate limit per user on the tool name header, log method, tool, user and duration, and cap request sizes. And run two replicas from day one, which proves the server really is stateless before anyone depends on it.
2:04 Step 3: tests
Step three is automated tests. The lesson's tests check that the company brief never exposes contact emails or phone numbers, that a five month performance range is rejected with a helpful message, and that logging the same activity twice doesn't create a duplicate. Security tests run against staging: no token and wrong audience both get four oh one, a read token can't log activities, marketing drafts to non consented contacts fail, and SQL characters in search queries are treated as plain text.
2:40 Simple example: three layers
A simple example of defense in depth. An analyst's token has only the read scope. Their AI app, influenced by an injected note in an activity record, tries to call log activity. First line of defense: the transport's required scopes. Second line: the require scope helper inside the tool rejects it with a clear message. Third line: even if both were misconfigured, the activity would be recorded with the analyst's real identity, so the audit log shows exactly who did what. Three independent layers, each simple.
3:17 Step 4: evals in two hosts
Step four is model evaluation in two hosts. Run your twenty five eval requests through the client bridge with two models, and manually in Claude and VS Code for a subset. Score tool selection, argument validity, calls per task and correctness against your fictional data, with thresholds agreed with the account team. When selection goes wrong, fix the tool descriptions, not host prompts, because descriptions travel with the server to every host.
3:48 Step 5: security review
Step five is the security review. Confirm the server publishes protected resource metadata and validates audience. No token passthrough, and the server uses its own least privilege database user. Scopes enforced per tool. No generic SQL, export, send or delete tools. Parameterized queries, caps and range limits. Minimal personal data in outputs and none in logs. Tool definitions hashed in the catalog. Dependency and image scanning. And a red team check that injected text inside activity notes never triggers writes in your agent eval.
4:25 Step 6: rollout
Step six is rollout. Add a catalog entry with owner, scopes, data classification and approved hosts. Pilot with five account managers and two analysts for two weeks, and feed every failure into the eval set. Then add it as an organization connector in Claude, share a VS Code snippet with developers, and give the reporting agent a read only machine token. Publish a changelog and book a review in six months.
4:56 Example: pilot findings
Here's how a pilot might go. Account managers love the company brief but ask for last contacted dates. Analysts hit the ninety two day limit for quarter over quarter comparisons. The team adds last contacted to the contact summary, with no new personal data, and adds guidance to the performance tool's description about using two calls for quarters. Eval accuracy stays stable, and the changelog records both changes. That's the loop: pilot, learn, adjust, re evaluate, document.
5:29 After launch
What happens after launch? Plan a short operating rhythm. Weekly for the first month: review errors, slow tools, denied scope attempts and a sample of real questions that went badly, and add them to the eval set. Monthly after that: review the catalog entry, dependency updates and costs. And every time a new host version or protocol version arrives, rerun the host tests. A server that's cared for stays trustworthy; one that's forgotten slowly drifts.
6:02 Deeper: the two-week pilot (illustrative)
Let's deepen the pilot for Crescent Growth with illustrative numbers. Five account managers and two analysts, two weeks. Account managers used the company brief most, often daily. Analysts hit the ninety two day limit a few times and adapted. Security monitoring showed a handful of denied write attempts from analysts, all accidental, all correctly refused. Two eval failures came from real questions phrased in Roman Urdu, which the model sometimes misread; the team added those phrasings to the eval set and clarified two tool descriptions. At the end of the pilot, the head of accounts approved rollout to all staff, with a six month review booked in the catalog.
6:49 Watch me do it: auth + tests
Watch me do it. Let's walk through the scope helper and the tests. Require scope calls get access token, which returns the verified token for the current HTTP request. If there's no token, it allows a local development user only when the unauthenticated flag is explicitly set, otherwise raises authentication required. If the scope is missing, it raises a clear tool error. Otherwise it returns the token's subject, or the client id. In log activity, created by now comes from require scope with crm write. Now the tests, run with the local flag and a test database. The first asserts that every contact in a brief has exactly three fields: id, name and role. The second asks for a five month range and checks the ninety two day message. The third logs the same activity twice and checks for, already logged. I run pytest, three passes. Then, against staging over HTTP, the negative auth tests: no token, wrong audience and read only token all fail as expected.
8:02 Try this now
Try this now. Deploy crescent CRM to staging with OAuth and two replicas. Run the protocol tests, the security tests and your twenty five request eval in two hosts. Fill in the security checklist with evidence links, write the catalog entry, and plan a two week pilot. Put all of it in one folder: that's your launch pack. When a friend who hasn't seen the project can read it and explain how to switch the server off, it's done.
8:36 Deliverable and course recap
Your final deliverable is a launch pack: the design doc, the repository with tests, the security checklist with evidence, eval results in two hosts, the catalog entry, the rollout plan and the changelog. It's exactly what real platform teams produce. Across this course you've learned why MCP exists, its architecture and primitives, transports, building servers in Python and TypeScript, connecting hosts and clients, OAuth and security, testing, deployment, registries and enterprise rollout. Assemble the pack, then take the final exam.
From working to shippable
crescent-crm works locally. Shipping it to the agency means: remote deployment over Streamable HTTP, OAuth with scopes per tool, real user identity in writes, automated tests and evals, a security review, and host rollout.
Step 1: add authentication and scopes
Reuse the JWT verifier from lesson 11 and require crm.read for every request. Then enforce crm.write inside write tools. In the Python SDK v2, get_access_token() returns the AccessToken your verifier built for the current HTTP request (and None over stdio or in-memory tests). A small helper keeps tools clean:
import os
from mcp.server.auth.middleware.auth_context import get_access_token
from mcp.server.mcpserver.exceptions import ToolError
ALLOW_LOCAL = os.environ.get("CRM_ALLOW_UNAUTHENTICATED") == "1" # dev/tests over stdio only; never in prod
def require_scope(scope: str) -> str:
"""Return the caller's identity if the token has the scope, else raise a model-readable error."""
token = get_access_token()
if token is None:
if ALLOW_LOCAL:
return "local-dev-user"
raise ToolError("Authentication required.")
if scope not in token.scopes:
raise ToolError(f"This action requires the {scope} scope. Ask your admin or re-authenticate.")
return token.subject or token.client_idIn crm_log_activity, replace the placeholder with created_by = require_scope("crm.write"), and call require_scope("crm.write") at the top of crm_draft_client_email. In production, prefer returning an HTTP 403 with a scope challenge so clients can trigger step-up authorization; the SDK's auth middleware handles the required-scope check at the transport level, and per-tool checks add defense in depth.
Step 2: deploy
- Containerize (lesson 14), run with
transport="streamable-http"behind TLS on a container platform or behind the company API gateway. - Environment:
CRM_DB(or a Postgres DSN),MCP_RESOURCE_URL,OAUTH_ISSUER,OAUTH_JWKS_URLfrom a secret manager. - Gateway: rate limit per user on
Mcp-Name, log method/tool/user/duration, block request bodies over a size limit. - Two replicas from day one to prove statelessness.
Step 3: automated tests
import pytest
from mcp import Client
from server import mcp
@pytest.fixture
def anyio_backend():
return "asyncio"
@pytest.mark.anyio
async def test_brief_hides_contact_details():
async with Client(mcp) as c:
companies = (await c.call_tool("crm_find_companies", {"query": "Noor"})).structured_content
cid = companies["companies"][0]["company_id"]
brief = (await c.call_tool("crm_company_brief", {"company_id": cid})).structured_content
for contact in brief["key_contacts"]:
assert set(contact) == {"contact_id", "full_name", "role"}
@pytest.mark.anyio
async def test_long_range_rejected():
async with Client(mcp) as c:
r = await c.call_tool("crm_campaign_performance", {"start": "2026-01-01", "end": "2026-06-01"})
assert r.is_error and "92 days" in r.content[0].text
@pytest.mark.anyio
async def test_log_activity_is_idempotent():
args = {"company_id": "co_noor", "type": "call", "note": "Wants proposal by Friday",
"idempotency_key": "test-key-0001"}
async with Client(mcp) as c:
await c.call_tool("crm_log_activity", args)
second = await c.call_tool("crm_log_activity", args)
assert "already logged" in second.content[0].text(Run these in-memory tests with CRM_ALLOW_UNAUTHENTICATED=1 and a test database; in-memory clients bypass the HTTP auth layer by design. Run auth tests over HTTP against staging.)
Security tests against staging: no token → 401; token with wrong audience → 401; crm.read token calling crm_log_activity → rejected; marketing draft to a non-consented contact → tool error; SQL metacharacters in query → treated as text.
Step 4: model evals in two hosts
Run your 25 eval requests through (a) the client bridge from lesson 10 with two models and (b) manually in Claude and VS Code for a subset. Score tool selection, argument validity, calls per task and answer correctness against the fictional data. Target: agreed thresholds with account-team representatives. Fix descriptions, not prompts, when selection is wrong.
Step 5: security review checklist
| Check | Status |
|---|---|
| Protected Resource Metadata served; audience validated | |
| No token passthrough; server uses its own DB credentials (read-mostly DB user) | |
Scopes enforced per tool; write tools need crm.write | |
| No generic SQL, export, send or delete tools | |
| Parameterized queries; row caps; range limits | |
| Personal data minimized in outputs and absent from logs | |
| Tool definitions hashed and recorded in the catalog | |
| Host/Origin validation for any local HTTP use | |
| Dependency and container image scanning | |
| Red-team: injection text in activity notes does not cause writes in the agent eval |
Step 6: rollout
- Add a catalog entry (lesson 15) with owner, scopes, data classification and hosts.
- Pilot with five account managers and two analysts for two weeks; collect failures into the eval set.
- Add as an org-managed connector in Claude; share a VS Code
mcp.jsonsnippet; issue a client-credentialscrm.readtoken to the reporting agent. - Publish a changelog; schedule a review in six months.
Worked example: the pilot findings
In the pilot (illustrative), account managers loved crm_company_brief but asked for "last contacted" per contact; analysts hit the 92-day limit for quarter-over-quarter comparisons. The team added last_contacted to ContactSummary (no new personal data), and added compare_previous guidance to the description for quarters (two 92-day calls). Tool-selection accuracy on the eval stayed stable; the changelog noted both changes.
Final deliverable
A launch pack: design doc (part 1), repository with tests, security checklist with evidence, eval results in two hosts, catalog entry, rollout plan and changelog. That pack is exactly what a real platform team produces, and it is the evidence behind this course's badge.
Key takeaways
- Shipping means remote deployment, OAuth with per-tool scopes, real identity in writes, tests, evals and review.
- Enforce scopes at the transport and inside write tools for defense in depth.
- Run at least two replicas from day one to prove statelessness.
- Evaluate in two hosts and fix tool descriptions when selection goes wrong.
- Ship with a launch pack: design, tests, security evidence, evals, catalog entry, rollout and changelog.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Deploy crescent-crm to staging with OAuth, run the protocol, security and eval suites, complete the security checklist, and assemble the launch pack.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.