Model Context Protocol (MCP): Connect AI to Your Tools and DataWhy MCP and how it works · Lesson 2 of 18
MCP architecture: hosts, clients, servers and the protocol
Video lecture
MCP architecture: hosts, clients, servers and the protocol
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 MCP architecture
To build or secure anything with MCP, you need a crisp mental model of who does what. In this lesson you'll learn the three roles, host, client and server, the two layers of the protocol, how discovery works in the old and new spec versions, and exactly what happens when a model calls a tool.
0:24 Why architecture matters
Why spend time on architecture? Because nearly every MCP security incident and most confusing bugs come from a fuzzy idea of who is responsible for what. Here's an analogy. A restaurant has diners, waiters and kitchens. The diner is the user. The host application is the waiter, who takes orders, checks allergies and decides what to bring to the table. Each kitchen is a server that prepares one kind of dish. The kitchen never sees the whole table's conversation. It only sees the order ticket the waiter hands it. Keep that picture in mind.
1:05 Three roles
The host is the AI application you interact with, like Claude Desktop, an IDE, ChatGPT or your own agent. It manages the model, the conversation, consent and security policy. Inside the host are clients, and each client maintains a connection to exactly one server. The server exposes capabilities, like tools, resources and prompts, either as a local process or a remote web service. One design principle matters a lot: servers are isolated. A server never sees the whole conversation or other servers' data. The host decides what each server receives, and that isolation is a security property worth protecting.
1:48 Two layers
MCP has two layers. The data layer is JSON RPC two point oh: requests, results, errors and notifications, with defined methods like tools list, tools call, resources read and prompts get. The transport layer is how those messages travel: standard input and output for local subprocesses, and Streamable HTTP for remote servers. We'll go deep on transports in lesson five.
2:14 Discovery: then and now
Discovery changed in twenty twenty six. In the twenty twenty five spec versions, every connection began with an initialize handshake, where client and server swapped protocol versions and capabilities, and HTTP could carry a session id. In the twenty twenty six version, the handshake and protocol sessions are gone. Every request carries its protocol version and client capabilities in a metadata field, and servers implement a discover method for clients that want capabilities up front. If a server needs state across calls, it hands the model an explicit handle, like a cart id, that gets passed back as an argument.
2:57 Simple example: calendar + email
A simple example. You ask an AI app, what's on my calendar tomorrow, and draft an email moving the budget review. The host has two clients: one connected to a calendar server, one to an email server. The model first calls the calendar tool. The calendar server sees only a date range, not your question about the budget. Then the model calls the email server's draft tool with the text it wrote. The email server never sees your calendar. The host saw everything and decided what each server received.
3:36 Capabilities
Capabilities tell each side what the other supports. Servers declare things like tools, resources, prompts, completions and optional extensions. Clients declare things like elicitation, which lets servers ask the user for input. Sampling and roots, from older versions, are now deprecated. The rule is simple: never use a feature the other side hasn't declared. Official SDKs handle both protocol eras, but when you evaluate a third party server or client, check which versions it supports.
4:09 A tool call over HTTP
Let's follow one tool call over HTTP. The client posts a JSON RPC request to the server's endpoint. The headers include the protocol version, plus two new ones, M C P method and M C P name, that let gateways and firewalls route and authorize without reading the body, and a bearer token. The body says tools call, with the tool name and arguments. The server replies with a result marked complete, containing human readable content and, optionally, structured content that matches the tool's output schema.
4:46 The host's loop
Now the host's side. It lists tools from every connected server and gives their names, descriptions and schemas to the model, often namespaced so two search tools don't collide. The model decides to call one. The host checks policy and consent, then the right client sends the call. The result goes back into the model's context. The spec says there should always be a human able to deny tool calls, and hosts should show which tools are exposed and when they run. That's host behavior. A server can't guarantee it.
5:25 Example: retail insights assistant
A worked example. A UK retailer's customer insights assistant is an internal web app, which is the host. Its backend runs three MCP clients connected to three remote servers: orders, reviews and a read only analytics database, all over Streamable HTTP with OAuth. Policy in the host exposes only read tools to analysts. Because servers are isolated, the reviews server never sees order data, and compromising one server doesn't expose the others' credentials. Documenting each server's role, transport, auth, data scope and protocol versions becomes the basis for every security review.
6:05 Deeper: the retail insights assistant
Let's deepen the UK retailer's insights assistant. Analysts ask questions like, which products with falling reviews also had delivery delays last month. The host sends a date range to the orders server and a product list to the reviews server, never the whole conversation. When the reviews vendor had a security incident, illustrative scenario, the retailer's exposure was limited to review text the host had chosen to request, because the orders server's credentials and data were never visible to the reviews server. Their security team later used the host, client and server map as the starting point for every new integration review, which cut review time because the data flows were already documented.
6:54 Watch me do it: reading the raw messages
Watch me do it. Let's read one real tool call in the Inspector's history, message by message. I call add with a equals one and b equals two. The request shows JSON RPC version two point oh, an id, the method tools call, and params with the tool name and arguments. Under params there's an underscore meta field with client information; on a twenty twenty six connection it also carries the protocol version and client capabilities. The response has the same id, so the client can match it. In the result, resultType says complete, content holds a text block with three, and structured content holds result three. There's no error field, and is error is false. If I call add with a string for b, the response still arrives as a result, but with is error true and a validation message the model could read. That's the whole protocol in two messages.
8:00 Try this now
Try this now. Run the demo server from lesson one under the Inspector and open the message history. Call the add tool and read a greeting resource. For each message, find the method name, the parameters, the metadata fields and the result. Then draw the host, client and server map for an AI assistant you use or plan to build. For each server, write its transport, how it authenticates, what data it can touch and which protocol versions it supports. That map is your first security document.
8:38 Recap
To recap: the host owns the model, consent and policy; each client talks to one isolated server; messages are JSON RPC over stdio or HTTP; and the twenty twenty six spec made requests self describing. Your next step: run the demo server under the Inspector, watch the raw messages, and then draw the host, client and server map for an assistant you use or plan.
The three roles
- Host: the AI application the user interacts with (Claude Desktop, Claude Code, an IDE, ChatGPT, your own agent). It manages the model, the conversation, user consent and security policy.
- Client: a component inside the host that maintains a connection to one server. A host typically runs many clients, one per connected server.
- Server: a program that exposes capabilities (tools, resources, prompts) over MCP. It can be a local subprocess or a remote web service.
A key design principle: servers should be easy to build and isolated from each other. A server never sees the whole conversation or other servers' data; the host decides what each server receives. That isolation is a security property; keep it.
The layers
- Data layer: JSON-RPC 2.0 messages (requests, results, errors, notifications) with defined methods such as
tools/list,tools/call,resources/read,prompts/get. - Transport layer: how messages move. stdio for local subprocesses, Streamable HTTP for remote servers (lesson 5).
Discovery and versions
In the 2025-era specs (2025-06-18, 2025-11-25), a connection starts with an initialize request where client and server exchange protocol versions and capabilities, followed by an initialized notification. Streamable HTTP could carry a session ID.
In 2026-07-28, that handshake is gone. Every request carries its protocol version, client capabilities and (recommended) client identity in _meta; servers identify themselves in each result's _meta. Servers must implement a new server/discover method so clients can learn supported versions and capabilities up front if they want. Protocol-level sessions and the Mcp-Session-Id header are removed; if a server needs state across calls, it mints an explicit handle (for example a cart_id) returned by one tool and passed as an argument to the next, which the model can see and reason about.
Official SDKs handle both eras: the v2 TypeScript and Python SDKs speak 2026-07-28 and serve older clients too. When you evaluate a third-party server or client, check which protocol versions it supports.
Capabilities
Each side advertises what it supports. Servers declare capabilities such as tools, resources, prompts, completions (argument autocompletion) and optional extensions. Clients declare capabilities such as elicitation (and, in older versions, sampling and roots, now deprecated in 2026-07-28). Neither side should use a feature the other hasn't declared.
A request, end to end (Streamable HTTP, 2026-07-28)
POST /mcp HTTP/1.1
Content-Type: application/json
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: search_contacts
Authorization: Bearer eyJ...
{"jsonrpc": "2.0", "id": 7, "method": "tools/call",
"params": {"name": "search_contacts", "arguments": {"query": "Acme"},
"_meta": {"io.modelcontextprotocol/clientInfo": {"name": "my-agent", "version": "1.4.0"}}}}The Mcp-Method and Mcp-Name headers let gateways, rate limiters and WAFs route and authorize without parsing the body. The server replies with a JSON result (or an SSE stream when it sends progress notifications first):
{"jsonrpc": "2.0", "id": 7, "result": {
"resultType": "complete",
"content": [{"type": "text", "text": "[{\"id\":\"c_101\",\"name\":\"Sara Khan\"}]"}],
"structuredContent": {"contacts": [{"id": "c_101", "name": "Sara Khan"}]},
"isError": false}}How the host uses the server
- The host lists tools from each connected server and passes their names, descriptions and schemas to the model (often namespaced, such as
crm__search_contacts). - The model decides to call a tool.
- The host checks policy and user consent, then its client sends
tools/callto the right server. - The result is added to the model's context.
The spec says there should always be a human in the loop with the ability to deny tool invocations, and hosts should show which tools are exposed and when they are called. That is host behavior, not something a server can guarantee.
Worked example: mapping a real deployment
A UK retailer's customer-insights assistant:
| Role | Component |
|---|---|
| Host | Internal web app with a chat UI and an agent loop |
| Clients | Three MCP clients inside the web app's backend |
| Servers | orders (remote, Streamable HTTP, OAuth), reviews (remote), analytics-sql (remote, read-only) |
| Policy | Host allows only read tools for analysts; no write tools exposed |
Because servers are isolated, the reviews server never sees order data, and a compromise of one server does not expose the others' credentials.
Hands-on: watch the protocol
Run the demo server from lesson 1 under the Inspector and open its message/history view. Call a tool and read a resource, then inspect the JSON-RPC requests and responses. Identify: the method, the params, _meta fields, structuredContent versus content, and any error objects.
Pitfalls
- Assuming a server can see the conversation (it cannot, unless the host sends data as tool arguments).
- Confusing host and client responsibilities; consent and policy live in the host.
- Ignoring protocol-version mismatches when mixing older clients with newer servers.
Measuring success
For your architecture, document each server's role, transport, auth, data scope and supported protocol versions. An up-to-date inventory is the foundation for security reviews in module 5.
Key takeaways
- Hosts run the model and policy; each client connects to one server; servers expose capabilities in isolation.
- MCP uses JSON-RPC 2.0 over stdio or Streamable HTTP.
- 2026-07-28 removed the initialize handshake and sessions; requests carry version and capabilities in _meta.
- Mcp-Method and Mcp-Name headers let gateways route and authorize without parsing bodies.
- Consent and human approval are host responsibilities; servers cannot guarantee them.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Draw the host, client and server map for an AI assistant you use or plan, including transport, auth and supported protocol versions for each server.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.