# Model Context Protocol

_Level 04 · Tool use · sourced_

A standard way to connect models to tools and data.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow access to a resource or tool through a common connection interface. Inspect what the server exposes, what the client is allowed to use, and how returned content enters the task.

**Assumptions:** A connected server may expose capabilities the user has not authorized for this task. Connection and authentication do not establish content trust.

**Design choices:** Choose the resources and tools needed for the job and preserve their schemas and provenance. Use a direct integration when it is simpler for a single fixed capability.

**Request:** Look up the AX-20 manual through a connected resource service.

**Starting evidence:** Mock server advertises manual search and inventory. Caller is authorized only for public manuals.

**Action and control:** Discover capabilities, choose manual search, check access, and receive resource content.

**Stage records (authored, not executed):**

### Input record

Mock server advertises manual search and inventory. Caller is authorized only for public manuals.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose the resources and tools needed for the job and preserve their schemas and provenance. Use a direct integration when it is simpler for a single fixed capability.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Discover capabilities, choose manual search, check access, and receive resource content.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Manual resource returned with version metadata. Inventory access was neither requested nor granted.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Capability discovery, a tool request/result, access refusal, and an outage, clearly separated from transport details.

If the result falls short:
When a server is unavailable or returns untrusted instructions, separate the access failure or content from the user's request. Do not quietly substitute a different authority.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Adapt the example to document stores, project systems, or engineering services. Specify the integration's actual read and write scope rather than assuming a standard protocol determines permissions.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Manual resource returned with version metadata. Inventory access was neither requested nor granted.

**Change something — Call the advertised inventory tool:** Server denies access. Discovery describes what exists, not what this caller may execute.

**Decision:** Does an advertised tool imply permission?

**Answer:** No; authorization is a separate check.

**Why:** Discovery is not permission; tool descriptions can be misleading and services can be unavailable.

**Review criteria:** Capability discovery, a tool request/result, access refusal, and an outage, clearly separated from transport details.

**Recovery:** When a server is unavailable or returns untrusted instructions, separate the access failure or content from the user's request. Do not quietly substitute a different authority.

**Adapt it:** Adapt the example to document stores, project systems, or engineering services. Specify the integration's actual read and write scope rather than assuming a standard protocol determines permissions.

The Model Context Protocol, or MCP, is a standard way for an AI application to connect to servers
that expose tools, resources and prompts, instead of a developer wiring each integration by hand.
The protocol's own architecture overview names three participants: an MCP host is "the AI
application that coordinates and manages one or multiple MCP clients"; an MCP client "maintains a
connection to an MCP server and obtains context from an MCP server for the MCP host to use"; an
MCP server is "a program that provides context to MCP clients"[1]. A host creates one
client per server[1].

MCP sits at level 4 for the same reason function calling does: a tool a server exposes is still
one fixed action your code carries out once the model picks it, inside a run your code bounds.
The line to level 5 falls in the same place too: one choice here, against a loop the model
leaves on its own in a [single agent](/gradient_ascent/techniques/single-agent/). What MCP
changes is not how much the model decides but where the tool definitions and the code that runs
them live: behind a protocol boundary any compliant host can speak, instead of wired into one
application's own tool-calling code.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 4 · MCP. The same steps are described in the sections below._

## Practical guidance

Look for "Connectors," "MCP servers" or "Integrations" in your chat app's or editor's settings,
usually with a directory to browse and add from. Claude connectors work this way, and workflow
builders such as Langflow and Gumloop let a flow call tools from a connected server the same way a
function-calling flow calls one defined inline: the point of the protocol is that the same server
works with any of these hosts.

Before adding one, read what it offers: the protocol treats three kinds of thing differently, and
only one of them can act on its own. "Tools" are "functions that your LLM can actively call," and
the model decides when to use them; "Resources" are "passive data sources that provide read-only
access to information for context," fetched by the application rather than requested by the model;
"Prompts" are "pre-built instruction templates," invoked by you[2]. A server offering
only resources exposes data through read operations, but that does not make it safe to add
without review. Check the server's trustworthiness, what data it can access, and where that data
is sent. Resource permissions and access controls still matter[11], and resource text
must be treated as untrusted input rather than instructions. One offering tools can take actions on your
behalf, described by whoever built the server, not by the app you are using, and the specification
itself says "clients MUST consider tool annotations to be untrusted unless they come from trusted
servers"[3].

A safe tool next to a dangerous one on the same server can look identical on the surface: both are
just a name and a sentence. Read the sentence for what the tool actually does, not how it is
phrased.

To test whether a connected tool is actually being called rather than answered from memory, ask
for something only that tool could know: look up one specific record through the connector and
name the exact field the answer came from. A real call shows you the tool name and its arguments
before or alongside the answer; the specification lists, among the behaviors it wants from a
client, "Show tool inputs to the user before calling the server, to avoid malicious or accidental
data exfiltration"[3], so a product that goes straight to a final answer with nothing
shown in between is skipping a step it is supposed to offer you.

If nothing you use regularly calls out to another system, you have no reason to add a connector at
all: a plain chat answers as well and reads nothing else.

## Implementation details

The example makes the same one decision `examples/function_calling` does (which tool, with what
arguments) through a small stand-in of an MCP client and server instead of a Python function
this file defines inline. It borrows the shape of a few things the specification defines and
skips almost everything else:

`examples/mcp/run.py` (lines 52-107)

```python
class StandInServer:
    """Handles `tools/list` and `tools/call`, JSON-RPC-shaped. Not a real MCP server: no
    capability negotiation, no other methods, one hard-coded tool."""

    def __init__(self, sections: dict[str, Section]) -> None:
        self._sections = sections

    def handle(self, request: dict) -> dict:
        method = request.get("method")
        if method == "tools/list":
            return _rpc_result(request, {"tools": [SEARCH_TOOL]})
        if method == "tools/call":
            return self._call_tool(request)
        return _rpc_error(request, f"unknown method: {method}")

    def _call_tool(self, request: dict) -> dict:
        params = request.get("params", {})
        if params.get("name") != SEARCH_TOOL["name"]:
            return _rpc_error(request, f"unknown tool: {params.get('name')}")
        query = params.get("arguments", {}).get("query", "")
        hits = bm25_search(self._sections, query, k=SEARCH_K)
        text = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s, _ in hits) or "no results"
        return _rpc_result(request, {"content": [{"type": "text", "text": text}], "isError": False})

def _rpc_result(request: dict, result: dict) -> dict:
    return {"jsonrpc": "2.0", "id": request.get("id"), "result": {"resultType": "complete", **result}}

def _rpc_error(request: dict, message: str) -> dict:
    # -32602 is the spec's own example code for "Unknown tool" in its protocol-errors example.
    return {"jsonrpc": "2.0", "id": request.get("id"), "error": {"code": -32602, "message": message}}

def connect(server: StandInServer) -> Transport:
    """The in-memory pipe: calling `transport(request)` is this example's entire substitute for
    serializing a JSON-RPC message onto stdio or a Streamable HTTP request and reading the reply
    back off it. A real client does that serialization; here the message dict just changes hands
    inside one process."""
    return server.handle

def list_tools(transport: Transport) -> list[dict]:
    response = transport({"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}})
    return response["result"]["tools"]

def call_tool(transport: Transport, name: str, arguments: dict) -> tuple[str, bool]:
    """Returns (text, is_error). `is_error` covers both the spec's tool-execution errors
    (`isError: true` in a normal result) and this stand-in's one protocol error."""
    response = transport({"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": name, "arguments": arguments}})
    if "error" in response:
        return response["error"]["message"], True
    content = response["result"]["content"]
    text = "\n".join(block["text"] for block in content if block.get("type") == "text")
    return text, response["result"].get("isError", False)
```

`_rpc_result` and `_rpc_error` mirror two shapes straight from the specification's own examples:
a result wrapped in `"resultType": "complete"`, and an error whose code `-32602` is the exact code
the spec's own example uses for "Unknown tool"[3]. `connect` is the whole stand-in for a
transport: the specification says a real one "defines how messages are framed and delivered" over
something like standard input and output or an HTTP request[4]; here a request dict and
a Python function call take the place of both, which is exactly what this file's docstring and the
page's `README.md` say plainly is not the real thing.

The stand-in mirrors revision 2026-07-28, and one thing it does not skip is worth naming: there is
no connection to open and no handshake to perform before the first request, because the current
specification removed both: see Stateless MCP below.

`run` itself looks almost identical to `examples/function_calling`'s: list what is available, let
the model pick at most one, run it, ask once more.

`examples/mcp/run.py` (lines 125-145)

```python
    del embedder  # this level retrieves through the MCP server's tool, not a vector index
    sections = load_sections(corpus_dir)
    transport = connect(StandInServer(sections))

    mcp_tools = list_tools(transport)
    tracer.record(kind="code", decided_by="code", title="List tools from the MCP server", detail=mcp_tools[0]["name"])

    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
    first = model.complete(messages, tools=[_to_model_tool(t) for t in mcp_tools], max_tokens=300)

    if not first.tool_calls:
        tracer.record(
            kind="model",
            decided_by="model",
            title="Model answers directly, no tool call",
            detail=first.text[:200],
            tokens_in=first.tokens_in,
            tokens_out=first.tokens_out,
            ms=first.ms,
        )
        return Answer.from_text(first.text)
```

The difference is that `mcp_tools` came from `list_tools(transport)` (a round trip through the
stand-in server) rather than a Python list this file wrote. A real client would send the same
`tools/list` request over stdio or Streamable HTTP and get the same shape of answer back from a
server it may not have written or even trust. Run it yourself:

`examples/mcp/README.md` (lines 20-20)

```text
python -m examples.mcp --model stub:scripted
```

## When you do not need this

Try [function calling](/gradient_ascent/techniques/function-calling/) first if the tools your
code needs are the ones your own application defines: a protocol adds a client, a server
boundary and a message format to maintain, for no benefit if nothing outside your own codebase
will ever call those tools a different way.

Move up to MCP once more than one application needs the same tools, or the tools should come from
someone else's server (a database, a ticketing system, a search index) that you would rather
connect to than reimplement.

## Stateless MCP

"Stateless" is the specification's own word for the current revision, not a paraphrase. Its
changelog lists "Make MCP stateless: remove the initialize/notifications/initialized handshake.
Every request now carries its protocol version and client capabilities in _meta"[6]
among the major changes since the previous revision, 2025-11-25[6]. The maintainers'
roadmap, updated August 22, 2026, refers to the 2026-07-28 revision as a release already made,
not a proposal[9].

Two changes shipped together. The handshake is gone: earlier revisions opened a connection with
an `initialize` call and kept capabilities for as long as it lasted; every request now states its
own protocol version and capabilities instead. Protocol-level sessions are gone with it: the same
changelog entry removes "protocol-level sessions and the Mcp-Session-Id header from the
Streamable HTTP transport", and says "Servers that need cross-call state use explicit,
server-minted handles passed as ordinary tool arguments"[6]. The transport
specification is direct about what that took away: servers used to assign a session via a
header, a client could open a standalone stream for server-initiated messages, and a broken
stream could resume; "None of these mechanisms are part of this revision"[5], and a
server that gets an `Mcp-Session-Id` header from an older client is told to "ignore it, and do
not mint or echo session IDs"[5].

The proposal that became this, SEP-2575, gives the reason: the old handshake created "significant
challenges for scalability (load balancing requires sticky sessions), resilience (server failure
loses session state), and implementation complexity (both client and server must manage session
lifecycles)"[7]. A protocol that asks nothing to be remembered between requests has
nothing to lose when the backend that handled the last one is gone. The companion proposal,
SEP-2567, describes what replaces a session for a server that genuinely needs one: "Stateful
workflows use `create_*() -> handle` + threaded parameters (guidance, not a protocol
construct)"[8]: an ordinary tool call hands back an identifier, and later calls pass it
back as an argument like any other, rather than something the transport carries.

Both proposals are merged (SEP-2575 on May 11, 2026, SEP-2567 on May 7, 2026), so both changes
are in the 2026-07-28 revision: released, not proposed. What is still open is narrower: the
roadmap lists "capability scoping for tool lists after SEP-2575" among what maintainers want to
look at next[9], a refinement of what a stateless request can express, not a
reconsideration of removing sessions in the first place.

## Authorization

Connecting to somebody else's server raises a question the rest of this page does not: how a
server you did not write learns who is asking. The specification has a chapter for it, scoped
narrowly. The protocol "provides authorization capabilities at the transport level, enabling MCP
clients to make requests to restricted MCP servers on behalf of resource owners", and "This
specification defines the authorization flow for HTTP-based transports"[10]. Nothing
here is invented for MCP: the chapter builds on OAuth 2.1 and a list of related RFCs, of which it
implements a selected subset[10].

None of that moves a decision, which is why it belongs to level 4 rather than a rung of its own.
The model still picks one action and your code still carries it out. Authorization is the separate
question of whether your code is allowed to carry it out, and on whose behalf.
[Guardrails](/gradient_ascent/techniques/guardrails/) limits what may be done; this settles who
the server thinks is asking.

You do not need it for a server you run yourself on your own machine: "Implementations using an
STDIO transport SHOULD NOT follow this specification, and instead retrieve credentials from the
environment"[10]. It is optional in general, too: "Authorization is OPTIONAL for MCP
implementations"[10]. A remote server that asks for nothing still conforms, so "it
speaks MCP" says nothing about who may call it.

The failure mode the specification spends requirements on is a token reaching the wrong server.
"MCP servers MUST validate that access tokens were issued specifically for them as the intended
audience", and "MCP servers MUST NOT accept or transit any other tokens"[10]. A host
that passes one server's token along to another is the case those two sentences exist to stop.

## Failure modes

### A tool's own description steers the model somewhere it shouldn't go

- **How to notice it:** A connected server's tool description reads like an instruction rather than documentation ('always call this first' or 'ignore prior instructions and') and the model follows it, because the specification lets a description do exactly what it says: steer which tool the model picks.
- **How to test for it:** Before connecting a new server, read every tool's name and description as if it were untrusted text, the way the specification itself says to treat tool annotations from a server you have not verified.

### A tool with a side effect runs with no visible confirmation

- **How to notice it:** The model calls a tool that changes something (files a ticket, sends a message) and the host shows nothing before or after, so there is no point at which a person could have said no.
- **How to test for it:** Check whether the host shows the tool name and its arguments before the call runs, the way the specification asks clients to; a chat bubble with only the final answer is not that.

### The tool list changes and the code still has the old one

- **How to notice it:** A server adds, removes or changes a tool, and a client that cached the list from tools/list keeps offering the model a tool that no longer exists, or the old shape of one that does.
- **How to test for it:** This example cannot show the failure: it calls list_tools on every run and caches nothing. Check a real client instead. Does it list once at startup or per request, does it honor the freshness hint the specification puts on a tool list, and does it subscribe to the list-changed notification the specification defines?

### An unexpected method or a hallucinated tool name is treated as a crash instead of an answer

- **How to notice it:** The model asks for a tool that does not exist on the server, or the code sends a request the server does not implement, and the whole run fails instead of the model getting a chance to recover.
- **How to test for it:** Send a `tools/call` naming a tool the server never registered and confirm the server returns a protocol error the caller can read, rather than raising; see tests/test_example_mcp.py.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, tool used:** 2
- **Model calls, no tool needed:** 1
- **Protocol round trips per call:** 2
- **Tokens in, tool-call turn:** ~210

**Compared with Function calling (level 4).** The model-facing shape and the model call count are identical to function calling. What MCP adds is the list-tools round trip and the client/server boundary the call goes through, which cost time and, over a real transport, network latency that an in-process function call does not pay.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

The example answers the same kind of question `function_calling` does, with the same citations,
so it is scored against the site's own 60-question set the same way (see `docs/EVALS.md`): exact
or rubric match, plus citation hit rate. Level 4 adds the same check function calling describes:
`model_decided_steps` should equal the number of questions run, exactly one per question.

The number specific to this technique, beyond accuracy, is how much the protocol layer itself
costs: compare `tokens_in` and `wall_time_s` on a result file here against `function_calling`'s
own, on the same questions, to see what routing the same one decision through a client and server
boundary adds over calling a Python function directly. In this example that boundary is an
in-process function call, so the difference is the `tools/list` round trip and nothing else; over
a real transport it also carries network time this measurement will not show.

No result file exists for MCP yet. Run `python scripts/eval_run.py --example mcp --model <spec>
--dry` to project the cost of a real run before spending anything on one.

## Run it

**What to monitor.** Which servers and tools are actually being called versus merely listed, and how often the model declines every offered tool. A tool nobody ever calls is either miscategorized or badly described, the same failure routing's dead fallback route is, one level up.

**Cost at volume.** Every question pays for at least the tools/list round trip and one model call; a second model call only happens when a tool actually ran. Over a real transport, each round trip also pays network latency an in-process call does not, so cost tracks server count and network distance, not just question count.

**How it fails in production.** A server's tool descriptions change and the host is still working from a cached list, so the model is offered a tool that no longer behaves the way its old description said. A server outside your control changes what a tool's description tells the model to do, and nothing downstream re-checks it.

**What to log.** Which server and tool were listed and called, the raw request and result at the protocol boundary, and whether the call succeeded or came back as a protocol or tool-execution error, so a wrong answer traces back to a bad tool choice, a bad argument, or the server itself.

## Try it

1. **Use it.** Find a product that lets you connect an MCP server. Before adding one, read every tool it would expose, and ask whether you would trust a stranger to write instructions your assistant follows automatically.
2. **Build it.** Run python -m examples.mcp --model stub:scripted from the repo root: the server lists its one tool, the model calls it, and the answer cites the section with the price in it. Then, in a Python shell, build the transport yourself and call call_tool(transport, "search_docs", {"query": "pump"}): a tool name the server never registered. Back comes ("unknown tool: search_docs", True), a protocol error the caller can read, not an exception. Which layer decided that, and what would a real host do with it?
3. **Either lane.** Compare this page’s run to function calling’s. Which steps are identical, and which exist only because a protocol boundary sits between the model’s choice and the code that carries it out?
4. **Build it.** Read examples/common/bench.py (docs/THE-BENCH.md describes it): READ_ONLY_HEADERS and is_read_only() sort an instrument’s commands into a query an agent may run on its own, a command that sets a value, and a command that energizes the board, which needs a person’s Approval naming the set point. Exposed as MCP tools rather than Python functions, which class could the specification’s own "Tools" primitive leave to the model alone, and which would still need the human-in-the-loop step this page’s Use it lane describes?


## Sources

1. [Architecture overview](https://modelcontextprotocol.io/docs/2026-07-28/learn/architecture) — Model Context Protocol (specification site) (accessed 2026-09-19)
2. [Understanding MCP servers](https://modelcontextprotocol.io/docs/2026-07-28/learn/server-concepts) — Model Context Protocol (specification site) (accessed 2026-09-19)
3. [Tools](https://modelcontextprotocol.io/specification/2026-07-28/server/tools) — Model Context Protocol (specification) (accessed 2026-09-19)
4. [Transports Overview](https://modelcontextprotocol.io/specification/2026-07-28/basic/transports) — Model Context Protocol (specification) (accessed 2026-09-19)
5. [Streamable HTTP](https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/streamable-http) — Model Context Protocol (specification) (accessed 2026-09-19)
6. [Key Changes](https://modelcontextprotocol.io/specification/2026-07-28/changelog) — Model Context Protocol (specification changelog) (accessed 2026-09-19)
7. [SEP-2575: Make MCP Stateless](https://github.com/modelcontextprotocol/modelcontextprotocol/pull/2575) — Model Context Protocol (GitHub, merged proposal), 2026-05-11 (accessed 2026-09-19)
8. [SEP-2567: Sessionless MCP via Explicit State Handles](https://github.com/modelcontextprotocol/modelcontextprotocol/pull/2567) — Model Context Protocol (GitHub, merged proposal), 2026-05-07 (accessed 2026-09-19)
9. [Roadmap](https://modelcontextprotocol.io/development/roadmap) — Model Context Protocol (specification site), 2026-08-22 (accessed 2026-09-19)
10. [Authorization](https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization) — Model Context Protocol (specification) (accessed 2026-09-19)
11. [Resources: security considerations](https://modelcontextprotocol.io/specification/2026-07-28/server/resources#security-considerations) — Model Context Protocol (specification) (accessed 2026-09-19)


Last reviewed 2026-09-19.
