A standard way to connect models to tools and data.
Sourced
Concept at a glance
Give tools and data a common connection.
SequenceConceptual illustration
Read the connections in words
AI application → MCP connection: Discover and call tools.
MCP connection → MCP server: Exposes tools and data.
Key idea
MCP describes the connection; the host still owns permissions and execution policy.
A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.
GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions
Model Context Protocol: see it in practice.
A protocol for exposing tools, resources, and other capabilities to compatible AI applications.
What you’ll walk through
Follow access to a resource or tool through a common connection interface. Inspect what the server exposes, what the client is allowed to use, and how returned content enters the task.
The task in this version
Look up the AX-20 manual through a connected resource service.
What you’ll learn to check
Capability discovery, a tool request/result, access refusal, and an outage, clearly separated from transport details.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example
Look up the AX-20 manual through a connected resource service.
Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Mock server advertises manual search and inventory. Caller is authorized only for public manuals.
What changed: Establish the facts supplied for this version of the task.
WHY THIS MATTERS
What this case assumes
A connected server may expose capabilities the user has not authorized for this task. Connection and authentication do not establish content trust.
1 / 6
Apply this to your project
Describe your task to your own model and use Model Context Protocol as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Go deeper: practical guidance, failure modes, and implementation
The Model Context Protocol, or MCP, is a standard way for an AI application to connect to servers
that expose tools, resources and prompts, instead of a developer wiring each integration by hand.
The protocol’s own architecture overview names three participants: an MCP host is “the AI
application that coordinates and manages one or multiple MCP clients”; an MCP client “maintains a
connection to an MCP server and obtains context from an MCP server for the MCP host to use”; an
MCP server is “a program that provides context to MCP clients”[1]. A host creates one
client per server[1].
MCP sits at level 4 for the same reason function calling does: a tool a server exposes is still
one fixed action your code carries out once the model picks it, inside a run your code bounds.
The line to level 5 falls in the same place too: one choice here, against a loop the model
leaves on its own in a single agent. What MCP
changes is not how much the model decides but where the tool definitions and the code that runs
them live: behind a protocol boundary any compliant host can speak, instead of wired into one
application’s own tool-calling code.
This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.
Optional: inspect the implementation trace
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Model Context Protocol
The model picks a tool listed by a server; a stand-in client and server carry the call.
Level 4 · Tool use
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step
The run, step by step
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
STEP 01 / 05Your code chose
The question arrives
"What does part HLV-2205 cost?"
0 tokens · 0 ms
Practical guidance
Look for “Connectors,” “MCP servers” or “Integrations” in your chat app’s or editor’s settings,
usually with a directory to browse and add from. Claude connectors work this way, and workflow
builders such as Langflow and Gumloop let a flow call tools from a connected server the same way a
function-calling flow calls one defined inline: the point of the protocol is that the same server
works with any of these hosts.
Before adding one, read what it offers: the protocol treats three kinds of thing differently, and
only one of them can act on its own. “Tools” are “functions that your LLM can actively call,” and
the model decides when to use them; “Resources” are “passive data sources that provide read-only
access to information for context,” fetched by the application rather than requested by the model;
“Prompts” are “pre-built instruction templates,” invoked by you[2]. A server offering
only resources exposes data through read operations, but that does not make it safe to add
without review. Check the server’s trustworthiness, what data it can access, and where that data
is sent. Resource permissions and access controls still matter[11], and resource text
must be treated as untrusted input rather than instructions. One offering tools can take actions on your
behalf, described by whoever built the server, not by the app you are using, and the specification
itself says “clients MUST consider tool annotations to be untrusted unless they come from trusted
servers”[3].
A safe tool next to a dangerous one on the same server can look identical on the surface: both are
just a name and a sentence. Read the sentence for what the tool actually does, not how it is
phrased.
To test whether a connected tool is actually being called rather than answered from memory, ask
for something only that tool could know: look up one specific record through the connector and
name the exact field the answer came from. A real call shows you the tool name and its arguments
before or alongside the answer; the specification lists, among the behaviors it wants from a
client, “Show tool inputs to the user before calling the server, to avoid malicious or accidental
data exfiltration”[3], so a product that goes straight to a final answer with nothing
shown in between is skipping a step it is supposed to offer you.
If nothing you use regularly calls out to another system, you have no reason to add a connector at
all: a plain chat answers as well and reads nothing else.
Implementation details
The example makes the same one decision examples/function_calling does (which tool, with what
arguments) through a small stand-in of an MCP client and server instead of a Python function
this file defines inline. It borrows the shape of a few things the specification defines and
skips almost everything else:
examples/mcp/run.py · lines 52–107
class StandInServer:
"""Handles `tools/list` and `tools/call`, JSON-RPC-shaped. Not a real MCP server: no
capability negotiation, no other methods, one hard-coded tool."""
def __init__(self, sections: dict[str, Section]) -> None:
self._sections = sections
def handle(self, request: dict) -> dict:
method = request.get("method")
if method == "tools/list":
return _rpc_result(request, {"tools": [SEARCH_TOOL]})
if method == "tools/call":
return self._call_tool(request)
return _rpc_error(request, f"unknown method: {method}")
def _call_tool(self, request: dict) -> dict:
params = request.get("params", {})
if params.get("name") != SEARCH_TOOL["name"]:
return _rpc_error(request, f"unknown tool: {params.get('name')}")
query = params.get("arguments", {}).get("query", "")
hits = bm25_search(self._sections, query, k=SEARCH_K)
text = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s, _ in hits) or "no results"
return _rpc_result(request, {"content": [{"type": "text", "text": text}], "isError": False})
def _rpc_result(request: dict, result: dict) -> dict:
return {"jsonrpc": "2.0", "id": request.get("id"), "result": {"resultType": "complete", **result}}
def _rpc_error(request: dict, message: str) -> dict:
# -32602 is the spec's own example code for "Unknown tool" in its protocol-errors example.
return {"jsonrpc": "2.0", "id": request.get("id"), "error": {"code": -32602, "message": message}}
def connect(server: StandInServer) -> Transport:
"""The in-memory pipe: calling `transport(request)` is this example's entire substitute for
serializing a JSON-RPC message onto stdio or a Streamable HTTP request and reading the reply
back off it. A real client does that serialization; here the message dict just changes hands
inside one process."""
return server.handle
def list_tools(transport: Transport) -> list[dict]:
response = transport({"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}})
return response["result"]["tools"]
def call_tool(transport: Transport, name: str, arguments: dict) -> tuple[str, bool]:
"""Returns (text, is_error). `is_error` covers both the spec's tool-execution errors
(`isError: true` in a normal result) and this stand-in's one protocol error."""
response = transport({"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": name, "arguments": arguments}})
if "error" in response:
return response["error"]["message"], True
content = response["result"]["content"]
text = "\n".join(block["text"] for block in content if block.get("type") == "text")
return text, response["result"].get("isError", False)
_rpc_result and _rpc_error mirror two shapes straight from the specification’s own examples:
a result wrapped in "resultType": "complete", and an error whose code -32602 is the exact code
the spec’s own example uses for “Unknown tool”[3]. connect is the whole stand-in for a
transport: the specification says a real one “defines how messages are framed and delivered” over
something like standard input and output or an HTTP request[4]; here a request dict and
a Python function call take the place of both, which is exactly what this file’s docstring and the
page’s README.md say plainly is not the real thing.
The stand-in mirrors revision 2026-07-28, and one thing it does not skip is worth naming: there is
no connection to open and no handshake to perform before the first request, because the current
specification removed both: see Stateless MCP below.
run itself looks almost identical to examples/function_calling’s: list what is available, let
the model pick at most one, run it, ask once more.
examples/mcp/run.py · lines 125–145
del embedder # this level retrieves through the MCP server's tool, not a vector index
sections = load_sections(corpus_dir)
transport = connect(StandInServer(sections))
mcp_tools = list_tools(transport)
tracer.record(kind="code", decided_by="code", title="List tools from the MCP server", detail=mcp_tools[0]["name"])
messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
first = model.complete(messages, tools=[_to_model_tool(t) for t in mcp_tools], max_tokens=300)
if not first.tool_calls:
tracer.record(
kind="model",
decided_by="model",
title="Model answers directly, no tool call",
detail=first.text[:200],
tokens_in=first.tokens_in,
tokens_out=first.tokens_out,
ms=first.ms,
)
return Answer.from_text(first.text)
The difference is that mcp_tools came from list_tools(transport) (a round trip through the
stand-in server) rather than a Python list this file wrote. A real client would send the same
tools/list request over stdio or Streamable HTTP and get the same shape of answer back from a
server it may not have written or even trust. Run it yourself:
examples/mcp/README.md · lines 20–20
python -m examples.mcp --model stub:scripted
When you do not need this
Try function calling first if the tools your
code needs are the ones your own application defines: a protocol adds a client, a server
boundary and a message format to maintain, for no benefit if nothing outside your own codebase
will ever call those tools a different way.
Move up to MCP once more than one application needs the same tools, or the tools should come from
someone else’s server (a database, a ticketing system, a search index) that you would rather
connect to than reimplement.
Stateless MCP
“Stateless” is the specification’s own word for the current revision, not a paraphrase. Its
changelog lists “Make MCP stateless: remove the initialize/notifications/initialized handshake.
Every request now carries its protocol version and client capabilities in _meta”[6]
among the major changes since the previous revision, 2025-11-25[6]. The maintainers’
roadmap, updated August 22, 2026, refers to the 2026-07-28 revision as a release already made,
not a proposal[9].
Two changes shipped together. The handshake is gone: earlier revisions opened a connection with
an initialize call and kept capabilities for as long as it lasted; every request now states its
own protocol version and capabilities instead. Protocol-level sessions are gone with it: the same
changelog entry removes “protocol-level sessions and the Mcp-Session-Id header from the
Streamable HTTP transport”, and says “Servers that need cross-call state use explicit,
server-minted handles passed as ordinary tool arguments”[6]. The transport
specification is direct about what that took away: servers used to assign a session via a
header, a client could open a standalone stream for server-initiated messages, and a broken
stream could resume; “None of these mechanisms are part of this revision”[5], and a
server that gets an Mcp-Session-Id header from an older client is told to “ignore it, and do
not mint or echo session IDs”[5].
The proposal that became this, SEP-2575, gives the reason: the old handshake created “significant
challenges for scalability (load balancing requires sticky sessions), resilience (server failure
loses session state), and implementation complexity (both client and server must manage session
lifecycles)”[7]. A protocol that asks nothing to be remembered between requests has
nothing to lose when the backend that handled the last one is gone. The companion proposal,
SEP-2567, describes what replaces a session for a server that genuinely needs one: “Stateful
workflows use create_*() -> handle + threaded parameters (guidance, not a protocol
construct)”[8]: an ordinary tool call hands back an identifier, and later calls pass it
back as an argument like any other, rather than something the transport carries.
Both proposals are merged (SEP-2575 on May 11, 2026, SEP-2567 on May 7, 2026), so both changes
are in the 2026-07-28 revision: released, not proposed. What is still open is narrower: the
roadmap lists “capability scoping for tool lists after SEP-2575” among what maintainers want to
look at next[9], a refinement of what a stateless request can express, not a
reconsideration of removing sessions in the first place.
Authorization
Connecting to somebody else’s server raises a question the rest of this page does not: how a
server you did not write learns who is asking. The specification has a chapter for it, scoped
narrowly. The protocol “provides authorization capabilities at the transport level, enabling MCP
clients to make requests to restricted MCP servers on behalf of resource owners”, and “This
specification defines the authorization flow for HTTP-based transports”[10]. Nothing
here is invented for MCP: the chapter builds on OAuth 2.1 and a list of related RFCs, of which it
implements a selected subset[10].
None of that moves a decision, which is why it belongs to level 4 rather than a rung of its own.
The model still picks one action and your code still carries it out. Authorization is the separate
question of whether your code is allowed to carry it out, and on whose behalf.
Guardrails limits what may be done; this settles who
the server thinks is asking.
You do not need it for a server you run yourself on your own machine: “Implementations using an
STDIO transport SHOULD NOT follow this specification, and instead retrieve credentials from the
environment”[10]. It is optional in general, too: “Authorization is OPTIONAL for MCP
implementations”[10]. A remote server that asks for nothing still conforms, so “it
speaks MCP” says nothing about who may call it.
The failure mode the specification spends requirements on is a token reaching the wrong server.
“MCP servers MUST validate that access tokens were issued specifically for them as the intended
audience”, and “MCP servers MUST NOT accept or transit any other tokens”[10]. A host
that passes one server’s token along to another is the case those two sentences exist to stop.
Failure modes
A tool's own description steers the model somewhere it shouldn't go
How to notice it
A connected server's tool description reads like an instruction rather than documentation ('always call this first' or 'ignore prior instructions and') and the model follows it, because the specification lets a description do exactly what it says: steer which tool the model picks.
How to test for it
Before connecting a new server, read every tool's name and description as if it were untrusted text, the way the specification itself says to treat tool annotations from a server you have not verified.
A tool with a side effect runs with no visible confirmation
How to notice it
The model calls a tool that changes something (files a ticket, sends a message) and the host shows nothing before or after, so there is no point at which a person could have said no.
How to test for it
Check whether the host shows the tool name and its arguments before the call runs, the way the specification asks clients to; a chat bubble with only the final answer is not that.
The tool list changes and the code still has the old one
How to notice it
A server adds, removes or changes a tool, and a client that cached the list from tools/list keeps offering the model a tool that no longer exists, or the old shape of one that does.
How to test for it
This example cannot show the failure: it calls list_tools on every run and caches nothing. Check a real client instead. Does it list once at startup or per request, does it honor the freshness hint the specification puts on a tool list, and does it subscribe to the list-changed notification the specification defines?
An unexpected method or a hallucinated tool name is treated as a crash instead of an answer
How to notice it
The model asks for a tool that does not exist on the server, or the code sends a request the server does not implement, and the whole run fails instead of the model getting a chance to recover.
How to test for it
Send a `tools/call` naming a tool the server never registered and confirm the server returns a protocol error the caller can read, rather than raising; see tests/test_example_mcp.py.
Cost and latency
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
2Model calls, tool used
1Model calls, no tool needed
2Protocol round trips per call
~210Tokens in, tool-call turn
Compared with Function calling (level 4)The model-facing shape and the model call count are identical to function calling. What MCP adds is the list-tools round trip and the client/server boundary the call goes through, which cost time and, over a real transport, network latency that an in-process function call does not pay.
The example answers the same kind of question function_calling does, with the same citations,
so it is scored against the site’s own 60-question set the same way (see docs/EVALS.md): exact
or rubric match, plus citation hit rate. Level 4 adds the same check function calling describes:
model_decided_steps should equal the number of questions run, exactly one per question.
The number specific to this technique, beyond accuracy, is how much the protocol layer itself
costs: compare tokens_in and wall_time_s on a result file here against function_calling’s
own, on the same questions, to see what routing the same one decision through a client and server
boundary adds over calling a Python function directly. In this example that boundary is an
in-process function call, so the difference is the tools/list round trip and nothing else; over
a real transport it also carries network time this measurement will not show.
No result file exists for MCP yet. Run python scripts/eval_run.py --example mcp --model <spec> --dry to project the cost of a real run before spending anything on one.
Run it
What to monitor
Which servers and tools are actually being called versus merely listed, and how often the model declines every offered tool. A tool nobody ever calls is either miscategorized or badly described, the same failure routing's dead fallback route is, one level up.
Cost at volume
Every question pays for at least the tools/list round trip and one model call; a second model call only happens when a tool actually ran. Over a real transport, each round trip also pays network latency an in-process call does not, so cost tracks server count and network distance, not just question count.
How it fails in production
A server's tool descriptions change and the host is still working from a cached list, so the model is offered a tool that no longer behaves the way its old description said. A server outside your control changes what a tool's description tells the model to do, and nothing downstream re-checks it.
What to log
Which server and tool were listed and called, the raw request and result at the protocol boundary, and whether the call succeeded or came back as a protocol or tool-execution error, so a wrong answer traces back to a bad tool choice, a bad argument, or the server itself.
Try it
Use it
Find a product that lets you connect an MCP server. Before adding one, read every tool it would expose, and ask whether you would trust a stranger to write instructions your assistant follows automatically.
Build it
Run python -m examples.mcp --model stub:scripted from the repo root: the server lists its one tool, the model calls it, and the answer cites the section with the price in it. Then, in a Python shell, build the transport yourself and call call_tool(transport, "search_docs", {"query": "pump"}): a tool name the server never registered. Back comes ("unknown tool: search_docs", True), a protocol error the caller can read, not an exception. Which layer decided that, and what would a real host do with it?
Either lane
Compare this page’s run to function calling’s. Which steps are identical, and which exist only because a protocol boundary sits between the model’s choice and the code that carries it out?
Build it
Read examples/common/bench.py (docs/THE-BENCH.md describes it): READ_ONLY_HEADERS and is_read_only() sort an instrument’s commands into a query an agent may run on its own, a command that sets a value, and a command that energizes the board, which needs a person’s Approval naming the set point. Exposed as MCP tools rather than Python functions, which class could the specification’s own "Tools" primitive leave to the model alone, and which would still need the human-in-the-loop step this page’s Use it lane describes?