Level 04 · Tool use

Model Context Protocol

A standard way to connect models to tools and data.

Sourced

Concept at a glance

Give tools and data a common connection.

SequenceConceptual illustration
Give tools and data a common connection.AI application leads to MCP connection. MCP connection leads to MCP server. MCP describes the connection; the host still owns permissions and execution policy.AI applicationHost and MCP clientMCP connectionDiscover and call toolsMCP serverExposes tools and dataGive tools and data a common connection.AI application leads to MCP connection. MCP connection leads to MCP server. MCP describes the connection; the host still owns permissions and execution policy.AI applicationHost and MCP clientMCP connectionDiscover and call toolsMCP serverExposes tools and data
Read the connections in words
  • AI application → MCP connection: Discover and call tools.
  • MCP connection → MCP server: Exposes tools and data.
Key idea

MCP describes the connection; the host still owns permissions and execution policy.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Model Context Protocol: see it in practice.

A protocol for exposing tools, resources, and other capabilities to compatible AI applications.

What you’ll walk through

Follow access to a resource or tool through a common connection interface. Inspect what the server exposes, what the client is allowed to use, and how returned content enters the task.

The task in this version

Look up the AX-20 manual through a connected resource service.

What you’ll learn to check

Capability discovery, a tool request/result, access refusal, and an outage, clearly separated from transport details.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Look up the AX-20 manual through a connected resource service.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Mock server advertises manual search and inventory. Caller is authorized only for public manuals.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

A connected server may expose capabilities the user has not authorized for this task. Connection and authentication do not establish content trust.

1 / 6

Apply this to your project

Describe your task to your own model and use Model Context Protocol as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

The Model Context Protocol, or MCP, is a standard way for an AI application to connect to servers that expose tools, resources and prompts, instead of a developer wiring each integration by hand. The protocol’s own architecture overview names three participants: an MCP host is “the AI application that coordinates and manages one or multiple MCP clients”; an MCP client “maintains a connection to an MCP server and obtains context from an MCP server for the MCP host to use”; an MCP server is “a program that provides context to MCP clients”[1]. A host creates one client per server[1].

MCP sits at level 4 for the same reason function calling does: a tool a server exposes is still one fixed action your code carries out once the model picks it, inside a run your code bounds. The line to level 5 falls in the same place too: one choice here, against a loop the model leaves on its own in a single agent. What MCP changes is not how much the model decides but where the tool definitions and the code that runs them live: behind a protocol boundary any compliant host can speak, instead of wired into one application’s own tool-calling code.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Model Context Protocol

The model picks a tool listed by a server; a stand-in client and server carry the call.

Level 4 · Tool use
QuestionQuestionlist tools (tools/list)list tools(tools/list)MODELpicks a tool, or answerspicks a tool,or answersTOOLMCP server (tools/call)MCP server(tools/call)MODELanswers using the resultanswers usingthe resultAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

The question arrives

"What does part HLV-2205 cost?"
0 tokens · 0 ms

Practical guidance

Look for “Connectors,” “MCP servers” or “Integrations” in your chat app’s or editor’s settings, usually with a directory to browse and add from. Claude connectors work this way, and workflow builders such as Langflow and Gumloop let a flow call tools from a connected server the same way a function-calling flow calls one defined inline: the point of the protocol is that the same server works with any of these hosts.

Before adding one, read what it offers: the protocol treats three kinds of thing differently, and only one of them can act on its own. “Tools” are “functions that your LLM can actively call,” and the model decides when to use them; “Resources” are “passive data sources that provide read-only access to information for context,” fetched by the application rather than requested by the model; “Prompts” are “pre-built instruction templates,” invoked by you[2]. A server offering only resources exposes data through read operations, but that does not make it safe to add without review. Check the server’s trustworthiness, what data it can access, and where that data is sent. Resource permissions and access controls still matter[11], and resource text must be treated as untrusted input rather than instructions. One offering tools can take actions on your behalf, described by whoever built the server, not by the app you are using, and the specification itself says “clients MUST consider tool annotations to be untrusted unless they come from trusted servers”[3].

A safe tool next to a dangerous one on the same server can look identical on the surface: both are just a name and a sentence. Read the sentence for what the tool actually does, not how it is phrased.

To test whether a connected tool is actually being called rather than answered from memory, ask for something only that tool could know: look up one specific record through the connector and name the exact field the answer came from. A real call shows you the tool name and its arguments before or alongside the answer; the specification lists, among the behaviors it wants from a client, “Show tool inputs to the user before calling the server, to avoid malicious or accidental data exfiltration”[3], so a product that goes straight to a final answer with nothing shown in between is skipping a step it is supposed to offer you.

If nothing you use regularly calls out to another system, you have no reason to add a connector at all: a plain chat answers as well and reads nothing else.

Implementation details

The example makes the same one decision examples/function_calling does (which tool, with what arguments) through a small stand-in of an MCP client and server instead of a Python function this file defines inline. It borrows the shape of a few things the specification defines and skips almost everything else:

examples/mcp/run.py · lines 52–107
class StandInServer:
    """Handles `tools/list` and `tools/call`, JSON-RPC-shaped. Not a real MCP server: no
    capability negotiation, no other methods, one hard-coded tool."""

    def __init__(self, sections: dict[str, Section]) -> None:
        self._sections = sections

    def handle(self, request: dict) -> dict:
        method = request.get("method")
        if method == "tools/list":
            return _rpc_result(request, {"tools": [SEARCH_TOOL]})
        if method == "tools/call":
            return self._call_tool(request)
        return _rpc_error(request, f"unknown method: {method}")

    def _call_tool(self, request: dict) -> dict:
        params = request.get("params", {})
        if params.get("name") != SEARCH_TOOL["name"]:
            return _rpc_error(request, f"unknown tool: {params.get('name')}")
        query = params.get("arguments", {}).get("query", "")
        hits = bm25_search(self._sections, query, k=SEARCH_K)
        text = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s, _ in hits) or "no results"
        return _rpc_result(request, {"content": [{"type": "text", "text": text}], "isError": False})


def _rpc_result(request: dict, result: dict) -> dict:
    return {"jsonrpc": "2.0", "id": request.get("id"), "result": {"resultType": "complete", **result}}


def _rpc_error(request: dict, message: str) -> dict:
    # -32602 is the spec's own example code for "Unknown tool" in its protocol-errors example.
    return {"jsonrpc": "2.0", "id": request.get("id"), "error": {"code": -32602, "message": message}}


def connect(server: StandInServer) -> Transport:
    """The in-memory pipe: calling `transport(request)` is this example's entire substitute for
    serializing a JSON-RPC message onto stdio or a Streamable HTTP request and reading the reply
    back off it. A real client does that serialization; here the message dict just changes hands
    inside one process."""
    return server.handle


def list_tools(transport: Transport) -> list[dict]:
    response = transport({"jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {}})
    return response["result"]["tools"]


def call_tool(transport: Transport, name: str, arguments: dict) -> tuple[str, bool]:
    """Returns (text, is_error). `is_error` covers both the spec's tool-execution errors
    (`isError: true` in a normal result) and this stand-in's one protocol error."""
    response = transport({"jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": {"name": name, "arguments": arguments}})
    if "error" in response:
        return response["error"]["message"], True
    content = response["result"]["content"]
    text = "\n".join(block["text"] for block in content if block.get("type") == "text")
    return text, response["result"].get("isError", False)

_rpc_result and _rpc_error mirror two shapes straight from the specification’s own examples: a result wrapped in "resultType": "complete", and an error whose code -32602 is the exact code the spec’s own example uses for “Unknown tool”[3]. connect is the whole stand-in for a transport: the specification says a real one “defines how messages are framed and delivered” over something like standard input and output or an HTTP request[4]; here a request dict and a Python function call take the place of both, which is exactly what this file’s docstring and the page’s README.md say plainly is not the real thing.

The stand-in mirrors revision 2026-07-28, and one thing it does not skip is worth naming: there is no connection to open and no handshake to perform before the first request, because the current specification removed both: see Stateless MCP below.

run itself looks almost identical to examples/function_calling’s: list what is available, let the model pick at most one, run it, ask once more.

examples/mcp/run.py · lines 125–145
    del embedder  # this level retrieves through the MCP server's tool, not a vector index
    sections = load_sections(corpus_dir)
    transport = connect(StandInServer(sections))

    mcp_tools = list_tools(transport)
    tracer.record(kind="code", decided_by="code", title="List tools from the MCP server", detail=mcp_tools[0]["name"])

    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
    first = model.complete(messages, tools=[_to_model_tool(t) for t in mcp_tools], max_tokens=300)

    if not first.tool_calls:
        tracer.record(
            kind="model",
            decided_by="model",
            title="Model answers directly, no tool call",
            detail=first.text[:200],
            tokens_in=first.tokens_in,
            tokens_out=first.tokens_out,
            ms=first.ms,
        )
        return Answer.from_text(first.text)

The difference is that mcp_tools came from list_tools(transport) (a round trip through the stand-in server) rather than a Python list this file wrote. A real client would send the same tools/list request over stdio or Streamable HTTP and get the same shape of answer back from a server it may not have written or even trust. Run it yourself:

examples/mcp/README.md · lines 20–20
python -m examples.mcp --model stub:scripted
When you do not need this

Try function calling first if the tools your code needs are the ones your own application defines: a protocol adds a client, a server boundary and a message format to maintain, for no benefit if nothing outside your own codebase will ever call those tools a different way.

Move up to MCP once more than one application needs the same tools, or the tools should come from someone else’s server (a database, a ticketing system, a search index) that you would rather connect to than reimplement.

Stateless MCP

“Stateless” is the specification’s own word for the current revision, not a paraphrase. Its changelog lists “Make MCP stateless: remove the initialize/notifications/initialized handshake. Every request now carries its protocol version and client capabilities in _meta”[6] among the major changes since the previous revision, 2025-11-25[6]. The maintainers’ roadmap, updated August 22, 2026, refers to the 2026-07-28 revision as a release already made, not a proposal[9].

Two changes shipped together. The handshake is gone: earlier revisions opened a connection with an initialize call and kept capabilities for as long as it lasted; every request now states its own protocol version and capabilities instead. Protocol-level sessions are gone with it: the same changelog entry removes “protocol-level sessions and the Mcp-Session-Id header from the Streamable HTTP transport”, and says “Servers that need cross-call state use explicit, server-minted handles passed as ordinary tool arguments”[6]. The transport specification is direct about what that took away: servers used to assign a session via a header, a client could open a standalone stream for server-initiated messages, and a broken stream could resume; “None of these mechanisms are part of this revision”[5], and a server that gets an Mcp-Session-Id header from an older client is told to “ignore it, and do not mint or echo session IDs”[5].

The proposal that became this, SEP-2575, gives the reason: the old handshake created “significant challenges for scalability (load balancing requires sticky sessions), resilience (server failure loses session state), and implementation complexity (both client and server must manage session lifecycles)”[7]. A protocol that asks nothing to be remembered between requests has nothing to lose when the backend that handled the last one is gone. The companion proposal, SEP-2567, describes what replaces a session for a server that genuinely needs one: “Stateful workflows use create_*() -> handle + threaded parameters (guidance, not a protocol construct)”[8]: an ordinary tool call hands back an identifier, and later calls pass it back as an argument like any other, rather than something the transport carries.

Both proposals are merged (SEP-2575 on May 11, 2026, SEP-2567 on May 7, 2026), so both changes are in the 2026-07-28 revision: released, not proposed. What is still open is narrower: the roadmap lists “capability scoping for tool lists after SEP-2575” among what maintainers want to look at next[9], a refinement of what a stateless request can express, not a reconsideration of removing sessions in the first place.

Authorization

Connecting to somebody else’s server raises a question the rest of this page does not: how a server you did not write learns who is asking. The specification has a chapter for it, scoped narrowly. The protocol “provides authorization capabilities at the transport level, enabling MCP clients to make requests to restricted MCP servers on behalf of resource owners”, and “This specification defines the authorization flow for HTTP-based transports”[10]. Nothing here is invented for MCP: the chapter builds on OAuth 2.1 and a list of related RFCs, of which it implements a selected subset[10].

None of that moves a decision, which is why it belongs to level 4 rather than a rung of its own. The model still picks one action and your code still carries it out. Authorization is the separate question of whether your code is allowed to carry it out, and on whose behalf. Guardrails limits what may be done; this settles who the server thinks is asking.

You do not need it for a server you run yourself on your own machine: “Implementations using an STDIO transport SHOULD NOT follow this specification, and instead retrieve credentials from the environment”[10]. It is optional in general, too: “Authorization is OPTIONAL for MCP implementations”[10]. A remote server that asks for nothing still conforms, so “it speaks MCP” says nothing about who may call it.

The failure mode the specification spends requirements on is a token reaching the wrong server. “MCP servers MUST validate that access tokens were issued specifically for them as the intended audience”, and “MCP servers MUST NOT accept or transit any other tokens”[10]. A host that passes one server’s token along to another is the case those two sentences exist to stop.

Failure modes

A tool's own description steers the model somewhere it shouldn't go

How to notice it
A connected server's tool description reads like an instruction rather than documentation ('always call this first' or 'ignore prior instructions and') and the model follows it, because the specification lets a description do exactly what it says: steer which tool the model picks.
How to test for it
Before connecting a new server, read every tool's name and description as if it were untrusted text, the way the specification itself says to treat tool annotations from a server you have not verified.

A tool with a side effect runs with no visible confirmation

How to notice it
The model calls a tool that changes something (files a ticket, sends a message) and the host shows nothing before or after, so there is no point at which a person could have said no.
How to test for it
Check whether the host shows the tool name and its arguments before the call runs, the way the specification asks clients to; a chat bubble with only the final answer is not that.

The tool list changes and the code still has the old one

How to notice it
A server adds, removes or changes a tool, and a client that cached the list from tools/list keeps offering the model a tool that no longer exists, or the old shape of one that does.
How to test for it
This example cannot show the failure: it calls list_tools on every run and caches nothing. Check a real client instead. Does it list once at startup or per request, does it honor the freshness hint the specification puts on a tool list, and does it subscribe to the list-changed notification the specification defines?

An unexpected method or a hallucinated tool name is treated as a crash instead of an answer

How to notice it
The model asks for a tool that does not exist on the server, or the code sends a request the server does not implement, and the whole run fails instead of the model getting a chance to recover.
How to test for it
Send a `tools/call` naming a tool the server never registered and confirm the server returns a protocol error the caller can read, rather than raising; see tests/test_example_mcp.py.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

2Model calls, tool used
1Model calls, no tool needed
2Protocol round trips per call
~210Tokens in, tool-call turn
Compared with Function calling (level 4)The model-facing shape and the model call count are identical to function calling. What MCP adds is the list-tools round trip and the client/server boundary the call goes through, which cost time and, over a real transport, network latency that an in-process function call does not pay.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

The example answers the same kind of question function_calling does, with the same citations, so it is scored against the site’s own 60-question set the same way (see docs/EVALS.md): exact or rubric match, plus citation hit rate. Level 4 adds the same check function calling describes: model_decided_steps should equal the number of questions run, exactly one per question.

The number specific to this technique, beyond accuracy, is how much the protocol layer itself costs: compare tokens_in and wall_time_s on a result file here against function_calling’s own, on the same questions, to see what routing the same one decision through a client and server boundary adds over calling a Python function directly. In this example that boundary is an in-process function call, so the difference is the tools/list round trip and nothing else; over a real transport it also carries network time this measurement will not show.

No result file exists for MCP yet. Run python scripts/eval_run.py --example mcp --model <spec> --dry to project the cost of a real run before spending anything on one.

Run it

What to monitor

Which servers and tools are actually being called versus merely listed, and how often the model declines every offered tool. A tool nobody ever calls is either miscategorized or badly described, the same failure routing's dead fallback route is, one level up.

Cost at volume

Every question pays for at least the tools/list round trip and one model call; a second model call only happens when a tool actually ran. Over a real transport, each round trip also pays network latency an in-process call does not, so cost tracks server count and network distance, not just question count.

How it fails in production

A server's tool descriptions change and the host is still working from a cached list, so the model is offered a tool that no longer behaves the way its old description said. A server outside your control changes what a tool's description tells the model to do, and nothing downstream re-checks it.

What to log

Which server and tool were listed and called, the raw request and result at the protocol boundary, and whether the call succeeded or came back as a protocol or tool-execution error, so a wrong answer traces back to a bad tool choice, a bad argument, or the server itself.

Try it

  1. Use it

    Find a product that lets you connect an MCP server. Before adding one, read every tool it would expose, and ask whether you would trust a stranger to write instructions your assistant follows automatically.

  2. Build it

    Run python -m examples.mcp --model stub:scripted from the repo root: the server lists its one tool, the model calls it, and the answer cites the section with the price in it. Then, in a Python shell, build the transport yourself and call call_tool(transport, "search_docs", {"query": "pump"}): a tool name the server never registered. Back comes ("unknown tool: search_docs", True), a protocol error the caller can read, not an exception. Which layer decided that, and what would a real host do with it?

  3. Either lane

    Compare this page’s run to function calling’s. Which steps are identical, and which exist only because a protocol boundary sits between the model’s choice and the code that carries it out?

  4. Build it

    Read examples/common/bench.py (docs/THE-BENCH.md describes it): READ_ONLY_HEADERS and is_read_only() sort an instrument’s commands into a query an agent may run on its own, a command that sets a value, and a command that energizes the board, which needs a person’s Approval naming the set point. Exposed as MCP tools rather than Python functions, which class could the specification’s own "Tools" primitive leave to the model alone, and which would still need the human-in-the-loop step this page’s Use it lane describes?

How it connects

Before, after and instead of this

Move up when

  • The agent harnessThe tools are connected and what is missing is everything around the model: the loop, the permissions, the caps and the sandbox.

Instead of

Optional: products, tools, and models

6 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Connect an assistant to a document store

An MCP server exposes search and read tools; the host application decides which calls it permits.

Out there

Named products, tools and models

Products3
  • Claude connectorsAnthropic · tool connections in a chat app
  • GumloopGumloop · visual workflow builder
  • LangflowIBM · visual workflow builder
Tools3
  • ComposioComposio · prebuilt tool connections
  • Model Context Protocolopen standard · protocol for tools and data
  • Strands AgentsStrands Agents · agent harness SDK

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Architecture overview · Model Context Protocol (specification site) (accessed 09/19/2026)
  2. Understanding MCP servers · Model Context Protocol (specification site) (accessed 09/19/2026)
  3. Tools · Model Context Protocol (specification) (accessed 09/19/2026)
  4. Transports Overview · Model Context Protocol (specification) (accessed 09/19/2026)
  5. Streamable HTTP · Model Context Protocol (specification) (accessed 09/19/2026)
  6. Key Changes · Model Context Protocol (specification changelog) (accessed 09/19/2026)
  7. SEP-2575: Make MCP Stateless · Model Context Protocol (GitHub, merged proposal), 05/11/2026 (accessed 09/19/2026)
  8. SEP-2567: Sessionless MCP via Explicit State Handles · Model Context Protocol (GitHub, merged proposal), 05/07/2026 (accessed 09/19/2026)
  9. Roadmap · Model Context Protocol (specification site), 08/22/2026 (accessed 09/19/2026)
  10. Authorization · Model Context Protocol (specification) (accessed 09/19/2026)
  11. Resources: security considerations · Model Context Protocol (specification) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page