Level 04 · Tool use

Function calling

Letting the model call functions that you define.

Sourced

Concept at a glance

The model requests a tool. Code runs it.

SequenceConceptual illustration
The model requests a tool. Code runs it.Model request leads to Your code. Your code leads to Tool result. A tool call is a structured request for an action, not the action itself.Model requestTool name + argumentsYour codeValidate and run the toolTool resultReturn data for the answerThe model requests a tool. Code runs it.Model request leads to Your code. Your code leads to Tool result. A tool call is a structured request for an action, not the action itself.Model requestTool name + argumentsYour codeValidate and run the toolTool resultReturn data for the answer
Read the connections in words
  • Model request → Your code: Validate and run the tool.
  • Your code → Tool result: Return data for the answer.
Key idea

A tool call is a structured request for an action, not the action itself.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Function calling: see it in practice.

A model proposing a named tool and structured arguments that application code validates and executes.

What you’ll walk through

Follow an English request into a proposed call to an application function, then inspect the result returned to the model. Distinguish choosing a function from the application actually executing it.

The task in this version

Check whether three replacement filters can be reserved.

What you’ll learn to check

English request, optional argument view, validation result, read-only lookup, and a separately approved reservation.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Check whether three replacement filters can be reserved.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
lookup_stock(part) returns five F2 filters. Reservation is a separate write action.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

The function contract defines arguments and results. The application must still enforce access and handle invalid or failed calls.

1 / 6

Apply this to your project

Describe your task to your own model and use Function calling as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Function calling gives the model a fixed list of actions your code defined, each with a name, a description and an argument schema, and lets it choose whether to use one, which one, and what to put in the arguments. Anthropic calls the same mechanism tool use: “Claude determines when to call a tool based on the user’s request and the tool’s description. It then returns a structured call that your application executes (client tools) or that Anthropic executes (server tools)”[1]. The model produces a call, not an effect; this page is about the first of those two cases, where your own code runs it.

Function calling sits at level 4, tools. The one real choice in a run is the model’s: which action, if any, and with what arguments: the decided_by: "model" step in the run below. Your code decides which actions to offer, runs whichever one gets chosen, and asks for the final answer. That is the line to level 5: here the model chooses once, inside a run your code bounds; a single agent chooses again after every result, and leaves the loop only when it decides to.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Function calling

The model picks one action and its arguments; your code runs it and asks once more.

Level 4 · Tool use
QuestionQuestionMODELpicks a tool, or answerspicks a tool,or answersTOOLlookup_part(num)lookup_part(num)MODELanswers using the resultanswers usingthe resultAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 04Your code chose

The question arrives

"What does part HLV-2205 cost?"
0 tokens · 0 ms

Practical guidance

Look in your chat app’s settings, or an icon near the message box, for connectors, plugins, actions or tools: the feature that lets the model do something beyond answering from what it already knows. Custom GPTs with actions, Claude connectors and Gemini connected apps are the same mechanism wearing a product name: the app offers the model a list of actions, the model picks one when the request calls for it, and the product carries it out. OpenAI is retiring one of those: its own FAQ names a “Scheduled retirement” on which “Custom GPTs stop running” for affected workspaces, though “the dates are subject to change”[6], so check what a connector migrates to before building a habit around it.

Before turning one on, read what it can actually do, not just its name. A tool that only looks something up (today’s weather, an account balance) fails safely if the model reaches for it by mistake. A tool with a real consequence sits right next to it on the same permission screen and looks just as ordinary: OpenAI’s own examples range from “Get today’s weather for a location” to “Issue refunds for a lost order”[4]. Read every tool in the list this way, not only the one you meant to add.

To check whether a connector actually ran rather than the model answering from memory, ask something it could only get right by checking: what is actually on your calendar on a specific day, not what a typical day would look like. Whether anything gets called at all is a judgment call the app makes on its own: it “calls a tool when the request maps to that tool’s described capability and the answer isn’t already in context”[1]. A fact that could only have come from your real data means it worked; a generic-sounding answer with no sign anything ran means it guessed instead, and the fix is usually to ask more specifically, naming the exact thing to check.

For anything with a side effect, look for a screen that shows the action and its arguments before it runs, not just the final answer; skipping that is skipping the one point where you could still say no.

If the single thing you need is already a button in the app itself, skip the connector and use the button.

Implementation details

All three makers’ APIs in this page’s sources have the same two halves, so the code below is the shape you write against any of them. Google states the division plainly: “The model doesn’t execute the function itself. Extract the name and args and execute in your application”[5].

The example offers two tools, search(query) and lookup_part(part_number), and lets the model call at most one of them. The system prompt says so directly, and the code enforces it a second way that does not depend on the model following instructions: only first.tool_calls[0] is ever run, and the follow-up call that asks for a final answer is not given the tools list at all, so there is nothing left for the model to call even if it wanted to. That second fact is where level 4 stops and level 5 starts: capping the run at one action is a decision your code made in advance, not one the model makes about when to stop.

examples/function_calling/run.py · lines 39–95
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
    del embedder  # level 4 retrieves through its tools, not a vector index
    sections = load_sections(corpus_dir)
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
    tracer.record(kind="code", decided_by="code", title="Build prompt with tool definitions", detail="search, lookup_part")

    first = model.complete(messages, tools=TOOLS, max_tokens=300)
    if not first.tool_calls:
        tracer.record(
            kind="model",
            decided_by="model",
            title="Model answers directly, no tool call",
            detail=first.text[:200],
            tokens_in=first.tokens_in,
            tokens_out=first.tokens_out,
            ms=first.ms,
        )
        return Answer.from_text(first.text)

    call = first.tool_calls[0]
    # the prompt allows one call; if the model asked for more, the code drops the rest, and the
    # trace has to say so rather than quietly showing a tidier run than the one that happened
    dropped = "" if len(first.tool_calls) == 1 else f" (dropped {len(first.tool_calls) - 1} further call(s))"
    tracer.record(
        kind="model",
        decided_by="model",
        title=f"Model calls {call.name}",
        detail=json.dumps(call.arguments, sort_keys=True) + dropped,
        tokens_in=first.tokens_in,
        tokens_out=first.tokens_out,
        ms=first.ms,
    )
    result_text, citations = _run_tool(call, sections)
    tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])

    follow_up = messages + [
        Message(role="assistant", content=f"[called {call.name}({json.dumps(call.arguments)})]"),
        Message(role="user", content=f"Tool result:\n{result_text}\n\nNow answer the question: {question}"),
    ]
    final = model.complete(follow_up, max_tokens=400)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask for a final answer",
        detail=final.text[:200],
        tokens_in=final.tokens_in,
        tokens_out=final.tokens_out,
        ms=final.ms,
    )
    return Answer.from_text(final.text, retrieved_sources=citations)

If the model’s first response carries no tool call, that is still the one model decision this level records: it chose to answer directly rather than to act, the same kind of choice as the stop at level 5. If it asks for more than one tool at once (Anthropic notes that “by default, Claude may call multiple tools in a single response”[3]) the code drops every call after the first and says so in the trace, rather than silently running one and hiding that a second was requested.

_run_tool is the closest thing here to validating a call before acting on it: it checks the tool’s name against the two it knows and falls through to unknown_tool for anything else, which reports the mismatch as text instead of raising. It does not check the shape of the arguments (part_number is read with a plain default, not checked against the schema) which is what a maker’s own schema-conformance feature is for. Anthropic’s tip on the same page as the quote above: “Add strict: true to your custom tool definitions to ensure Claude’s tool calls always match your schema exactly”[1]. OpenAI documents the same idea for its own schema: “Setting strict to true will ensure function calls reliably adhere to the function schema, instead of being best effort”[4].

This site’s electronics-test bench (docs/THE-BENCH.md) draws the same line one step further. examples/common/bench.py sorts every instrument command into three classes, not two: READ_ONLY_HEADERS and is_read_only() let an agent run any query on its own because a query changes nothing; a command that sets a value needs code to check it against a safety envelope first; and a command that energizes the board needs that check plus a person’s Approval naming the set point, the same across a production test, an engineering sweep or a precise measurement. _run_tool’s two-tool whitelist is only the first of those three lines, drawn for a lookup and a part search; a tool that could turn something on would need the other two as well. Run it yourself:

examples/function_calling/README.md · lines 16–16
python -m examples.function_calling --model stub:scripted
When you do not need this

Try routing first if you already know, from the question’s surface form, which single action applies: the model does not need to choose an action your code can already tell apart. Try prompt chaining, or a plain conditional, if the set of actions is small and always runs in the same order regardless of what the model says.

Move up to function calling once the right action depends on something only the model can judge from open-ended input (which of several tools applies, or whether none does) and getting that judgment wrong sometimes is cheap enough to tolerate.

Two siblings at this level change where the action comes from rather than who chooses it. Reach for MCP when more than one application needs the same tools, or the tools should come from a server you did not write: the model’s decision is identical, and what moves is the boundary the call crosses. Reach for code execution when the thing you need done cannot be enumerated in advance as a list of named actions at all.

Failure modes

Confident call, wrong or missing argument

How to notice it
The model calls the right tool but the argument does not match anything real (a part number that was never in the parts list, a query that does not resemble the question) and the tool answers with whatever it was actually handed rather than what the reader meant.
How to test for it
Ask about a part number that does not exist and confirm the tool reports it as not found, rather than the model inventing a price to go with a citation that never backed one.

Extra tool calls silently dropped

How to notice it
The model asks for more than one action in a single turn, and only the first one visibly happens, with nothing telling you a second request existed at all.
How to test for it
Script a model response with two tool calls and check the trace records that the extra one was dropped, not just that the first one ran.

A tool result is treated as trustworthy text

How to notice it
A search result or a lookup can carry text written to look like an instruction, and nothing about being a tool result rather than a user message stops the model from reading it as one.
How to test for it
Anthropic's own guidance: "an attacker who can influence it may embed instructions that try to redirect Claude (indirect prompt injection)". Add a document section with an embedded instruction to the corpus and see whether a search that surfaces it changes the answer to match it.

The model stops reaching for the tool at all

How to notice it
Across many similar questions, the share that get a tool call drops toward zero even though the documents still hold the answer, because the tool’s description drifted out of sync with what people actually ask.
How to test for it
Track how often decided_by: "model" ends in a tool call versus a direct answer over a batch of known-lookup questions; a falling rate with no change in the questions is a description problem, not a model problem.

A tool is more powerful than the question needed

How to notice it
The call that ran was the right one, on the right input, but the tool itself could do more than this question ever required, so a future mistaken call has a bigger blast radius than a wrong answer.
How to test for it
List every tool offered for a given prompt and check whether each one's effect (what it can change, not just what it can read) matches what that prompt's questions actually need.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

2Model calls, tool used
1Model calls, no tool needed
~190Tokens in, tool-call turn
~14Tokens out, tool-call turn
Compared with RAG (level 2)RAG always retrieves and always calls the model once, so every question costs the same. Function calling asks the model first, so a question it can already answer costs one call instead of two; the second call only happens when the model decides a tool is needed.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

Function calling is one of the examples the site’s own 60-question set scores directly (see docs/EVALS.md), the same way RAG and routing are: it answers a question about the documents and cites what it used. Level 4 adds one more thing worth checking beyond the answer itself: a result file’s model_decided_steps should equal the number of questions run, exactly: one model decision per question, never zero and never more than one. A number outside that range means the trace is wrong before the answer is even graded.

Unlike routing, whose --dry projection follows the cheapest branch because the token-counting stand-in cannot produce a label the code recognizes, this example’s branch depends on whether tools were offered, not on parsed text, and the stand-in calls the first tool every time tools are offered, so --dry already projects the tool-calling branch, the more expensive one.

No result file exists for function calling yet. Run python scripts/eval_run.py --example function_calling --model <spec> --dry to project the cost of a real run before spending anything on one.

Run it

What to monitor

The share of questions where decided_by is model that end in a tool call versus a direct answer, tracked over time against a batch of questions you know should call a tool. A falling rate with no change in the traffic is the tool description going stale, not the model getting worse.

Cost at volume

Every question pays for at least one call. Only the ones where the model reaches for a tool pay for the second, so cost tracks how often real traffic actually needs an action, not a fixed number per question.

How it fails in production

A tool with a side effect runs on an argument the model half-guessed, because nothing between the model's call and the tool's execution checked that the argument was real. A second tool call the model asked for is dropped with no record, and a person debugging a wrong answer has no way to know one was ever requested.

What to log

The full list of tools offered, which one (if any) was called and with what arguments, the raw tool result, and whether decided_by was model or code for every step, so a wrong answer traces back to the wrong tool, the wrong argument, or the model declining to act at all.

Try it

  1. Use it

    Find a chat app feature that can act beyond answering: a connector, a plugin, a custom action. Ask it something the action does not cover: does it say it could not act, or quietly answer from memory?

  2. Build it

    Run python -m examples.function_calling --model stub:scripted from the repo root: the model calls lookup_part on HLV-2205, the code runs it, the answer cites parts-list#2. Change that part number in SCRIPTED (examples/function_calling/__main__.py) to HLV-9999: the tool reports it not found, the next reply still prices it at $52.00, and the citations line empties.

  3. Either lane

    That is the first failure mode above; cause another the same way, from evals/corpus/ only.

How it connects

Before, after and instead of this

Move up when

  • Single agentThe next action depends on what the last one returned, so the model has to choose again and decide when to stop.

Often used with

Instead of

Optional: products, tools, and models

9 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 3 more examples
In practice

Look up a part price

The model requests lookup_part with a part number. Your code runs the lookup and returns the result.

Out there

Named products, tools and models

Products4
  • Claude connectorsAnthropic · tool connections in a chat app
  • Custom GPTs with actionsOpenAI · tool calling in a chat appRetires 2026-12-11
  • Gemini connected appsGoogle · tool connections in a chat app · formerly Gemini Extensions
  • Mistral VibeMistral AI · ai agent for work and coding · formerly Le Chat
Tools6
  • AI SDKVercel · TypeScript AI and agent SDK
  • Claude APIAnthropic · model API
  • ComposioComposio · prebuilt tool connections
  • Gemini APIGoogle · model API
  • OpenAI APIOpenAI · model API
  • Strands AgentsStrands Agents · agent harness SDK

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Tool use with Claude · Anthropic (accessed 09/19/2026)
  2. Handle tool calls · Anthropic (accessed 09/19/2026)
  3. Parallel tool use · Anthropic (accessed 09/19/2026)
  4. Function calling · OpenAI (API documentation) (accessed 09/19/2026)
  5. Function calling with the Gemini API (archived copy) · Google (Gemini API documentation, via the Internet Archive) (accessed 09/19/2026)
  6. Custom GPT retirement and migration FAQ (archived copy) · OpenAI (help center, via the Internet Archive) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page