# Function calling

_Level 04 · Tool use · sourced_

Letting the model call functions that you define.


## Try this in a recipe
- [Approve the exact change before it happens](/gradient_ascent/recipes/assistant-team.md): Draft a calendar change, bind review to the exact proposal, and detect stale or repeated approvals.

## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an English request into a proposed call to an application function, then inspect the result returned to the model. Distinguish choosing a function from the application actually executing it.

**Assumptions:** The function contract defines arguments and results. The application must still enforce access and handle invalid or failed calls.

**Design choices:** Expose narrow, useful operations rather than forcing the model to assemble fragile low-level steps. Validate arguments and distinguish lookup operations from actions with side effects.

**Request:** Check whether three replacement filters can be reserved.

**Starting evidence:** lookup_stock(part) returns five F2 filters. Reservation is a separate write action.

**Action and control:** Model proposes a named lookup; application validates arguments and runs the fixture lookup.

**Stage records (authored, not executed):**

### Input record

lookup_stock(part) returns five F2 filters. Reservation is a separate write action.

What changed: Establish the facts supplied for this version of the task.

### Design note

Expose narrow, useful operations rather than forcing the model to assemble fragile low-level steps. Validate arguments and distinguish lookup operations from actions with side effects.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Model proposes a named lookup; application validates arguments and runs the fixture lookup.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Five available. Proposed reservation: three. Inventory remains unchanged pending an authorized reservation.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

English request, optional argument view, validation result, read-only lookup, and a separately approved reservation.

If the result falls short:
Return a clear error or uncertain outcome. Retry only when the operation is safe to repeat; a timeout is not proof that a reservation or update failed.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use your existing APIs, business functions, or instrument abstractions. The tool schema should express the operation your application can reliably support.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Five available. Proposed reservation: three. Inventory remains unchanged pending an authorized reservation.

**Change something — Pass a negative reservation quantity:** Validation rejects -3 before execution. A valid tool name does not make arguments valid.

**Decision:** Does checking availability authorize a reservation?

**Answer:** No; separate lookup and write authority.

**Why:** A tool call is a proposal, not authorization; handle invalid arguments, absent stock, and tool failure.

**Review criteria:** English request, optional argument view, validation result, read-only lookup, and a separately approved reservation.

**Recovery:** Return a clear error or uncertain outcome. Retry only when the operation is safe to repeat; a timeout is not proof that a reservation or update failed.

**Adapt it:** Use your existing APIs, business functions, or instrument abstractions. The tool schema should express the operation your application can reliably support.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an English request into a proposed call to an application function, then inspect the result returned to the model. Distinguish choosing a function from the application actually executing it.

**Assumptions:** The function contract defines arguments and results. The application must still enforce access and handle invalid or failed calls.

**Design choices:** Expose narrow, useful operations rather than forcing the model to assemble fragile low-level steps. Validate arguments and distinguish lookup operations from actions with side effects.

**Request:** Check whether the library has this book, without placing a hold.

**Starting evidence:** Mock tools: search_catalog and place_hold. Catalog shows one copy available.

**Action and control:** Propose search_catalog with title/author, validate arguments, and return the lookup result.

**Stage records (authored, not executed):**

### Input record

Mock tools: search_catalog and place_hold. Catalog shows one copy available.

What changed: Establish the facts supplied for this version of the task.

### Design note

Expose narrow, useful operations rather than forcing the model to assemble fragile low-level steps. Validate arguments and distinguish lookup operations from actions with side effects.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Propose search_catalog with title/author, validate arguments, and return the lookup result.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

One copy listed as available. A hold would be a separate authorized action.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Inspect tool name, arguments, read/write classification, and action authorization.

If the result falls short:
Return a clear error or uncertain outcome. Retry only when the operation is safe to repeat; a timeout is not proof that a reservation or update failed.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use your existing APIs, business functions, or instrument abstractions. The tool schema should express the operation your application can reliably support.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** One copy listed as available. A hold would be a separate authorized action.

**Change something — Automatically call place_hold after lookup:** That exceeds the stated scope. The application should refuse the write action.

**Decision:** Does a read request imply authority to reserve the book?

**Answer:** No; ask before the separate action.

**Why:** Tool selection does not itself establish permission to execute the tool.

**Review criteria:** Inspect tool name, arguments, read/write classification, and action authorization.

**Recovery:** Return a clear error or uncertain outcome. Retry only when the operation is safe to repeat; a timeout is not proof that a reservation or update failed.

**Adapt it:** Use your existing APIs, business functions, or instrument abstractions. The tool schema should express the operation your application can reliably support.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an English request into a proposed call to an application function, then inspect the result returned to the model. Distinguish choosing a function from the application actually executing it.

**Assumptions:** The function contract defines arguments and results. The application must still enforce access and handle invalid or failed calls.

**Design choices:** Expose narrow, useful operations rather than forcing the model to assemble fragile low-level steps. Validate arguments and distinguish lookup operations from actions with side effects.

**Request:** Read the latest archived temperature measurement for channel C.

**Starting evidence:** Allowed tool: query_archived_reading. Mock response: 24.1 °C recorded at 09:00. Live instrument tool exists but is not authorized.

**Action and control:** Validate channel and query the stored record, preserving timestamp and units.

**Stage records (authored, not executed):**

### Input record

Allowed tool: query_archived_reading. Mock response: 24.1 °C recorded at 09:00. Live instrument tool exists but is not authorized.

What changed: Establish the facts supplied for this version of the task.

### Design note

Expose narrow, useful operations rather than forcing the model to assemble fragile low-level steps. Validate arguments and distinguish lookup operations from actions with side effects.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Validate channel and query the stored record, preserving timestamp and units.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Archive returns 24.1 °C at 09:00. This is not a live measurement.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Check tool identity, channel, timestamp, units, and absence of a live instrument call.

If the result falls short:
Return a clear error or uncertain outcome. Retry only when the operation is safe to repeat; a timeout is not proof that a reservation or update failed.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use your existing APIs, business functions, or instrument abstractions. The tool schema should express the operation your application can reliably support.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Archive returns 24.1 °C at 09:00. This is not a live measurement.

**Change something — Model proposes live_read_temperature instead:** Reject the unauthorized tool even if it could provide fresher data. Explain the freshness limit.

**Decision:** Does the desire for fresh data authorize instrument access?

**Answer:** No; retain the tool and permission boundary.

**Why:** A helpful tool proposal remains subject to execution policy.

**Review criteria:** Check tool identity, channel, timestamp, units, and absence of a live instrument call.

**Recovery:** Return a clear error or uncertain outcome. Retry only when the operation is safe to repeat; a timeout is not proof that a reservation or update failed.

**Adapt it:** Use your existing APIs, business functions, or instrument abstractions. The tool schema should express the operation your application can reliably support.

Function calling gives the model a fixed list of actions your code defined, each with a name, a
description and an argument schema, and lets it choose whether to use one, which one, and what to
put in the arguments. Anthropic calls the same mechanism tool use: "Claude determines
when to call a tool based on the user's request and the tool's description. It then returns a
structured call that your application executes (client tools) or that Anthropic executes (server
tools)"[1]. The model produces a call, not an effect; this page is about the first of
those two cases, where your own code runs it.

Function calling sits at level 4, tools. The one real choice in a run is the model's: which
action, if any, and with what arguments: the `decided_by: "model"` step in the run below. Your
code decides which actions to offer, runs whichever one gets chosen, and asks for the final
answer. That is the line to level 5: here the model chooses once, inside a run your code bounds;
a single agent chooses again after every result, and leaves the loop only when it decides to.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 4 · Function calling. The same steps are described in the sections below._

## Practical guidance

Look in your chat app's settings, or an icon near the message box, for connectors, plugins,
actions or tools: the feature that lets the model do something beyond answering from what it
already knows. Custom GPTs with actions, Claude connectors and Gemini connected apps are the same
mechanism wearing a product name: the app offers the model a list of actions, the model picks one
when the request calls for it, and the product carries it out. OpenAI is retiring one of those:
its own FAQ names a "Scheduled retirement" on which "Custom GPTs stop running" for affected
workspaces, though "the dates are subject to change"[6], so check what a connector
migrates to before building a habit around it.

Before turning one on, read what it can actually do, not just its name. A tool that only looks
something up (today's weather, an account balance) fails safely if the model reaches for it by
mistake. A tool with a real consequence sits right next to it on the same permission screen and
looks just as ordinary: OpenAI's own examples range from "Get today's weather for a location" to
"Issue refunds for a lost order"[4]. Read every tool in the list this way, not only the
one you meant to add.

To check whether a connector actually ran rather than the model answering from memory, ask
something it could only get right by checking: what is actually on your calendar on a specific
day, not what a typical day would look like. Whether anything gets called at all is a judgment
call the app makes on its own: it "calls a tool when the request maps to that tool's described
capability and the answer isn't already in context"[1]. A fact that could only have come
from your real data means it worked; a generic-sounding answer with no sign anything ran means it
guessed instead, and the fix is usually to ask more specifically, naming the exact thing to check.

For anything with a side effect, look for a screen that shows the action and its arguments before
it runs, not just the final answer; skipping that is skipping the one point where you could still
say no.

If the single thing you need is already a button in the app itself, skip the connector and use
the button.

## Implementation details

All three makers' APIs in this page's sources have the same two halves, so the code below is the
shape you write against any of them. Google states the division plainly: "The model doesn't
execute the function itself. Extract the name and args and execute in your
application"[5].

The example offers two tools, `search(query)` and `lookup_part(part_number)`, and lets the model
call at most one of them. The system prompt says so directly, and the code enforces it a second
way that does not depend on the model following instructions: only `first.tool_calls[0]` is ever
run, and the follow-up call that asks for a final answer is not given the `tools` list at all, so
there is nothing left for the model to call even if it wanted to. That second fact is where level
4 stops and level 5 starts: capping the run at one action is a decision your code made in
advance, not one the model makes about when to stop.

`examples/function_calling/run.py` (lines 39-95)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
    del embedder  # level 4 retrieves through its tools, not a vector index
    sections = load_sections(corpus_dir)
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
    tracer.record(kind="code", decided_by="code", title="Build prompt with tool definitions", detail="search, lookup_part")

    first = model.complete(messages, tools=TOOLS, max_tokens=300)
    if not first.tool_calls:
        tracer.record(
            kind="model",
            decided_by="model",
            title="Model answers directly, no tool call",
            detail=first.text[:200],
            tokens_in=first.tokens_in,
            tokens_out=first.tokens_out,
            ms=first.ms,
        )
        return Answer.from_text(first.text)

    call = first.tool_calls[0]
    # the prompt allows one call; if the model asked for more, the code drops the rest, and the
    # trace has to say so rather than quietly showing a tidier run than the one that happened
    dropped = "" if len(first.tool_calls) == 1 else f" (dropped {len(first.tool_calls) - 1} further call(s))"
    tracer.record(
        kind="model",
        decided_by="model",
        title=f"Model calls {call.name}",
        detail=json.dumps(call.arguments, sort_keys=True) + dropped,
        tokens_in=first.tokens_in,
        tokens_out=first.tokens_out,
        ms=first.ms,
    )
    result_text, citations = _run_tool(call, sections)
    tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])

    follow_up = messages + [
        Message(role="assistant", content=f"[called {call.name}({json.dumps(call.arguments)})]"),
        Message(role="user", content=f"Tool result:\n{result_text}\n\nNow answer the question: {question}"),
    ]
    final = model.complete(follow_up, max_tokens=400)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask for a final answer",
        detail=final.text[:200],
        tokens_in=final.tokens_in,
        tokens_out=final.tokens_out,
        ms=final.ms,
    )
    return Answer.from_text(final.text, retrieved_sources=citations)
```

If the model's first response carries no tool call, that is still the one model decision this
level records: it chose to answer directly rather than to act, the same kind of choice as the
stop at level 5. If it asks for more than one tool at once (Anthropic notes that "by default,
Claude may call multiple tools in a single response"[3]) the code drops every call
after the first and says so in the trace, rather than silently running one and hiding that a
second was requested.

`_run_tool` is the closest thing here to validating a call before acting on it: it checks the
tool's name against the two it knows and falls through to `unknown_tool` for anything else,
which reports the mismatch as text instead of raising. It does not check the *shape* of the
arguments (`part_number` is read with a plain default, not checked against the schema) which
is what a maker's own schema-conformance feature is for. Anthropic's tip on the same page as the
quote above: "Add `strict: true` to your custom tool definitions to ensure Claude's tool calls
always match your schema exactly"[1]. OpenAI documents the same idea for its own
schema: "Setting `strict` to `true` will ensure function calls reliably adhere to the function
schema, instead of being best effort"[4].

This site's electronics-test bench (`docs/THE-BENCH.md`) draws the same line one step further.
`examples/common/bench.py` sorts every instrument command into three classes, not two:
`READ_ONLY_HEADERS` and `is_read_only()` let an agent run any query on its own because a query
changes nothing; a command that sets a value needs code to check it against a safety envelope
first; and a command that energizes the board needs that check plus a person's `Approval` naming
the set point, the same across a production test, an engineering sweep or a precise measurement.
`_run_tool`'s two-tool whitelist is only the first of those three lines, drawn for a lookup and a
part search; a tool that could turn something on would need the other two as well. Run it
yourself:

`examples/function_calling/README.md` (lines 16-16)

```text
python -m examples.function_calling --model stub:scripted
```

## When you do not need this

Try [routing](/gradient_ascent/techniques/routing/) first if you already know, from the
question's surface form, which single action applies: the model does not need to choose an
action your code can already tell apart. Try [prompt
chaining](/gradient_ascent/techniques/prompt-chaining/), or a plain conditional, if the set of actions is small and always runs in the
same order regardless of what the model says.

Move up to function calling once the right action depends on something only the model can judge
from open-ended input (which of several tools applies, or whether none does) and getting that
judgment wrong sometimes is cheap enough to tolerate.

Two siblings at this level change where the action comes from rather than who chooses it. Reach
for [MCP](/gradient_ascent/techniques/mcp/) when more than one application needs the same tools,
or the tools should come from a server you did not write: the model's decision is identical, and
what moves is the boundary the call crosses. Reach for [code execution](/gradient_ascent/techniques/code-execution/) when the thing you need done cannot be
enumerated in advance as a list of named actions at all.

## Failure modes

### Confident call, wrong or missing argument

- **How to notice it:** The model calls the right tool but the argument does not match anything real (a part number that was never in the parts list, a query that does not resemble the question) and the tool answers with whatever it was actually handed rather than what the reader meant.
- **How to test for it:** Ask about a part number that does not exist and confirm the tool reports it as not found, rather than the model inventing a price to go with a citation that never backed one.

### Extra tool calls silently dropped

- **How to notice it:** The model asks for more than one action in a single turn, and only the first one visibly happens, with nothing telling you a second request existed at all.
- **How to test for it:** Script a model response with two tool calls and check the trace records that the extra one was dropped, not just that the first one ran.

### A tool result is treated as trustworthy text

- **How to notice it:** A search result or a lookup can carry text written to look like an instruction, and nothing about being a tool result rather than a user message stops the model from reading it as one.
- **How to test for it:** Anthropic's own guidance: "an attacker who can influence it may embed instructions that try to redirect Claude (indirect prompt injection)". Add a document section with an embedded instruction to the corpus and see whether a search that surfaces it changes the answer to match it.

### The model stops reaching for the tool at all

- **How to notice it:** Across many similar questions, the share that get a tool call drops toward zero even though the documents still hold the answer, because the tool’s description drifted out of sync with what people actually ask.
- **How to test for it:** Track how often decided_by: "model" ends in a tool call versus a direct answer over a batch of known-lookup questions; a falling rate with no change in the questions is a description problem, not a model problem.

### A tool is more powerful than the question needed

- **How to notice it:** The call that ran was the right one, on the right input, but the tool itself could do more than this question ever required, so a future mistaken call has a bigger blast radius than a wrong answer.
- **How to test for it:** List every tool offered for a given prompt and check whether each one's effect (what it can change, not just what it can read) matches what that prompt's questions actually need.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, tool used:** 2
- **Model calls, no tool needed:** 1
- **Tokens in, tool-call turn:** ~190
- **Tokens out, tool-call turn:** ~14

**Compared with RAG (level 2).** RAG always retrieves and always calls the model once, so every question costs the same. Function calling asks the model first, so a question it can already answer costs one call instead of two; the second call only happens when the model decides a tool is needed.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Function calling is one of the examples the site's own 60-question set scores directly (see
`docs/EVALS.md`), the same way RAG and routing are: it answers a question about the documents and
cites what it used. Level 4 adds one more thing worth checking beyond the answer itself: a result
file's `model_decided_steps` should equal the number of questions run, exactly: one model
decision per question, never zero and never more than one. A number outside that range means the
trace is wrong before the answer is even graded.

Unlike routing, whose `--dry` projection follows the cheapest branch because the token-counting
stand-in cannot produce a label the code recognizes, this example's branch depends on whether
`tools` were offered, not on parsed text, and the stand-in calls the first tool every time tools
are offered, so `--dry` already projects the tool-calling branch, the more expensive one.

No result file exists for function calling yet. Run `python scripts/eval_run.py --example
function_calling --model <spec> --dry` to project the cost of a real run before spending anything
on one.

## Run it

**What to monitor.** The share of questions where decided_by is model that end in a tool call versus a direct answer, tracked over time against a batch of questions you know should call a tool. A falling rate with no change in the traffic is the tool description going stale, not the model getting worse.

**Cost at volume.** Every question pays for at least one call. Only the ones where the model reaches for a tool pay for the second, so cost tracks how often real traffic actually needs an action, not a fixed number per question.

**How it fails in production.** A tool with a side effect runs on an argument the model half-guessed, because nothing between the model's call and the tool's execution checked that the argument was real. A second tool call the model asked for is dropped with no record, and a person debugging a wrong answer has no way to know one was ever requested.

**What to log.** The full list of tools offered, which one (if any) was called and with what arguments, the raw tool result, and whether decided_by was model or code for every step, so a wrong answer traces back to the wrong tool, the wrong argument, or the model declining to act at all.

## Try it

1. **Use it.** Find a chat app feature that can act beyond answering: a connector, a plugin, a custom action. Ask it something the action does not cover: does it say it could not act, or quietly answer from memory?
2. **Build it.** Run python -m examples.function_calling --model stub:scripted from the repo root: the model calls lookup_part on HLV-2205, the code runs it, the answer cites parts-list#2. Change that part number in SCRIPTED (examples/function_calling/__main__.py) to HLV-9999: the tool reports it not found, the next reply still prices it at $52.00, and the citations line empties.
3. **Either lane.** That is the first failure mode above; cause another the same way, from evals/corpus/ only.


## Sources

1. [Tool use with Claude](https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview) — Anthropic (accessed 2026-09-19)
2. [Handle tool calls](https://platform.claude.com/docs/en/agents-and-tools/tool-use/handle-tool-calls) — Anthropic (accessed 2026-09-19)
3. [Parallel tool use](https://platform.claude.com/docs/en/agents-and-tools/tool-use/parallel-tool-use) — Anthropic (accessed 2026-09-19)
4. [Function calling](https://developers.openai.com/api/docs/guides/function-calling) — OpenAI (API documentation) (accessed 2026-09-19)
5. [Function calling with the Gemini API (archived copy)](https://web.archive.org/web/20260915180415id_/https://ai.google.dev/gemini-api/docs/function-calling) — Google (Gemini API documentation, via the Internet Archive) (accessed 2026-09-19)
6. [Custom GPT retirement and migration FAQ (archived copy)](https://web.archive.org/web/20260918151544id_/https://help.openai.com/en/articles/20001519-custom-gpt-retirement-and-migration-faq) — OpenAI (help center, via the Internet Archive) (accessed 2026-09-19)


Last reviewed 2026-09-19.
