All three makers’ APIs in this page’s sources have the same two halves, so the code below is the
shape you write against any of them. Google states the division plainly: “The model doesn’t
execute the function itself. Extract the name and args and execute in your
application”[5].
The example offers two tools, search(query) and lookup_part(part_number), and lets the model
call at most one of them. The system prompt says so directly, and the code enforces it a second
way that does not depend on the model following instructions: only first.tool_calls[0] is ever
run, and the follow-up call that asks for a final answer is not given the tools list at all, so
there is nothing left for the model to call even if it wanted to. That second fact is where level
4 stops and level 5 starts: capping the run at one action is a decision your code made in
advance, not one the model makes about when to stop.
examples/function_calling/run.py · lines 39–95
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
) -> Answer:
del embedder # level 4 retrieves through its tools, not a vector index
sections = load_sections(corpus_dir)
messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
tracer.record(kind="code", decided_by="code", title="Build prompt with tool definitions", detail="search, lookup_part")
first = model.complete(messages, tools=TOOLS, max_tokens=300)
if not first.tool_calls:
tracer.record(
kind="model",
decided_by="model",
title="Model answers directly, no tool call",
detail=first.text[:200],
tokens_in=first.tokens_in,
tokens_out=first.tokens_out,
ms=first.ms,
)
return Answer.from_text(first.text)
call = first.tool_calls[0]
# the prompt allows one call; if the model asked for more, the code drops the rest, and the
# trace has to say so rather than quietly showing a tidier run than the one that happened
dropped = "" if len(first.tool_calls) == 1 else f" (dropped {len(first.tool_calls) - 1} further call(s))"
tracer.record(
kind="model",
decided_by="model",
title=f"Model calls {call.name}",
detail=json.dumps(call.arguments, sort_keys=True) + dropped,
tokens_in=first.tokens_in,
tokens_out=first.tokens_out,
ms=first.ms,
)
result_text, citations = _run_tool(call, sections)
tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
follow_up = messages + [
Message(role="assistant", content=f"[called {call.name}({json.dumps(call.arguments)})]"),
Message(role="user", content=f"Tool result:\n{result_text}\n\nNow answer the question: {question}"),
]
final = model.complete(follow_up, max_tokens=400)
tracer.record(
kind="model",
decided_by="code",
title="Ask for a final answer",
detail=final.text[:200],
tokens_in=final.tokens_in,
tokens_out=final.tokens_out,
ms=final.ms,
)
return Answer.from_text(final.text, retrieved_sources=citations)
If the model’s first response carries no tool call, that is still the one model decision this
level records: it chose to answer directly rather than to act, the same kind of choice as the
stop at level 5. If it asks for more than one tool at once (Anthropic notes that “by default,
Claude may call multiple tools in a single response”[3]) the code drops every call
after the first and says so in the trace, rather than silently running one and hiding that a
second was requested.
_run_tool is the closest thing here to validating a call before acting on it: it checks the
tool’s name against the two it knows and falls through to unknown_tool for anything else,
which reports the mismatch as text instead of raising. It does not check the shape of the
arguments (part_number is read with a plain default, not checked against the schema) which
is what a maker’s own schema-conformance feature is for. Anthropic’s tip on the same page as the
quote above: “Add strict: true to your custom tool definitions to ensure Claude’s tool calls
always match your schema exactly”[1]. OpenAI documents the same idea for its own
schema: “Setting strict to true will ensure function calls reliably adhere to the function
schema, instead of being best effort”[4].
This site’s electronics-test bench (docs/THE-BENCH.md) draws the same line one step further.
examples/common/bench.py sorts every instrument command into three classes, not two:
READ_ONLY_HEADERS and is_read_only() let an agent run any query on its own because a query
changes nothing; a command that sets a value needs code to check it against a safety envelope
first; and a command that energizes the board needs that check plus a person’s Approval naming
the set point, the same across a production test, an engineering sweep or a precise measurement.
_run_tool’s two-tool whitelist is only the first of those three lines, drawn for a lookup and a
part search; a tool that could turn something on would need the other two as well. Run it
yourself:
examples/function_calling/README.md · lines 16–16
python -m examples.function_calling --model stub:scripted