Level 04 · Tool use

Code execution

Letting the model write code and run it in a sandbox.

Sourced

Concept at a glance

Turn a generated program into a checked result.

SequenceConceptual illustration
Turn a generated program into a checked result.Model writes code leads to Sandbox runs it. Sandbox runs it leads to Result + errors. The execution environment needs its own limits; generated code is not automatically safe.Model writes codeA program for the taskSandbox runs itBounded executionResult + errorsInspect the actual outputTurn a generated program into a checked result.Model writes code leads to Sandbox runs it. Sandbox runs it leads to Result + errors. The execution environment needs its own limits; generated code is not automatically safe.Model writes codeA program for the taskSandbox runs itBounded executionResult + errorsInspect the actual output
Read the connections in words
  • Model writes code → Sandbox runs it: Bounded execution.
  • Sandbox runs it → Result + errors: Inspect the actual output.
Key idea

The execution environment needs its own limits; generated code is not automatically safe.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Code execution: see it in practice.

Running generated code in a constrained environment to perform computation or transform data.

What you’ll walk through

Follow a computation from supplied data through a small program into a checkable result. Inspect units, assumptions, and evidence of execution separately from generated code.

The task in this version

Compute total energy from these readings and show the units.

What you’ll learn to check

Input preview, inspectable Python, deterministic totals, a planted unit error, and denied file/network access in a mock boundary demonstration.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Compute total energy from these readings and show the units.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
CSV fixture: 500 Wh, 750 Wh, 250 Wh. Output unit: kWh.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Inputs and units must be defined. Code that looks plausible may not have run, and running successfully does not prove the calculation is appropriate.

1 / 6

Apply this to your project

Describe your task to your own model and use Code execution as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Code execution lets the model write a small program instead of choosing among named tools, and a sandbox (not the model) runs it. The shape is the same as function calling: the model’s output picks what happens next, and your code always carries it out. Anthropic describes its own version this way: the tool “allows Claude to run Bash commands and manipulate files, including writing code, in a secure, sandboxed environment”[1]. What the model writes is data your code hands to an interpreter, never text your code trusts and runs directly.

Code execution sits at level 4, tools. The model’s one real choice is what code to write; the decided_by: "model" step below is picking that content, not deciding whether to run it: your code always runs whatever it wrote, inside a fixed sandbox, and always asks once more for an answer once the result is back. That fixed shape, one write-and-run cycle bounded from outside, is the line to level 5: a coding agent keeps writing and running code in a loop it exits on its own, reading each result before deciding what to write next.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Code execution

The model writes one arithmetic expression; a restricted sandbox always runs it.

Level 4 · Tool use
QuestionQuestionsearch for numberssearch for numbersMODELwrites an expressionwrites anexpressionTOOLevaluates the expressionevaluates theexpressionMODELanswers using the resultanswers usingthe resultAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

The question arrives

"What is the total price to replace the heating elements
on both a DW-300 and a DW-480?"
0 tokens · 0 ms

Practical guidance

This is the feature in ChatGPT’s data analysis, Gemini Notebook and Mistral Vibe that runs actual Python on data you give it, instead of guessing a number from the words in your message. Paste a table or upload a spreadsheet, then ask something concrete: “Add up the total in the amount column, and show me a bar chart of totals by month.” A model asked to just compute that from the text of your message can get arithmetic wrong; a model that writes and runs code cannot, because the number comes from the code executing, not from a token prediction.

What makes this safe to try at all is what the sandbox refuses to do, not what it lets the model write. Anthropic’s own container has “Internet access: Completely disabled for security”[1]; OpenAI’s runs the same way, inside “a fully sandboxed virtual machine that the model can run Python code in”[2], with a fixed memory limit and a session that expires “if it is not used for 20 minutes”[2]. Neither can reach your email, your bank, or anything else on the internet, whatever the code says.

Get in the habit of asking to see the code, not just the number or the chart: most of these products have a button or an expandable section for it. You do not need to read Python fluently to check the shape of it: does it use the column you actually asked about, and does the final number come from a calculation in the code rather than a sentence typed after it? A total that changes when you ask the same question twice, or a chart with no code shown next to it, is a sign the answer was written rather than computed.

The same habit answers the question of trust for a credential, too: if a connected tool can reach a service that needs a password, that key has to live somewhere the code cannot read and only the network call can use, the way a sandbox product outside the chat apps, E2B, describes its own “Secrets vault” as “Keys your agent can use but never read”[3].

If the file is small enough to eyeball, or the calculation is one you would trust a spreadsheet formula to do, that is faster than typing a prompt for it: open the spreadsheet.

Implementation details

The example answers a numeric question by writing one arithmetic expression instead of prose. The model never gets Python; it gets a system prompt that asks for exactly one line (an expression using numbers, + - * / and parentheses, or the word NONE if the retrieved passages don’t have what it needs) and the code is the only thing that ever runs it:

examples/code_execution/run.py · lines 25–31
SYSTEM_PROMPT = (
    "You answer numeric questions about Halvorsen appliances using only the numbered sources "
    "below. Reply with exactly one line: a Python arithmetic expression using only numbers, "
    "+ - * / and parentheses, that computes the answer -- no words, no units, no code fences. "
    f"If the sources do not contain the numbers you would need, reply with the single word "
    f"{DECLINE} instead."
)

This is an allow-list, not a blocklist. A blocklist has to name every dangerous spelling in advance; _eval_node instead names the handful of node types arithmetic actually needs (constants, the four operators, unary plus and minus) and falls through to UnsafeExpression for anything else. A function call, a name lookup, an attribute access and an import all fail the same way, because none of them is a node type this function ever matches. That is why the whitelist is the point: it does not have to recognize an attack to stop it.

The whitelist itself is one function, and it is the whole defense, so here it is rather than a description of it. Every node type the task needs is named; anything else raises before it can run, which is why the check does not have to recognize an attack in order to stop one:

examples/code_execution/run.py · lines 71–80
def _eval_node(node: ast.AST, depth: int = 0) -> float:
    if depth > MAX_DEPTH:
        raise UnsafeExpression(f"expression nests deeper than {MAX_DEPTH} levels")
    if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)) and not isinstance(node.value, bool):
        return node.value
    if isinstance(node, ast.BinOp) and type(node.op) in _BINOPS:
        return _BINOPS[type(node.op)](_eval_node(node.left, depth + 1), _eval_node(node.right, depth + 1))
    if isinstance(node, ast.UnaryOp) and type(node.op) in _UNARYOPS:
        return _UNARYOPS[type(node.op)](_eval_node(node.operand, depth + 1))
    raise UnsafeExpression(f"{type(node).__name__} is not on the arithmetic whitelist")

Three bounds sit around that whitelist, because the node type is not the only way an expression can be a problem. The text is capped before it is parsed. The walk is capped by depth, so a long chain (1+1+1+...) is refused rather than running Python out of stack and raising an error this function never promised. And a result that is not a finite number is refused, so 1e400 and 1e308 * 1e308 come back as refusals rather than as inf:

examples/code_execution/run.py · lines 59–68
    if len(expr) > MAX_EXPRESSION_CHARS:
        raise UnsafeExpression(f"expression is {len(expr)} characters; the limit is {MAX_EXPRESSION_CHARS}")
    try:
        tree = ast.parse(expr, mode="eval")
    except (SyntaxError, ValueError) as exc:
        raise UnsafeExpression(f"not a valid expression: {exc}") from exc
    value = _eval_node(tree.body)
    if isinstance(value, float) and not math.isfinite(value):
        raise UnsafeExpression(f"result is not a finite number: {value}")
    return value

That last bound is on the result, not on the arithmetic, and the difference is worth knowing before you copy it: 1/1e400 overflows in the middle and comes back as 0.0, a finite number, so it is returned like any other. Nothing here is wrong with the figure (it is the value Python computes) but a reader is not told that an intermediate went to infinity on the way.

The run itself checks the model’s reply, evaluates it if it looks like one, and hands back either a computed answer or a plain refusal:

examples/code_execution/run.py · lines 112–143
    first = model.complete(messages, max_tokens=60)
    expr = first.text.strip()

    if not expr or expr.upper() == DECLINE:
        tracer.record(
            kind="model",
            decided_by="model",
            title="Model declines: not enough numbers in the passages",
            detail=expr or "(empty)",
            tokens_in=first.tokens_in,
            tokens_out=first.tokens_out,
            ms=first.ms,
        )
        return Answer(text="The documents don't give enough numbers to compute that.", citations=[])

    tracer.record(
        kind="model",
        decided_by="model",
        title="Model writes an expression",
        detail=expr,
        tokens_in=first.tokens_in,
        tokens_out=first.tokens_out,
        ms=first.ms,
    )

    try:
        value = safe_eval(expr)
    except (UnsafeExpression, ArithmeticError) as exc:
        tracer.record(kind="code", decided_by="code", title="Sandbox refused the expression", detail=str(exc))
        return Answer(text=f"Could not safely evaluate that expression: {exc}", citations=[])

    tracer.record(kind="code", decided_by="code", title="Sandbox evaluates the expression", detail=f"{expr} = {value}")

Declining (NONE) and writing a real expression are both the one decided_by: "model" step this level records. This is the same rule as function calling’s tool-or-not choice. A refusal from the sandbox is decided_by: "code", same as running a tool: the model already made its one decision by the time the sandbox looks at what it wrote, so rejecting the content is code’s call, not a second model decision. Run it yourself:

examples/code_execution/README.md · lines 17–17
python -m examples.code_execution --model stub:scripted
When you do not need this

Try function calling first if the set of things you would want the model to do can be named and typed in advance as a short list of tools: most of the time it can, and a named tool is easier to log, test and limit than an open-ended expression. Try a fixed formula in your own code if the calculation itself never varies with the question.

Move up to code execution once the computation depends on numbers or an operation you cannot enumerate in advance as a fixed set of named actions: arbitrary arithmetic, reshaping a table, a chart built from data given at question time.

Failure modes

The expression is safe but uses the wrong numbers

How to notice it
The sandbox happily evaluates 38.50 + 46.00 (a real result, cited to a real section) but the two numbers came from different products than the question asked about, because nothing checks that an expression only uses numbers the retrieved passages actually named for that product.
How to test for it
Ask about two similarly priced parts from different models and check that the cited section actually contains both numbers used, not just numbers that happen to appear somewhere in the retrieved passages.

The needed numbers were never retrieved

How to notice it
The model declines, correctly, because the passages it was given do not contain a number it needs, but a different search would have found it. A safe decline still means a right answer nobody got.
How to test for it
Ask a numeric question whose figures live in a section a keyword search ranks below the cutoff, and check whether the run declines instead of computing a wrong number from partial information.

The whitelist is too strict for a legitimate question

How to notice it
A question that genuinely needs an operation outside plus, minus, times and divide (a percentage, a power, a square root) gets a decline that looks like a security block but is really a missing feature. The length and depth bounds do the same for a legitimate sum with too many terms in it.
How to test for it
Ask a question whose arithmetic needs a percentage or an exponent and confirm the run declines cleanly, rather than the model trying to fake the operation with what is allowed. Read the refusal text: it names which bound was hit, so a missing operator and an over-long expression do not look alike in a log.

A real sandbox is given more reach than the question needs

How to notice it
This example's evaluator cannot make a network call or read a file no matter what the model writes, because those node types are not on the whitelist at all. A real code-execution sandbox that runs actual Python can, unless its own network and filesystem limits are configured as tightly as the question needs.
How to test for it
For a real sandbox, check its documented network and file-access limits directly rather than assuming a model's own caution will substitute for them.

An expression built to escape the whitelist

How to notice it
A written expression tries to reach a name, a call or an attribute (the pattern behind most real sandbox escapes) and has to be refused the same way a harmless typo is, before anything runs.
How to test for it
Script a model response that writes __import__('os').system(...) and confirm the run refuses it and never calls Python's own eval or exec; see tests/test_example_code_execution.py.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

2Model calls, valid expression
1Model calls, declines
~230Tokens in, expression turn
~8Tokens out, expression turn
Compared with Function calling (level 4)Function calling picks from a short, named list of actions. Code execution picks the content of an expression instead, so the same one-decision shape can answer a much wider range of questions without a new tool being defined for each one.

How to Evaluate It

12 questionsnumeric

Code execution is registered with the site’s runner as not scored against the shared 60-question set, with the reason (see docs/EVALS.md). Asking scripts/eval_run.py for it by name prints that reason and stops, rather than producing a number about a task the technique was never built to do. What it would be scored on instead: correctness on the 12 numeric questions in the shared set, the only kind this technique answers by design, plus two numbers the shared harness does not otherwise track: how often a deliberately malicious or malformed expression is refused rather than evaluated, and how cleanly the other four question kinds (lookup, multi-hop, unanswerable, conflicting sources) are declined instead of forced through an arithmetic answer that was never going to fit them.

Run it

What to monitor

The refusal rate on the sandbox step, tracked separately from the decline rate on the model step. A refusal means the model wrote something outside the whitelist; a decline means it correctly said the passages did not have enough numbers. Conflating the two hides whether a rising number is a prompting problem or a retrieval problem.

Cost at volume

Every question pays for at least one call. A second call, and the sandbox's own negligible cost, only happen when the model actually wrote something worth running, so cost tracks how often real traffic needs a computation, not a fixed number per question.

How it fails in production

A question whose real answer needs an operation off the whitelist (a percentage, a root) gets declined and reads as the documents lacking an answer, when the documents had everything but the sandbox lacked the operator. A retrieved passage has the wrong document's numbers in it, and the expression computes a real, wrong number with a real-looking citation.

What to log

The retrieved passages and their citations, the raw text the model wrote, whether it parsed as a safe expression or was refused and why, the computed value, and the final answer, so a wrong number traces back to retrieval, to the expression, or to the sandbox's own arithmetic.

Try it

  1. Use it

    Ask a chat app's data-analysis feature a question that needs a calculation across numbers you give it, then ask it to show you the code it ran. Does the code actually compute what the final answer claims?

  2. Build it

    Run python -m examples.code_execution --model stub:scripted from the repo root. Retrieval finds the two prices, the model writes 38.50 + 41.00, the sandbox evaluates it, and the answer is $79.50 with its citation. Run it again with --model stub: the echoed question goes to the sandbox instead of an expression and is refused as invalid syntax, which is the point, since nothing the model returns is trusted to be arithmetic. Then open a Python shell in the repo root and call examples.code_execution.run.safe_eval on a few strings of your own: "38.50 + 41.00", "2 ** 10", "__import__('os').listdir('.')", and "1" plus "+1" forty times. Read which bound each one hits.

  3. Either lane

    Pick one of the failure modes above and try to cause it on purpose, using only the synthetic documents in evals/corpus/.

  4. Build it

    Read examples/bench_test_data_by_conversation/README.md: on this site's electronics-test bench, a model writes a short snippet against a retest export whose volts column is actually millivolts under a header that says volts, and a sandbox runs it. Before that snippet or the model ever sees the table, code range-checks the column against the widest node the board has anywhere and corrects the mislabeled unit. Why does that check have to happen in code, on every table, whether or not the model's tool gets called at all, rather than being one more thing the written snippet is trusted to do?

How it connects

Before, after and instead of this

Move up when

  • Coding agentsThe code has to be run, read and rewritten until it works, with the model deciding when it is done.

Pages that need this one

Optional: products, tools, and models

6 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Calculate from a test log

The model writes a small calculation, a constrained environment runs it, and you inspect the result and errors.

Out there

Named products, tools and models

Products4
  • ChatGPT data analysisOpenAI · code execution in a chat app
  • Gemini NotebookGoogle · research notebook · formerly NotebookLM
  • Mistral VibeMistral AI · ai agent for work and coding · formerly Le Chat
  • n8nn8n · automation service, self-hostable
Tools2
  • E2BE2B · code sandbox
  • ModalModal · code sandbox and compute

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Code execution tool · Anthropic (accessed 09/19/2026)
  2. Code Interpreter · OpenAI (API documentation) (accessed 09/19/2026)
  3. E2B · E2B (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page