# Code execution

_Level 04 · Tool use · sourced_

Letting the model write code and run it in a sandbox.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a computation from supplied data through a small program into a checkable result. Inspect units, assumptions, and evidence of execution separately from generated code.

**Assumptions:** Inputs and units must be defined. Code that looks plausible may not have run, and running successfully does not prove the calculation is appropriate.

**Design choices:** Use code for repeatable calculation and transformations. Choose libraries and an execution environment suited to the data and allowed side effects.

**Request:** Compute total energy from these readings and show the units.

**Starting evidence:** CSV fixture: 500 Wh, 750 Wh, 250 Wh. Output unit: kWh.

**Action and control:** Illustrative Python sums 1500 Wh and divides by 1000. This UI does not run arbitrary user code.

**Stage records (authored, not executed):**

### Input record

CSV fixture: 500 Wh, 750 Wh, 250 Wh. Output unit: kWh.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use code for repeatable calculation and transformations. Choose libraries and an execution environment suited to the data and allowed side effects.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Illustrative Python sums 1500 Wh and divides by 1000. This UI does not run arbitrary user code.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Total: 1.5 kWh; verify as 0.5 + 0.75 + 0.25. No filesystem or network execution occurs.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Input preview, inspectable Python, deterministic totals, a planted unit error, and denied file/network access in a mock boundary demonstration.

If the result falls short:
If inputs are inconsistent or the result is implausible, inspect intermediate values and compare with a simple independent check. Do not hide a failed execution behind a predicted result.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to analysis, conversions, parsing, or plotting. Specify allowed files and operations according to the task; an isolated calculator needs fewer controls than a system-changing script.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Total: 1.5 kWh; verify as 0.5 + 0.75 + 0.25. No filesystem or network execution occurs.

**Change something — Mislabel 750 kWh as 750 Wh:** Correct arithmetic on incorrectly interpreted units is wrong. Resolve units before calculating.

**Decision:** Does sandboxing establish numerical correctness?

**Answer:** No; check inputs, units, and arithmetic.

**Why:** Malformed timestamps, missing rows, and unit mismatches affect results; a sandbox bounds access but does not ensure correct math.

**Review criteria:** Input preview, inspectable Python, deterministic totals, a planted unit error, and denied file/network access in a mock boundary demonstration.

**Recovery:** If inputs are inconsistent or the result is implausible, inspect intermediate values and compare with a simple independent check. Do not hide a failed execution behind a predicted result.

**Adapt it:** Apply this to analysis, conversions, parsing, or plotting. Specify allowed files and operations according to the task; an isolated calculator needs fewer controls than a system-changing script.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a computation from supplied data through a small program into a checkable result. Inspect units, assumptions, and evidence of execution separately from generated code.

**Assumptions:** Inputs and units must be defined. Code that looks plausible may not have run, and running successfully does not prove the calculation is appropriate.

**Design choices:** Use code for repeatable calculation and transformations. Choose libraries and an execution environment suited to the data and allowed side effects.

**Request:** Compare unit prices from these grocery package sizes.

**Starting evidence:** A: 500 g for $3. B: 750 g for $4.50. Ignore promotions not supplied.

**Action and control:** Compute price per kilogram using explicit conversions; the calculation shown is a fixture, not arbitrary code execution.

**Stage records (authored, not executed):**

### Input record

A: 500 g for $3. B: 750 g for $4.50. Ignore promotions not supplied.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use code for repeatable calculation and transformations. Choose libraries and an execution environment suited to the data and allowed side effects.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Compute price per kilogram using explicit conversions; the calculation shown is a fixture, not arbitrary code execution.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Both cost $6/kg. Package size alone does not make one a better price.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Recompute the conversions and compare with the original package labels.

If the result falls short:
If inputs are inconsistent or the result is implausible, inspect intermediate values and compare with a simple independent check. Do not hide a failed execution behind a predicted result.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to analysis, conversions, parsing, or plotting. Specify allowed files and operations according to the task; an isolated calculator needs fewer controls than a system-changing script.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Both cost $6/kg. Package size alone does not make one a better price.

**Change something — Read the 750 g label as 750 kg:** The calculation yields a nonsensical price. Validate input units rather than trust a runnable formula.

**Decision:** Does successful execution establish sensible inputs?

**Answer:** No; check units and plausibility.

**Why:** Execution can accurately compute the wrong problem.

**Review criteria:** Recompute the conversions and compare with the original package labels.

**Recovery:** If inputs are inconsistent or the result is implausible, inspect intermediate values and compare with a simple independent check. Do not hide a failed execution behind a predicted result.

**Adapt it:** Apply this to analysis, conversions, parsing, or plotting. Specify allowed files and operations according to the task; an isolated calculator needs fewer controls than a system-changing script.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a computation from supplied data through a small program into a checkable result. Inspect units, assumptions, and evidence of execution separately from generated code.

**Assumptions:** Inputs and units must be defined. Code that looks plausible may not have run, and running successfully does not prove the calculation is appropriate.

**Design choices:** Use code for repeatable calculation and transformations. Choose libraries and an execution environment suited to the data and allowed side effects.

**Request:** Calculate the portfolio's total approved budget from a CSV.

**Starting evidence:** Rows: Atlas $100k, Beacon $80k, Cedar unknown. Amounts are approved budgets, not actual spending.

**Action and control:** Parse units and missing values before aggregation; keep unknown separate from zero.

**Stage records (authored, not executed):**

### Input record

Rows: Atlas $100k, Beacon $80k, Cedar unknown. Amounts are approved budgets, not actual spending.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use code for repeatable calculation and transformations. Choose libraries and an execution environment suited to the data and allowed side effects.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Parse units and missing values before aggregation; keep unknown separate from zero.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Known approved budget totals $180k; portfolio total is incomplete because Cedar is unknown.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Check column meaning, currency/unit consistency, missing rows, and arithmetic.

If the result falls short:
If inputs are inconsistent or the result is implausible, inspect intermediate values and compare with a simple independent check. Do not hide a failed execution behind a predicted result.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to analysis, conversions, parsing, or plotting. Specify allowed files and operations according to the task; an isolated calculator needs fewer controls than a system-changing script.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Known approved budget totals $180k; portfolio total is incomplete because Cedar is unknown.

**Change something — Replace missing Cedar with zero:** The total now looks complete but makes an unsupported assumption.

**Decision:** Should a missing budget be converted to zero silently?

**Answer:** No; preserve the missing value and incomplete total.

**Why:** Numerical pipelines need explicit missing-data semantics.

**Review criteria:** Check column meaning, currency/unit consistency, missing rows, and arithmetic.

**Recovery:** If inputs are inconsistent or the result is implausible, inspect intermediate values and compare with a simple independent check. Do not hide a failed execution behind a predicted result.

**Adapt it:** Apply this to analysis, conversions, parsing, or plotting. Specify allowed files and operations according to the task; an isolated calculator needs fewer controls than a system-changing script.

Code execution lets the model write a small program instead of choosing among named tools, and a
sandbox (not the model) runs it. The shape is the same as function calling: the model's output
picks what happens next, and your code always carries it out. Anthropic describes its own version
this way: the tool "allows Claude to run Bash commands and manipulate files, including writing
code, in a secure, sandboxed environment"[1]. What the model writes is data your code
hands to an interpreter, never text your code trusts and runs directly.

Code execution sits at level 4, tools. The model's one real choice is what code to write; the
`decided_by: "model"` step below is picking that content, not deciding whether to run it: your
code always runs whatever it wrote, inside a fixed sandbox, and always asks once more for an
answer once the result is back. That fixed shape, one write-and-run cycle bounded from outside, is
the line to level 5: a coding agent keeps writing and running code in a loop it exits on its own,
reading each result before deciding what to write next.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 4 · Code execution. The same steps are described in the sections below._

## Practical guidance

This is the feature in ChatGPT's data analysis, Gemini Notebook and Mistral Vibe that runs actual
Python on data you give it, instead of guessing a number from the words in your message. Paste a
table or upload a spreadsheet, then ask something concrete: "Add up the total in the amount
column, and show me a bar chart of totals by month." A model asked to just compute that from the
text of your message can get arithmetic wrong; a model that writes and runs code cannot, because
the number comes from the code executing, not from a token prediction.

What makes this safe to try at all is what the sandbox refuses to do, not what it lets the model
write. Anthropic's own container has "Internet access: Completely disabled for security"[1]; OpenAI's runs the same way, inside "a fully sandboxed virtual machine that the model can
run Python code in"[2], with a fixed memory limit and a session that expires "if it is
not used for 20 minutes"[2]. Neither can reach your email, your bank, or anything else on
the internet, whatever the code says.

Get in the habit of asking to see the code, not just the number or the chart: most of these
products have a button or an expandable section for it. You do not need to read Python fluently to
check the shape of it: does it use the column you actually asked about, and does the final number
come from a calculation in the code rather than a sentence typed after it? A total that changes
when you ask the same question twice, or a chart with no code shown next to it, is a sign the
answer was written rather than computed.

The same habit answers the question of trust for a credential, too: if a connected tool can reach
a service that needs a password, that key has to live somewhere the code cannot read and only the
network call can use, the way a sandbox product outside the chat apps, E2B, describes its own
"Secrets vault" as "Keys your agent can use but never read"[3].

If the file is small enough to eyeball, or the calculation is one you would trust a spreadsheet
formula to do, that is faster than typing a prompt for it: open the spreadsheet.

## Implementation details

The example answers a numeric question by writing one arithmetic expression instead of prose. The
model never gets Python; it gets a system prompt that asks for exactly one line (an expression
using numbers, `+ - * /` and parentheses, or the word `NONE` if the retrieved passages don't have
what it needs) and the code is the only thing that ever runs it:

`examples/code_execution/run.py` (lines 25-31)

```python
SYSTEM_PROMPT = (
    "You answer numeric questions about Halvorsen appliances using only the numbered sources "
    "below. Reply with exactly one line: a Python arithmetic expression using only numbers, "
    "+ - * / and parentheses, that computes the answer -- no words, no units, no code fences. "
    f"If the sources do not contain the numbers you would need, reply with the single word "
    f"{DECLINE} instead."
)
```

This is an allow-list, not a blocklist. A blocklist has to name every dangerous spelling in
advance; `_eval_node` instead names the handful of node types arithmetic actually needs
(constants, the four operators, unary plus and minus) and falls through to `UnsafeExpression` for
anything else. A function call, a name lookup, an attribute access and an import all fail the
same way, because none of them is a node type this function ever matches. That is why the
whitelist is the point: it does not have to recognize an attack to stop it.

The whitelist itself is one function, and it is the whole defense, so here it is rather than a
description of it. Every node type the task needs is named; anything else raises before it can
run, which is why the check does not have to recognize an attack in order to stop one:

`examples/code_execution/run.py` (lines 71-80)

```python
def _eval_node(node: ast.AST, depth: int = 0) -> float:
    if depth > MAX_DEPTH:
        raise UnsafeExpression(f"expression nests deeper than {MAX_DEPTH} levels")
    if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)) and not isinstance(node.value, bool):
        return node.value
    if isinstance(node, ast.BinOp) and type(node.op) in _BINOPS:
        return _BINOPS[type(node.op)](_eval_node(node.left, depth + 1), _eval_node(node.right, depth + 1))
    if isinstance(node, ast.UnaryOp) and type(node.op) in _UNARYOPS:
        return _UNARYOPS[type(node.op)](_eval_node(node.operand, depth + 1))
    raise UnsafeExpression(f"{type(node).__name__} is not on the arithmetic whitelist")
```

Three bounds sit around that whitelist, because the node type is not the only way an expression
can be a problem. The text is capped before it is parsed. The walk is capped by depth, so a long
chain (`1+1+1+...`) is refused rather than running Python out of stack and raising an error this
function never promised. And a result that is not a finite number is refused, so `1e400` and
`1e308 * 1e308` come back as refusals rather than as `inf`:

`examples/code_execution/run.py` (lines 59-68)

```python
    if len(expr) > MAX_EXPRESSION_CHARS:
        raise UnsafeExpression(f"expression is {len(expr)} characters; the limit is {MAX_EXPRESSION_CHARS}")
    try:
        tree = ast.parse(expr, mode="eval")
    except (SyntaxError, ValueError) as exc:
        raise UnsafeExpression(f"not a valid expression: {exc}") from exc
    value = _eval_node(tree.body)
    if isinstance(value, float) and not math.isfinite(value):
        raise UnsafeExpression(f"result is not a finite number: {value}")
    return value
```

That last bound is on the result, not on the arithmetic, and the difference is worth knowing
before you copy it: `1/1e400` overflows in the middle and comes back as `0.0`, a finite number,
so it is returned like any other. Nothing here is wrong with the figure (it is the value Python
computes) but a reader is not told that an intermediate went to infinity on the way.

The run itself checks the model's reply, evaluates it if it looks like one, and hands back either
a computed answer or a plain refusal:

`examples/code_execution/run.py` (lines 112-143)

```python
    first = model.complete(messages, max_tokens=60)
    expr = first.text.strip()

    if not expr or expr.upper() == DECLINE:
        tracer.record(
            kind="model",
            decided_by="model",
            title="Model declines: not enough numbers in the passages",
            detail=expr or "(empty)",
            tokens_in=first.tokens_in,
            tokens_out=first.tokens_out,
            ms=first.ms,
        )
        return Answer(text="The documents don't give enough numbers to compute that.", citations=[])

    tracer.record(
        kind="model",
        decided_by="model",
        title="Model writes an expression",
        detail=expr,
        tokens_in=first.tokens_in,
        tokens_out=first.tokens_out,
        ms=first.ms,
    )

    try:
        value = safe_eval(expr)
    except (UnsafeExpression, ArithmeticError) as exc:
        tracer.record(kind="code", decided_by="code", title="Sandbox refused the expression", detail=str(exc))
        return Answer(text=f"Could not safely evaluate that expression: {exc}", citations=[])

    tracer.record(kind="code", decided_by="code", title="Sandbox evaluates the expression", detail=f"{expr} = {value}")
```

Declining (`NONE`) and writing a real expression are both the one `decided_by: "model"` step this
level records. This is the same rule as function calling's tool-or-not choice. A refusal from the
sandbox is `decided_by: "code"`, same as running a tool: the model already made its one decision
by the time the sandbox looks at what it wrote, so rejecting the content is code's call, not a
second model decision. Run it yourself:

`examples/code_execution/README.md` (lines 17-17)

```text
python -m examples.code_execution --model stub:scripted
```

## When you do not need this

Try [function calling](/gradient_ascent/techniques/function-calling/) first if the set of things
you would want the model to do can be named and typed in advance as a short list of tools: most
of the time it can, and a named tool is easier to log, test and limit than an open-ended
expression. Try a fixed formula in your own code if the calculation itself never varies with the
question.

Move up to code execution once the computation depends on numbers or an operation you cannot
enumerate in advance as a fixed set of named actions: arbitrary arithmetic, reshaping a table, a
chart built from data given at question time.

## Failure modes

### The expression is safe but uses the wrong numbers

- **How to notice it:** The sandbox happily evaluates 38.50 + 46.00 (a real result, cited to a real section) but the two numbers came from different products than the question asked about, because nothing checks that an expression only uses numbers the retrieved passages actually named for that product.
- **How to test for it:** Ask about two similarly priced parts from different models and check that the cited section actually contains both numbers used, not just numbers that happen to appear somewhere in the retrieved passages.

### The needed numbers were never retrieved

- **How to notice it:** The model declines, correctly, because the passages it was given do not contain a number it needs, but a different search would have found it. A safe decline still means a right answer nobody got.
- **How to test for it:** Ask a numeric question whose figures live in a section a keyword search ranks below the cutoff, and check whether the run declines instead of computing a wrong number from partial information.

### The whitelist is too strict for a legitimate question

- **How to notice it:** A question that genuinely needs an operation outside plus, minus, times and divide (a percentage, a power, a square root) gets a decline that looks like a security block but is really a missing feature. The length and depth bounds do the same for a legitimate sum with too many terms in it.
- **How to test for it:** Ask a question whose arithmetic needs a percentage or an exponent and confirm the run declines cleanly, rather than the model trying to fake the operation with what is allowed. Read the refusal text: it names which bound was hit, so a missing operator and an over-long expression do not look alike in a log.

### A real sandbox is given more reach than the question needs

- **How to notice it:** This example's evaluator cannot make a network call or read a file no matter what the model writes, because those node types are not on the whitelist at all. A real code-execution sandbox that runs actual Python can, unless its own network and filesystem limits are configured as tightly as the question needs.
- **How to test for it:** For a real sandbox, check its documented network and file-access limits directly rather than assuming a model's own caution will substitute for them.

### An expression built to escape the whitelist

- **How to notice it:** A written expression tries to reach a name, a call or an attribute (the pattern behind most real sandbox escapes) and has to be refused the same way a harmless typo is, before anything runs.
- **How to test for it:** Script a model response that writes __import__('os').system(...) and confirm the run refuses it and never calls Python's own eval or exec; see tests/test_example_code_execution.py.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, valid expression:** 2
- **Model calls, declines:** 1
- **Tokens in, expression turn:** ~230
- **Tokens out, expression turn:** ~8

**Compared with Function calling (level 4).** Function calling picks from a short, named list of actions. Code execution picks the content of an expression instead, so the same one-decision shape can answer a much wider range of questions without a new tool being defined for each one.

## How to Evaluate It

_Scored on 12 questions across kinds: numeric._

Code execution is registered with the site's runner as not scored against the shared 60-question
set, with the reason (see `docs/EVALS.md`). Asking `scripts/eval_run.py` for it by name prints
that reason and stops, rather than producing a number about a task the technique was never built
to do. What it would be scored on instead: correctness on the 12 `numeric` questions in the
shared set, the only kind this technique answers by design, plus two numbers the shared harness
does not otherwise track: how often a deliberately malicious or malformed expression is refused
rather than evaluated, and how cleanly the other four question kinds (lookup, multi-hop,
unanswerable, conflicting sources) are declined instead of forced through an arithmetic answer
that was never going to fit them.

## Run it

**What to monitor.** The refusal rate on the sandbox step, tracked separately from the decline rate on the model step. A refusal means the model wrote something outside the whitelist; a decline means it correctly said the passages did not have enough numbers. Conflating the two hides whether a rising number is a prompting problem or a retrieval problem.

**Cost at volume.** Every question pays for at least one call. A second call, and the sandbox's own negligible cost, only happen when the model actually wrote something worth running, so cost tracks how often real traffic needs a computation, not a fixed number per question.

**How it fails in production.** A question whose real answer needs an operation off the whitelist (a percentage, a root) gets declined and reads as the documents lacking an answer, when the documents had everything but the sandbox lacked the operator. A retrieved passage has the wrong document's numbers in it, and the expression computes a real, wrong number with a real-looking citation.

**What to log.** The retrieved passages and their citations, the raw text the model wrote, whether it parsed as a safe expression or was refused and why, the computed value, and the final answer, so a wrong number traces back to retrieval, to the expression, or to the sandbox's own arithmetic.

## Try it

1. **Use it.** Ask a chat app's data-analysis feature a question that needs a calculation across numbers you give it, then ask it to show you the code it ran. Does the code actually compute what the final answer claims?
2. **Build it.** Run python -m examples.code_execution --model stub:scripted from the repo root. Retrieval finds the two prices, the model writes 38.50 + 41.00, the sandbox evaluates it, and the answer is $79.50 with its citation. Run it again with --model stub: the echoed question goes to the sandbox instead of an expression and is refused as invalid syntax, which is the point, since nothing the model returns is trusted to be arithmetic. Then open a Python shell in the repo root and call examples.code_execution.run.safe_eval on a few strings of your own: "38.50 + 41.00", "2 ** 10", "__import__('os').listdir('.')", and "1" plus "+1" forty times. Read which bound each one hits.
3. **Either lane.** Pick one of the failure modes above and try to cause it on purpose, using only the synthetic documents in evals/corpus/.
4. **Build it.** Read examples/bench_test_data_by_conversation/README.md: on this site's electronics-test bench, a model writes a short snippet against a retest export whose volts column is actually millivolts under a header that says volts, and a sandbox runs it. Before that snippet or the model ever sees the table, code range-checks the column against the widest node the board has anywhere and corrects the mislabeled unit. Why does that check have to happen in code, on every table, whether or not the model's tool gets called at all, rather than being one more thing the written snippet is trusted to do?


## Sources

1. [Code execution tool](https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool) — Anthropic (accessed 2026-09-19)
2. [Code Interpreter](https://developers.openai.com/api/docs/guides/tools-code-interpreter) — OpenAI (API documentation) (accessed 2026-09-19)
3. [E2B](https://e2b.dev) — E2B (accessed 2026-09-19)


Last reviewed 2026-09-19.
