Primary sources
- Code execution tool · Anthropic (accessed 09/19/2026)
- Code Interpreter · OpenAI (API documentation) (accessed 09/19/2026)
- E2B · E2B (accessed 09/19/2026)
Letting the model write code and run it in a sandbox.
Sourced
Concept at a glance
The execution environment needs its own limits; generated code is not automatically safe.
Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.
Running generated code in a constrained environment to perform computation or transform data.
Follow a computation from supplied data through a small program into a checkable result. Inspect units, assumptions, and evidence of execution separately from generated code.
Compute total energy from these readings and show the units.
Input preview, inspectable Python, deterministic totals, a planted unit error, and denied file/network access in a mock boundary demonstration.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Compute total energy from these readings and show the units.
Authored case. Select any record below; nothing is sent to a model.What changed: Establish the facts supplied for this version of the task.
Inputs and units must be defined. Code that looks plausible may not have run, and running successfully does not prove the calculation is appropriate.
Describe your task to your own model and use Code execution as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Code execution lets the model write a small program instead of choosing among named tools, and a sandbox (not the model) runs it. The shape is the same as function calling: the model’s output picks what happens next, and your code always carries it out. Anthropic describes its own version this way: the tool “allows Claude to run Bash commands and manipulate files, including writing code, in a secure, sandboxed environment”[1]. What the model writes is data your code hands to an interpreter, never text your code trusts and runs directly.
Code execution sits at level 4, tools. The model’s one real choice is what code to write; the
decided_by: "model" step below is picking that content, not deciding whether to run it: your
code always runs whatever it wrote, inside a fixed sandbox, and always asks once more for an
answer once the result is back. That fixed shape, one write-and-run cycle bounded from outside, is
the line to level 5: a coding agent keeps writing and running code in a loop it exits on its own,
reading each result before deciding what to write next.
This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
The model writes one arithmetic expression; a restricted sandbox always runs it.
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
"What is the total price to replace the heating elements on both a DW-300 and a DW-480?"
This is the feature in ChatGPT’s data analysis, Gemini Notebook and Mistral Vibe that runs actual Python on data you give it, instead of guessing a number from the words in your message. Paste a table or upload a spreadsheet, then ask something concrete: “Add up the total in the amount column, and show me a bar chart of totals by month.” A model asked to just compute that from the text of your message can get arithmetic wrong; a model that writes and runs code cannot, because the number comes from the code executing, not from a token prediction.
What makes this safe to try at all is what the sandbox refuses to do, not what it lets the model write. Anthropic’s own container has “Internet access: Completely disabled for security”[1]; OpenAI’s runs the same way, inside “a fully sandboxed virtual machine that the model can run Python code in”[2], with a fixed memory limit and a session that expires “if it is not used for 20 minutes”[2]. Neither can reach your email, your bank, or anything else on the internet, whatever the code says.
Get in the habit of asking to see the code, not just the number or the chart: most of these products have a button or an expandable section for it. You do not need to read Python fluently to check the shape of it: does it use the column you actually asked about, and does the final number come from a calculation in the code rather than a sentence typed after it? A total that changes when you ask the same question twice, or a chart with no code shown next to it, is a sign the answer was written rather than computed.
The same habit answers the question of trust for a credential, too: if a connected tool can reach a service that needs a password, that key has to live somewhere the code cannot read and only the network call can use, the way a sandbox product outside the chat apps, E2B, describes its own “Secrets vault” as “Keys your agent can use but never read”[3].
If the file is small enough to eyeball, or the calculation is one you would trust a spreadsheet formula to do, that is faster than typing a prompt for it: open the spreadsheet.
The example answers a numeric question by writing one arithmetic expression instead of prose. The
model never gets Python; it gets a system prompt that asks for exactly one line (an expression
using numbers, + - * / and parentheses, or the word NONE if the retrieved passages don’t have
what it needs) and the code is the only thing that ever runs it:
SYSTEM_PROMPT = (
"You answer numeric questions about Halvorsen appliances using only the numbered sources "
"below. Reply with exactly one line: a Python arithmetic expression using only numbers, "
"+ - * / and parentheses, that computes the answer -- no words, no units, no code fences. "
f"If the sources do not contain the numbers you would need, reply with the single word "
f"{DECLINE} instead."
)This is an allow-list, not a blocklist. A blocklist has to name every dangerous spelling in
advance; _eval_node instead names the handful of node types arithmetic actually needs
(constants, the four operators, unary plus and minus) and falls through to UnsafeExpression for
anything else. A function call, a name lookup, an attribute access and an import all fail the
same way, because none of them is a node type this function ever matches. That is why the
whitelist is the point: it does not have to recognize an attack to stop it.
The whitelist itself is one function, and it is the whole defense, so here it is rather than a description of it. Every node type the task needs is named; anything else raises before it can run, which is why the check does not have to recognize an attack in order to stop one:
def _eval_node(node: ast.AST, depth: int = 0) -> float:
if depth > MAX_DEPTH:
raise UnsafeExpression(f"expression nests deeper than {MAX_DEPTH} levels")
if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)) and not isinstance(node.value, bool):
return node.value
if isinstance(node, ast.BinOp) and type(node.op) in _BINOPS:
return _BINOPS[type(node.op)](_eval_node(node.left, depth + 1), _eval_node(node.right, depth + 1))
if isinstance(node, ast.UnaryOp) and type(node.op) in _UNARYOPS:
return _UNARYOPS[type(node.op)](_eval_node(node.operand, depth + 1))
raise UnsafeExpression(f"{type(node).__name__} is not on the arithmetic whitelist")Three bounds sit around that whitelist, because the node type is not the only way an expression
can be a problem. The text is capped before it is parsed. The walk is capped by depth, so a long
chain (1+1+1+...) is refused rather than running Python out of stack and raising an error this
function never promised. And a result that is not a finite number is refused, so 1e400 and
1e308 * 1e308 come back as refusals rather than as inf:
if len(expr) > MAX_EXPRESSION_CHARS:
raise UnsafeExpression(f"expression is {len(expr)} characters; the limit is {MAX_EXPRESSION_CHARS}")
try:
tree = ast.parse(expr, mode="eval")
except (SyntaxError, ValueError) as exc:
raise UnsafeExpression(f"not a valid expression: {exc}") from exc
value = _eval_node(tree.body)
if isinstance(value, float) and not math.isfinite(value):
raise UnsafeExpression(f"result is not a finite number: {value}")
return valueThat last bound is on the result, not on the arithmetic, and the difference is worth knowing
before you copy it: 1/1e400 overflows in the middle and comes back as 0.0, a finite number,
so it is returned like any other. Nothing here is wrong with the figure (it is the value Python
computes) but a reader is not told that an intermediate went to infinity on the way.
The run itself checks the model’s reply, evaluates it if it looks like one, and hands back either a computed answer or a plain refusal:
first = model.complete(messages, max_tokens=60)
expr = first.text.strip()
if not expr or expr.upper() == DECLINE:
tracer.record(
kind="model",
decided_by="model",
title="Model declines: not enough numbers in the passages",
detail=expr or "(empty)",
tokens_in=first.tokens_in,
tokens_out=first.tokens_out,
ms=first.ms,
)
return Answer(text="The documents don't give enough numbers to compute that.", citations=[])
tracer.record(
kind="model",
decided_by="model",
title="Model writes an expression",
detail=expr,
tokens_in=first.tokens_in,
tokens_out=first.tokens_out,
ms=first.ms,
)
try:
value = safe_eval(expr)
except (UnsafeExpression, ArithmeticError) as exc:
tracer.record(kind="code", decided_by="code", title="Sandbox refused the expression", detail=str(exc))
return Answer(text=f"Could not safely evaluate that expression: {exc}", citations=[])
tracer.record(kind="code", decided_by="code", title="Sandbox evaluates the expression", detail=f"{expr} = {value}")Declining (NONE) and writing a real expression are both the one decided_by: "model" step this
level records. This is the same rule as function calling’s tool-or-not choice. A refusal from the
sandbox is decided_by: "code", same as running a tool: the model already made its one decision
by the time the sandbox looks at what it wrote, so rejecting the content is code’s call, not a
second model decision. Run it yourself:
python -m examples.code_execution --model stub:scriptedTry function calling first if the set of things you would want the model to do can be named and typed in advance as a short list of tools: most of the time it can, and a named tool is easier to log, test and limit than an open-ended expression. Try a fixed formula in your own code if the calculation itself never varies with the question.
Move up to code execution once the computation depends on numbers or an operation you cannot enumerate in advance as a fixed set of named actions: arbitrary arithmetic, reshaping a table, a chart built from data given at question time.
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
Code execution is registered with the site’s runner as not scored against the shared 60-question
set, with the reason (see docs/EVALS.md). Asking scripts/eval_run.py for it by name prints
that reason and stops, rather than producing a number about a task the technique was never built
to do. What it would be scored on instead: correctness on the 12 numeric questions in the
shared set, the only kind this technique answers by design, plus two numbers the shared harness
does not otherwise track: how often a deliberately malicious or malformed expression is refused
rather than evaluated, and how cleanly the other four question kinds (lookup, multi-hop,
unanswerable, conflicting sources) are declined instead of forced through an arithmetic answer
that was never going to fit them.
The refusal rate on the sandbox step, tracked separately from the decline rate on the model step. A refusal means the model wrote something outside the whitelist; a decline means it correctly said the passages did not have enough numbers. Conflating the two hides whether a rising number is a prompting problem or a retrieval problem.
Every question pays for at least one call. A second call, and the sandbox's own negligible cost, only happen when the model actually wrote something worth running, so cost tracks how often real traffic needs a computation, not a fixed number per question.
A question whose real answer needs an operation off the whitelist (a percentage, a root) gets declined and reads as the documents lacking an answer, when the documents had everything but the sandbox lacked the operator. A retrieved passage has the wrong document's numbers in it, and the expression computes a real, wrong number with a real-looking citation.
The retrieved passages and their citations, the raw text the model wrote, whether it parsed as a safe expression or was refused and why, the computed value, and the final answer, so a wrong number traces back to retrieval, to the expression, or to the sandbox's own arithmetic.
Ask a chat app's data-analysis feature a question that needs a calculation across numbers you give it, then ask it to show you the code it ran. Does the code actually compute what the final answer claims?
Run python -m examples.code_execution --model stub:scripted from the repo root. Retrieval finds the two prices, the model writes 38.50 + 41.00, the sandbox evaluates it, and the answer is $79.50 with its citation. Run it again with --model stub: the echoed question goes to the sandbox instead of an expression and is refused as invalid syntax, which is the point, since nothing the model returns is trusted to be arithmetic. Then open a Python shell in the repo root and call examples.code_execution.run.safe_eval on a few strings of your own: "38.50 + 41.00", "2 ** 10", "__import__('os').listdir('.')", and "1" plus "+1" forty times. Read which bound each one hits.
Pick one of the failure modes above and try to cause it on purpose, using only the synthetic documents in evals/corpus/.
Read examples/bench_test_data_by_conversation/README.md: on this site's electronics-test bench, a model writes a short snippet against a retest export whose volts column is actually millivolts under a header that says volts, and a sandbox runs it. Before that snippet or the model ever sees the table, code range-checks the column against the widest node the board has anywhere and corrects the mislabeled unit. Why does that check have to happen in code, on every table, whether or not the model's tool gets called at all, rather than being one more thing the written snippet is trusted to do?
6 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.
Code execution in a chat app
Maker’s documentation Checked 09/18/2026Research notebook
Maker’s documentation Checked 09/18/2026Ai agent for work and coding
Maker’s documentation Checked 09/19/2026Automation service, self-hostable
Maker’s documentation Checked 09/18/2026Code sandbox
Maker’s documentation Checked 09/18/2026Code sandbox and compute
Maker’s documentation Checked 09/18/2026The model writes a small calculation, a constrained environment runs it, and you inspect the result and errors.
Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page