The example is a propose-edit-run-test loop over one small function held in memory as a string,
not a real file: sum_evens sums the odd numbers instead of the even ones. The model calls
propose_edit with a full replacement; your code reads that source, decides whether it may run
at all, and only then runs it against four fixed test cases. The test result, pass or fail with a
reason, goes back to the model as the tool result, and the model decides whether to try again or
stop.
The deciding is check_source, and it is the part worth copying. An earlier version of this
example ran the proposed source through exec with an empty __builtins__ and called that a
sandbox. It is not one: an empty builtins dict does nothing about attribute access, and attribute
access alone reaches every loaded class and, through any of their __globals__, a real
__builtins__: an escape that uses no builtin name, so no list of forbidden names would catch
it. A review of this repository wrote that escape and it worked. What replaced it is a whitelist
of the AST node types the task actually needs, the same discipline
code execution’s arithmetic evaluator uses:
anything outside the list is refused by construction rather than by spelling.
Even that is a check inside the same interpreter, which is not what a real coding agent needs. A
real one restricts a whole filesystem and process: OpenAI documents Codex using “an OS-enforced
sandbox that limits what it can touch (typically to the current workspace), plus an approval
policy that controls when it must stop and ask you before acting”[4]. What carries over
from this example is the shape, not the boundary: the model never runs its own code directly,
your code always does, and the result the model sees is only ever what your code decided to
report back.
Every propose_edit call and the decision to stop are decided_by: "model"; the tests that run
in between are always decided_by: "code", the same split
single agent’s example makes. max_steps (3) and
max_tokens (2000) are the hard caps; hitting either forces a final answer that is decided_by: "code", and the returned citations are empty when the function was never actually fixed, so a
capped-out run cannot look like a real success by accident.
examples/coding_agents/run.py · lines 77–104
def check_source(source: str) -> str:
"""Why this proposed source may not be run, or "" if it may. Checked before `exec`.
An empty `__builtins__` is NOT a sandbox, and treating it as one is the mistake this check
exists to stop. Nothing in `{"__builtins__": {}}` removes attribute access, and attribute
access is all an escape needs: `().__class__.__mro__[-1].__subclasses__()` reaches every
loaded class, any one of their methods carries a `__globals__` holding a real `__builtins__`,
and from there `open` and `__import__` are back. No builtin name is used anywhere in that
chain, so no name-based check would see it coming.
So this uses the same discipline as `examples/code_execution`'s arithmetic evaluator: name
the node types the task actually needs and refuse everything else by construction, rather
than trying to list the dangerous spellings. Fixing `sum_evens` needs arithmetic, comparison,
a loop, a branch and a return; it needs no attribute access, no import and no global
statement, so none of those are on the list and the escape above has nowhere to start.
This is still not a substitute for running a real coding agent's edits in a real sandbox --
a separate process with its own filesystem and no network. It is the weakest check that makes
this example honest about the separation the module docstring claims.
"""
try:
tree = ast.parse(source)
except (SyntaxError, ValueError) as exc:
return f"not parseable Python: {exc}"
for node in ast.walk(tree):
if type(node) not in ALLOWED_NODES:
return f"{type(node).__name__} is not on the whitelist for a proposed edit"
return ""
Everything above is what happens before a single test runs. The loop itself is ordinary: ask,
check, run, report, repeat until the model stops or a cap ends it.
examples/coding_agents/run.py · lines 132–171
def run(
task: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
max_steps: int = MAX_STEPS,
max_tokens: int = MAX_TOKENS,
) -> Answer:
del embedder # this example has no documents to retrieve; the "test" is the only checker
passed, detail = _run_tests(BUGGY_SOURCE)
tracer.record(kind="code", decided_by="code", title="Run the failing test against the starting code", detail=detail)
messages = [
Message(role="system", content=SYSTEM_PROMPT),
Message(role="user", content=f"{task}\n\nCurrent source:\n{BUGGY_SOURCE}\nTest result: {detail}"),
]
tokens_used = 0
for _ in range(max_steps):
completion = model.complete(messages, tools=[PROPOSE_EDIT_TOOL], max_tokens=300)
tokens_used += completion.tokens_in + completion.tokens_out
if not completion.tool_calls:
record_completion(tracer, decided_by="model", title="Model stops and reports the fix", completion=completion)
return Answer(text=completion.text, citations=[FUNC_NAME] if passed else [])
new_source = str(completion.tool_calls[0].arguments.get("new_source", ""))
record_completion(tracer, decided_by="model", title="Model proposes an edit", completion=completion, detail=new_source[:200])
passed, detail = _run_tests(new_source)
tracer.record(kind="code", decided_by="code", title="Run the test against the proposed edit", detail=detail)
messages.append(Message(role="assistant", content=f"[proposed edit]\n{new_source}"))
messages.append(Message(role="user", content=f"Test result: {detail}"))
if tokens_used >= max_tokens:
reason = f"token budget reached: {tokens_used} >= {max_tokens}"
final = force_final(messages, model, tracer, reason=reason, max_tokens=300)
return Answer(text=final.text, citations=[FUNC_NAME] if passed else [])
final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=300)
return Answer(text=final.text, citations=[FUNC_NAME] if passed else [])
The check does not have to be a test suite. examples/bench_instrument_script_from_the_manual
runs the same propose-check-feedback shape against an instrument’s own documented command set
instead of tests: a drafted command that is not in the manual, or spelled the way a different
vendor’s firmware accepts it, comes back as that instrument’s own error rather than a passing
script, which is what actually happens when a script drafted for one vendor’s SCPI dialect is
pointed at another’s. Read it for the check, not for a coding agent: it is level 3, because code
alone decides when a draft is good enough, not the model.
Run it yourself:
examples/coding_agents/README.md · lines 18–18
python -m examples.coding_agents --model stub:scripted