# Coding agents

_Level 05 · Agent loops · sourced_

Agents that read, write, run and test code.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a code change from understanding the repository through a proposed edit and focused verification. Inspect whether the patch fixes the intended behavior without silently changing unrelated contracts.

**Assumptions:** Repository conventions and existing tests are context, not proof that the current behavior is correct. The issue needs a concrete expected outcome.

**Design choices:** Use the smallest coherent change, reuse existing abstractions, and choose checks tied to the bug or feature. Refactoring may be justified when it removes the cause, not just because the agent prefers it.

**Request:** Fix the inclusive end-date filter without unrelated changes.

**Starting evidence:** Event at 18:00 on the selected end date is excluded. Existing test covers midnight only.

**Action and control:** Inspect the comparison, propose an exclusive next-day boundary, and add an end-of-day regression; logs here are illustrative.

**Stage records (authored, not executed):**

### Input record

Event at 18:00 on the selected end date is excluded. Existing test covers midnight only.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use the smallest coherent change, reuse existing abstractions, and choose checks tied to the bug or feature. Refactoring may be justified when it removes the cause, not just because the agent prefers it.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Inspect the comparison, propose an exclusive next-day boundary, and add an end-of-day regression; logs here are illustrative.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Patch intent: include the full selected day. Review timezone and daylight-saving behavior. No repository is changed by this example.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A reproducible failing case, reviewed patch, deterministic tests, edge-case regression, and change summary.

If the result falls short:
If a test fails, distinguish a patch defect from an unrelated baseline failure. Preserve the diff and explain remaining uncertainty before expanding scope.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to a feature, bug, or migration. Agree on editable areas and review expectations appropriate to the repository rather than assuming every project has a protected shared framework.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Patch intent: include the full selected day. Review timezone and daylight-saving behavior. No repository is changed by this example.

**Change something — A timezone regression appears:** The original case passing is insufficient. Preserve the failing case and revise the patch.

**Decision:** Does one green regression test prove completion?

**Answer:** No; inspect relevant edge cases and scope.

**Why:** A passing test can miss regressions; keep edits scoped and distinguish simulated test logs from actually executed tests.

**Review criteria:** A reproducible failing case, reviewed patch, deterministic tests, edge-case regression, and change summary.

**Recovery:** If a test fails, distinguish a patch defect from an unrelated baseline failure. Preserve the diff and explain remaining uncertainty before expanding scope.

**Adapt it:** Apply this to a feature, bug, or migration. Agree on editable areas and review expectations appropriate to the repository rather than assuming every project has a protected shared framework.

A coding agent is a [single agent](/gradient_ascent/techniques/single-agent/) whose tools read
files, edit them, run commands and run tests, instead of searching documents. Code is a domain
this pattern fits unusually well, because a test is a checker a computer can run: a change either
passes its tests or it does not, so the loop has a real signal to act on instead of the model's own
sense of whether it is finished. AGENTS.md, an open format several coding agents read, puts it
directly: list your test commands and "the agent will attempt to execute relevant programmatic
checks and fix failures before finishing the task"[1].

Level 5 still means the model decides the action and the stop; your code still runs everything the
model proposes and enforces the caps. For a coding agent specifically, the caps that matter most
are the ones that limit how much damage a wrong step can do before a person sees it: which files
it can touch, whether it can reach the network, and how many turns it gets before it has to stop
and report.

This page is sourced, not measured: what these coding agents do comes from their makers' own
documentation, and no agent has been set on a repository and scored here. It is illustrated.

For a worked example using an established engineering framework, see
[creating a test automation project with an agent harness](/gradient_ascent/techniques/agent-harness/#worked-example-a-test-automation-framework).
It follows an English request through project structure, existing instrument drivers, unit conversion,
CSV output, plan approval, non-hardware checks, and user-led hardware testing.

## Build tools, then use them

A coding agent can help accomplish a task by creating or adapting a tool and then using it.
For example, to prepare a weekly project report, it could reuse existing connectors, write a
small collector for a missing source, validate the combined records, and run a report generator.
The deliverables are both the report draft and reusable tools, with evidence of what was checked.
This is an illustrative design pattern, not a claim that custom code is always the best choice.

The loop is **inspect existing capabilities → propose a tool or adaptation → obtain required
approval → build and test → execute within permissions → inspect the output → revise or hand off**.
A successful command is not enough: check source coverage, missing records, calculations, and the
result against the actual task. Generated code needs review; shared frameworks and external
systems remain subject to the same boundaries as any other action.

Distinguish **building the tool** from **running the finished tool**. The agent may choose its
own actions during development, while the resulting collector or report generator later runs
as ordinary software without a model. Keep an agent in the recurring process only where its
judgment or adaptation is useful. Compare development and maintenance effort with using an
existing product or a fixed workflow.

This pattern connects [code execution](/gradient_ascent/techniques/code-execution/),
[tool calls](/gradient_ascent/techniques/function-calling/),
[the agent harness](/gradient_ascent/techniques/agent-harness/), and
[human approval](/gradient_ascent/techniques/human-in-the-loop/).

_The web page for this technique includes an interactive step-through of Level 5 · Coding agents. The same steps are described in the sections below._

## Practical guidance

Claude Code, Codex, Cursor, GitHub Copilot and the other coding agents this page names all put a
permission setting in their settings menu, often called something like Auto, Plan, or Approval
mode. Start on the most cautious one: Anthropic's own plan mode is a state where "Claude explores
and plans without editing your source files; file edits are never auto-approved,"[2] so
you see the whole plan before anything on disk changes. Loosen it one step at a time as you learn
what a task actually needs to touch.

Before real work, write down, once, what you would otherwise repeat every session: the build
command, the test command, and anything it should never touch. AGENTS.md is the open convention
for this file, "a README for agents", holding "the extra, sometimes detailed context coding agents
need: build steps, tests, and conventions"[1]. A brief that names the test command gets
an agent that can tell for itself whether it succeeded; a vague one gets a vague result.

A first task with nothing at stake: point it at a small, already-broken piece of your own project
and ask it to fix that one thing and run the tests, nothing else. OpenAI documents its own default:
under Codex's Auto preset the agent "can read files, make edits, and run commands in the working
directory automatically", asking approval only to "edit files outside the workspace or to run
commands that require network access"[4]. Cross that boundary and, in OpenAI's words,
"the approval flow takes over"[3]: a request you answer, not something that happens
quietly.

Read what it changed the way you would read a colleague's work: does the diff touch only what you
asked for, and did the tests actually run, or does the agent's summary just say they should?
Claude Code keeps a checkpoint "before each prompt you send that starts a turn,"[5] so a
bad change can be undone, but the net has a hole: it "does not track files modified by Bash
commands"[5], only its own edits, and it is not a substitute for real version control.

A run that stops partway through is usually the approval boundary working: it reached something
outside the workspace, or the network, and is waiting on you rather than guessing.

If the fix is one obvious line, just make it: briefing an agent for that costs more than typing
the line.

## Implementation details

The example is a propose-edit-run-test loop over one small function held in memory as a string,
not a real file: `sum_evens` sums the odd numbers instead of the even ones. The model calls
`propose_edit` with a full replacement; your code reads that source, decides whether it may run
at all, and only then runs it against four fixed test cases. The test result, pass or fail with a
reason, goes back to the model as the tool result, and the model decides whether to try again or
stop.

The deciding is `check_source`, and it is the part worth copying. An earlier version of this
example ran the proposed source through `exec` with an empty `__builtins__` and called that a
sandbox. It is not one: an empty builtins dict does nothing about attribute access, and attribute
access alone reaches every loaded class and, through any of their `__globals__`, a real
`__builtins__`: an escape that uses no builtin name, so no list of forbidden names would catch
it. A review of this repository wrote that escape and it worked. What replaced it is a whitelist
of the AST node types the task actually needs, the same discipline
[code execution](/gradient_ascent/techniques/code-execution/)'s arithmetic evaluator uses:
anything outside the list is refused by construction rather than by spelling.

Even that is a check inside the same interpreter, which is not what a real coding agent needs. A
real one restricts a whole filesystem and process: OpenAI documents Codex using "an OS-enforced
sandbox that limits what it can touch (typically to the current workspace), plus an approval
policy that controls when it must stop and ask you before acting"[4]. What carries over
from this example is the shape, not the boundary: the model never runs its own code directly,
your code always does, and the result the model sees is only ever what your code decided to
report back.

Every `propose_edit` call and the decision to stop are `decided_by: "model"`; the tests that run
in between are always `decided_by: "code"`, the same split
[single agent](/gradient_ascent/techniques/single-agent/)'s example makes. `max_steps` (3) and
`max_tokens` (2000) are the hard caps; hitting either forces a final answer that is `decided_by:
"code"`, and the returned citations are empty when the function was never actually fixed, so a
capped-out run cannot look like a real success by accident.

`examples/coding_agents/run.py` (lines 77-104)

```python
def check_source(source: str) -> str:
    """Why this proposed source may not be run, or "" if it may. Checked before `exec`.

    An empty `__builtins__` is NOT a sandbox, and treating it as one is the mistake this check
    exists to stop. Nothing in `{"__builtins__": {}}` removes attribute access, and attribute
    access is all an escape needs: `().__class__.__mro__[-1].__subclasses__()` reaches every
    loaded class, any one of their methods carries a `__globals__` holding a real `__builtins__`,
    and from there `open` and `__import__` are back. No builtin name is used anywhere in that
    chain, so no name-based check would see it coming.

    So this uses the same discipline as `examples/code_execution`'s arithmetic evaluator: name
    the node types the task actually needs and refuse everything else by construction, rather
    than trying to list the dangerous spellings. Fixing `sum_evens` needs arithmetic, comparison,
    a loop, a branch and a return; it needs no attribute access, no import and no global
    statement, so none of those are on the list and the escape above has nowhere to start.

    This is still not a substitute for running a real coding agent's edits in a real sandbox --
    a separate process with its own filesystem and no network. It is the weakest check that makes
    this example honest about the separation the module docstring claims.
    """
    try:
        tree = ast.parse(source)
    except (SyntaxError, ValueError) as exc:
        return f"not parseable Python: {exc}"
    for node in ast.walk(tree):
        if type(node) not in ALLOWED_NODES:
            return f"{type(node).__name__} is not on the whitelist for a proposed edit"
    return ""
```

Everything above is what happens before a single test runs. The loop itself is ordinary: ask,
check, run, report, repeat until the model stops or a cap ends it.

`examples/coding_agents/run.py` (lines 132-171)

```python
def run(
    task: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    del embedder  # this example has no documents to retrieve; the "test" is the only checker
    passed, detail = _run_tests(BUGGY_SOURCE)
    tracer.record(kind="code", decided_by="code", title="Run the failing test against the starting code", detail=detail)
    messages = [
        Message(role="system", content=SYSTEM_PROMPT),
        Message(role="user", content=f"{task}\n\nCurrent source:\n{BUGGY_SOURCE}\nTest result: {detail}"),
    ]

    tokens_used = 0
    for _ in range(max_steps):
        completion = model.complete(messages, tools=[PROPOSE_EDIT_TOOL], max_tokens=300)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            record_completion(tracer, decided_by="model", title="Model stops and reports the fix", completion=completion)
            return Answer(text=completion.text, citations=[FUNC_NAME] if passed else [])

        new_source = str(completion.tool_calls[0].arguments.get("new_source", ""))
        record_completion(tracer, decided_by="model", title="Model proposes an edit", completion=completion, detail=new_source[:200])
        passed, detail = _run_tests(new_source)
        tracer.record(kind="code", decided_by="code", title="Run the test against the proposed edit", detail=detail)
        messages.append(Message(role="assistant", content=f"[proposed edit]\n{new_source}"))
        messages.append(Message(role="user", content=f"Test result: {detail}"))

        if tokens_used >= max_tokens:
            reason = f"token budget reached: {tokens_used} >= {max_tokens}"
            final = force_final(messages, model, tracer, reason=reason, max_tokens=300)
            return Answer(text=final.text, citations=[FUNC_NAME] if passed else [])

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=300)
    return Answer(text=final.text, citations=[FUNC_NAME] if passed else [])
```

The check does not have to be a test suite. `examples/bench_instrument_script_from_the_manual`
runs the same propose-check-feedback shape against an instrument's own documented command set
instead of tests: a drafted command that is not in the manual, or spelled the way a different
vendor's firmware accepts it, comes back as that instrument's own error rather than a passing
script, which is what actually happens when a script drafted for one vendor's SCPI dialect is
pointed at another's. Read it for the check, not for a coding agent: it is level 3, because code
alone decides when a draft is good enough, not the model.

Run it yourself:

`examples/coding_agents/README.md` (lines 18-18)

```text
python -m examples.coding_agents --model stub:scripted
```

## When you do not need this

Try [code execution](/gradient_ascent/techniques/code-execution/) first if the model only has to
write the fix once and your code can just run it and report the result; that costs one model call
instead of a loop, and there is no reason yet to expect a second try.

Try a fixed [workflow](/gradient_ascent/techniques/workflow-graphs/) (lint, autoformat, or a
single scripted patch) first if the fix is already known and the same every time; a workflow like
that is cheaper and never proposes something unexpected.

Try [function calling](/gradient_ascent/techniques/function-calling/) instead if one tool call
settles it, such as running a single existing test suite once and reporting the result with no
editing involved.

Move up to a coding agent once the number and shape of edits needed cannot be known before the
model reads the failing test, and the code has to be read, run and rewritten until it works, with
the model deciding when it is done. Fixing it might take one change or several, depending on what
the first attempt reveals.

## Failure modes

### A fix that passes the shown tests but breaks something else

- **How to notice it:** The edit makes the targeted test pass, but a test outside what the agent was told to run now fails, and nothing in the trace says so.
- **How to test for it:** Run the full test suite after the agent reports success, not just the test it was pointed at. AGENTS.md's own model is to run "relevant programmatic checks" before finishing; a check that was never listed is a check that never ran.

### Damage outside the intended boundary

- **How to notice it:** An overly permissive approval setting lets the agent edit or run something outside what the task needed, discovered after the fact rather than blocked at the time.
- **How to test for it:** Check which permission or approval mode the session actually ran under, not which one you meant to set, and confirm the boundary it enforced matches the task, not just the tool's default.

### Looping without progress

- **How to notice it:** The model proposes edits that address the same symptom in slightly different ways without ever reading why the previous attempt actually failed, until the step cap ends the run.
- **How to test for it:** Read the test-failure detail at each step in order: real progress narrows toward the actual bug; a loop repeats the same wrong theory.

### A diff that looks reasonable but was not actually re-checked

- **How to notice it:** A person reviews the code change, it reads as plausible, and it ships without the tests actually being re-run against it.
- **How to test for it:** Confirm the trace shows a test run after the final edit, not just after an earlier one. A plausible diff and a passing test are two different facts.

### The cap ships a still-broken fix

- **How to notice it:** The step or token cap is reached before the tests actually passed, and the run ends with an answer that reads like a normal report unless the forced-stop step is checked.
- **How to test for it:** Force a low cap on a task that needs more attempts than the cap allows (this page's own test suite does exactly this) and confirm the answer's citations are empty rather than claiming success.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, best case (one fix, then stop):** 2
- **Model calls, worst case (step cap reached):** 4
- **Tokens in, one attempt:** ~260–470
- **Wall time, one attempt:** ~0.7s

**Compared with a fixed autofix workflow (level 3).** A one-line fix costs close to what a single scripted patch-and-test step would; a fix that takes several attempts costs that many times over, for a problem a fixed workflow could not have known the shape of in advance.

## How to Evaluate It

`coding_agents` does not do the site's own question-answering task (there is no document, no
question and no citation to grade), so it is not part of the shared 60-question set, the same way
`embeddings_search` and `memory` are not (see `docs/EVALS.md`). What would actually be measured is
specific to this task: the share of attempts that reach a passing state at all, the number of
attempts it takes, whether a fix that passes its own test also passes every other test in the
suite it was not shown, and how closely the size of the diff matches the size of the actual bug.
`scripts/eval_run.py` knows `coding_agents` and refuses to score it, printing that reason; no
runner for the measures above exists yet.

## Run it

**What to monitor.** The share of runs that reach a passing state, the average number of attempts per task, and how often a change that passed its own test later broke something outside it.

**Cost at volume.** Cost per task varies with how many attempts a fix actually takes, the same as any agent loop; budget from the cap, and watch whether harder tasks are quietly consuming a disproportionate share of it.

**How it fails in production.** A permission or sandbox setting is looser than the task needed, so a change that should have stayed inside one file reaches further than intended before anyone reviews it.

**What to log.** Every proposed edit, every test result it produced, and which cap (if any) forced the final answer, so a bad merge traces back to the specific attempt that introduced it rather than to 'the agent did it.'

## Try it

1. **Use it.** Find a project's AGENTS.md, CLAUDE.md, or similar instructions file, if it has one. Does it name a test command a coding agent could actually run, or only prose a human would read?
2. **Build it.** Run python -m examples.coding_agents --model stub:scripted from the repo root. The first proposed edit counts the even numbers instead of summing them and the test catches it (3, not 12); the second passes all four cases. For the run that never gets there, read test_step_cap_forces_a_stop_when_the_model_never_fixes_it in tests/test_example_coding_agents.py.
3. **Either lane.** Open examples/coding_agents/run.py and change TEST_CASES to add a case the current BUGGY_SOURCE and the fix in the tests both need to handle differently. Does the existing scripted fix in the tests still pass?


## Sources

1. [AGENTS.md](https://agents.md/) — agents.md (stewarded by the Agentic AI Foundation, Linux Foundation) (accessed 2026-09-19)
2. [How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop) — Anthropic (Claude Agent SDK documentation) (accessed 2026-09-19)
3. [Sandbox](https://learn.chatgpt.com/docs/sandboxing) — OpenAI (Codex documentation) (accessed 2026-09-19)
4. [Agent approvals & security](https://learn.chatgpt.com/docs/agent-approvals-security) — OpenAI (Codex documentation) (accessed 2026-09-19)
5. [Checkpointing](https://code.claude.com/docs/en/checkpointing) — Anthropic (Claude Code documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
