Level 05 · Agent loops

Coding agents

Agents that read, write, run and test code.

Sourced

Concept at a glance

Edit code, run it, and learn from the result.

Feedback loopConceptual illustration
Edit code, run it, and learn from the result.Coding task leads to Agent. Agent leads to Edit + run tests. Edit + run tests leads to Diff + test output. Diff + test output leads to Agent as feedback. Tests and tool output feed the next edit; passing tests still need review.Coding taskGoal and repository contextAgentChoose the next changeEdit + run testsTools work on the codeDiff + test outputReview and iterateEdit code, run it, and learn from the result.Coding task leads to Agent. Agent leads to Edit + run tests. Edit + run tests leads to Diff + test output. Diff + test output leads to Agent as feedback. Tests and tool output feed the next edit; passing tests still need review.Coding taskGoal and repository contextAgentChoose the next changeEdit + run testsTools work on the codeDiff + test outputReview and iterate

Ending or continuingReturn the diff and test results for review when the task is done or blocked.

Read the connections in words
  • Coding task → Agent: Choose the next change.
  • Agent → Edit + run tests: Tools work on the code.
  • Edit + run tests → Diff + test output: Review and iterate.
  • Diff + test output → Agent: feedback informs another turn.
Key idea

Tests and tool output feed the next edit; passing tests still need review.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Coding agents: see it in practice.

Agents that inspect, edit, execute, and test code to complete a software task.

What you’ll walk through

Follow a code change from understanding the repository through a proposed edit and focused verification. Inspect whether the patch fixes the intended behavior without silently changing unrelated contracts.

The task in this version

Fix the inclusive end-date filter without unrelated changes.

What you’ll learn to check

A reproducible failing case, reviewed patch, deterministic tests, edge-case regression, and change summary.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Fix the inclusive end-date filter without unrelated changes.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Event at 18:00 on the selected end date is excluded. Existing test covers midnight only.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Repository conventions and existing tests are context, not proof that the current behavior is correct. The issue needs a concrete expected outcome.

1 / 6

Apply this to your project

Describe your task to your own model and use Coding agents as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

A coding agent is a single agent whose tools read files, edit them, run commands and run tests, instead of searching documents. Code is a domain this pattern fits unusually well, because a test is a checker a computer can run: a change either passes its tests or it does not, so the loop has a real signal to act on instead of the model’s own sense of whether it is finished. AGENTS.md, an open format several coding agents read, puts it directly: list your test commands and “the agent will attempt to execute relevant programmatic checks and fix failures before finishing the task”[1].

Level 5 still means the model decides the action and the stop; your code still runs everything the model proposes and enforces the caps. For a coding agent specifically, the caps that matter most are the ones that limit how much damage a wrong step can do before a person sees it: which files it can touch, whether it can reach the network, and how many turns it gets before it has to stop and report.

This page is sourced, not measured: what these coding agents do comes from their makers’ own documentation, and no agent has been set on a repository and scored here. It is illustrated.

For a worked example using an established engineering framework, see creating a test automation project with an agent harness. It follows an English request through project structure, existing instrument drivers, unit conversion, CSV output, plan approval, non-hardware checks, and user-led hardware testing.

Build tools, then use them

A coding agent can help accomplish a task by creating or adapting a tool and then using it. For example, to prepare a weekly project report, it could reuse existing connectors, write a small collector for a missing source, validate the combined records, and run a report generator. The deliverables are both the report draft and reusable tools, with evidence of what was checked. This is an illustrative design pattern, not a claim that custom code is always the best choice.

The loop is inspect existing capabilities → propose a tool or adaptation → obtain required approval → build and test → execute within permissions → inspect the output → revise or hand off. A successful command is not enough: check source coverage, missing records, calculations, and the result against the actual task. Generated code needs review; shared frameworks and external systems remain subject to the same boundaries as any other action.

Distinguish building the tool from running the finished tool. The agent may choose its own actions during development, while the resulting collector or report generator later runs as ordinary software without a model. Keep an agent in the recurring process only where its judgment or adaptation is useful. Compare development and maintenance effort with using an existing product or a fixed workflow.

This pattern connects code execution, tool calls, the agent harness, and human approval.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Coding agents

The model proposes an edit, your code runs the test, and the model decides whether to try again.

Level 5 · Agent loops
Task + buggy functionTask + buggyfunctionrun the failing testrun thefailing testMODELpicks the next actionpicks thenext actionTOOLpropose_edit(source)propose_edit(source)Fix reportedFix reported
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 07Your code chose

The task arrives with the buggy function attached

"Fix sum_evens so it returns the sum of the even
numbers in a list."
0 tokens · 0 ms

Practical guidance

Claude Code, Codex, Cursor, GitHub Copilot and the other coding agents this page names all put a permission setting in their settings menu, often called something like Auto, Plan, or Approval mode. Start on the most cautious one: Anthropic’s own plan mode is a state where “Claude explores and plans without editing your source files; file edits are never auto-approved,”[2] so you see the whole plan before anything on disk changes. Loosen it one step at a time as you learn what a task actually needs to touch.

Before real work, write down, once, what you would otherwise repeat every session: the build command, the test command, and anything it should never touch. AGENTS.md is the open convention for this file, “a README for agents”, holding “the extra, sometimes detailed context coding agents need: build steps, tests, and conventions”[1]. A brief that names the test command gets an agent that can tell for itself whether it succeeded; a vague one gets a vague result.

A first task with nothing at stake: point it at a small, already-broken piece of your own project and ask it to fix that one thing and run the tests, nothing else. OpenAI documents its own default: under Codex’s Auto preset the agent “can read files, make edits, and run commands in the working directory automatically”, asking approval only to “edit files outside the workspace or to run commands that require network access”[4]. Cross that boundary and, in OpenAI’s words, “the approval flow takes over”[3]: a request you answer, not something that happens quietly.

Read what it changed the way you would read a colleague’s work: does the diff touch only what you asked for, and did the tests actually run, or does the agent’s summary just say they should? Claude Code keeps a checkpoint “before each prompt you send that starts a turn,”[5] so a bad change can be undone, but the net has a hole: it “does not track files modified by Bash commands”[5], only its own edits, and it is not a substitute for real version control.

A run that stops partway through is usually the approval boundary working: it reached something outside the workspace, or the network, and is waiting on you rather than guessing.

If the fix is one obvious line, just make it: briefing an agent for that costs more than typing the line.

Implementation details

The example is a propose-edit-run-test loop over one small function held in memory as a string, not a real file: sum_evens sums the odd numbers instead of the even ones. The model calls propose_edit with a full replacement; your code reads that source, decides whether it may run at all, and only then runs it against four fixed test cases. The test result, pass or fail with a reason, goes back to the model as the tool result, and the model decides whether to try again or stop.

The deciding is check_source, and it is the part worth copying. An earlier version of this example ran the proposed source through exec with an empty __builtins__ and called that a sandbox. It is not one: an empty builtins dict does nothing about attribute access, and attribute access alone reaches every loaded class and, through any of their __globals__, a real __builtins__: an escape that uses no builtin name, so no list of forbidden names would catch it. A review of this repository wrote that escape and it worked. What replaced it is a whitelist of the AST node types the task actually needs, the same discipline code execution’s arithmetic evaluator uses: anything outside the list is refused by construction rather than by spelling.

Even that is a check inside the same interpreter, which is not what a real coding agent needs. A real one restricts a whole filesystem and process: OpenAI documents Codex using “an OS-enforced sandbox that limits what it can touch (typically to the current workspace), plus an approval policy that controls when it must stop and ask you before acting”[4]. What carries over from this example is the shape, not the boundary: the model never runs its own code directly, your code always does, and the result the model sees is only ever what your code decided to report back.

Every propose_edit call and the decision to stop are decided_by: "model"; the tests that run in between are always decided_by: "code", the same split single agent’s example makes. max_steps (3) and max_tokens (2000) are the hard caps; hitting either forces a final answer that is decided_by: "code", and the returned citations are empty when the function was never actually fixed, so a capped-out run cannot look like a real success by accident.

examples/coding_agents/run.py · lines 77–104
def check_source(source: str) -> str:
    """Why this proposed source may not be run, or "" if it may. Checked before `exec`.

    An empty `__builtins__` is NOT a sandbox, and treating it as one is the mistake this check
    exists to stop. Nothing in `{"__builtins__": {}}` removes attribute access, and attribute
    access is all an escape needs: `().__class__.__mro__[-1].__subclasses__()` reaches every
    loaded class, any one of their methods carries a `__globals__` holding a real `__builtins__`,
    and from there `open` and `__import__` are back. No builtin name is used anywhere in that
    chain, so no name-based check would see it coming.

    So this uses the same discipline as `examples/code_execution`'s arithmetic evaluator: name
    the node types the task actually needs and refuse everything else by construction, rather
    than trying to list the dangerous spellings. Fixing `sum_evens` needs arithmetic, comparison,
    a loop, a branch and a return; it needs no attribute access, no import and no global
    statement, so none of those are on the list and the escape above has nowhere to start.

    This is still not a substitute for running a real coding agent's edits in a real sandbox --
    a separate process with its own filesystem and no network. It is the weakest check that makes
    this example honest about the separation the module docstring claims.
    """
    try:
        tree = ast.parse(source)
    except (SyntaxError, ValueError) as exc:
        return f"not parseable Python: {exc}"
    for node in ast.walk(tree):
        if type(node) not in ALLOWED_NODES:
            return f"{type(node).__name__} is not on the whitelist for a proposed edit"
    return ""

Everything above is what happens before a single test runs. The loop itself is ordinary: ask, check, run, report, repeat until the model stops or a cap ends it.

examples/coding_agents/run.py · lines 132–171
def run(
    task: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    del embedder  # this example has no documents to retrieve; the "test" is the only checker
    passed, detail = _run_tests(BUGGY_SOURCE)
    tracer.record(kind="code", decided_by="code", title="Run the failing test against the starting code", detail=detail)
    messages = [
        Message(role="system", content=SYSTEM_PROMPT),
        Message(role="user", content=f"{task}\n\nCurrent source:\n{BUGGY_SOURCE}\nTest result: {detail}"),
    ]

    tokens_used = 0
    for _ in range(max_steps):
        completion = model.complete(messages, tools=[PROPOSE_EDIT_TOOL], max_tokens=300)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            record_completion(tracer, decided_by="model", title="Model stops and reports the fix", completion=completion)
            return Answer(text=completion.text, citations=[FUNC_NAME] if passed else [])

        new_source = str(completion.tool_calls[0].arguments.get("new_source", ""))
        record_completion(tracer, decided_by="model", title="Model proposes an edit", completion=completion, detail=new_source[:200])
        passed, detail = _run_tests(new_source)
        tracer.record(kind="code", decided_by="code", title="Run the test against the proposed edit", detail=detail)
        messages.append(Message(role="assistant", content=f"[proposed edit]\n{new_source}"))
        messages.append(Message(role="user", content=f"Test result: {detail}"))

        if tokens_used >= max_tokens:
            reason = f"token budget reached: {tokens_used} >= {max_tokens}"
            final = force_final(messages, model, tracer, reason=reason, max_tokens=300)
            return Answer(text=final.text, citations=[FUNC_NAME] if passed else [])

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=300)
    return Answer(text=final.text, citations=[FUNC_NAME] if passed else [])

The check does not have to be a test suite. examples/bench_instrument_script_from_the_manual runs the same propose-check-feedback shape against an instrument’s own documented command set instead of tests: a drafted command that is not in the manual, or spelled the way a different vendor’s firmware accepts it, comes back as that instrument’s own error rather than a passing script, which is what actually happens when a script drafted for one vendor’s SCPI dialect is pointed at another’s. Read it for the check, not for a coding agent: it is level 3, because code alone decides when a draft is good enough, not the model.

Run it yourself:

examples/coding_agents/README.md · lines 18–18
python -m examples.coding_agents --model stub:scripted
When you do not need this

Try code execution first if the model only has to write the fix once and your code can just run it and report the result; that costs one model call instead of a loop, and there is no reason yet to expect a second try.

Try a fixed workflow (lint, autoformat, or a single scripted patch) first if the fix is already known and the same every time; a workflow like that is cheaper and never proposes something unexpected.

Try function calling instead if one tool call settles it, such as running a single existing test suite once and reporting the result with no editing involved.

Move up to a coding agent once the number and shape of edits needed cannot be known before the model reads the failing test, and the code has to be read, run and rewritten until it works, with the model deciding when it is done. Fixing it might take one change or several, depending on what the first attempt reveals.

Failure modes

A fix that passes the shown tests but breaks something else

How to notice it
The edit makes the targeted test pass, but a test outside what the agent was told to run now fails, and nothing in the trace says so.
How to test for it
Run the full test suite after the agent reports success, not just the test it was pointed at. AGENTS.md's own model is to run "relevant programmatic checks" before finishing; a check that was never listed is a check that never ran.

Damage outside the intended boundary

How to notice it
An overly permissive approval setting lets the agent edit or run something outside what the task needed, discovered after the fact rather than blocked at the time.
How to test for it
Check which permission or approval mode the session actually ran under, not which one you meant to set, and confirm the boundary it enforced matches the task, not just the tool's default.

Looping without progress

How to notice it
The model proposes edits that address the same symptom in slightly different ways without ever reading why the previous attempt actually failed, until the step cap ends the run.
How to test for it
Read the test-failure detail at each step in order: real progress narrows toward the actual bug; a loop repeats the same wrong theory.

A diff that looks reasonable but was not actually re-checked

How to notice it
A person reviews the code change, it reads as plausible, and it ships without the tests actually being re-run against it.
How to test for it
Confirm the trace shows a test run after the final edit, not just after an earlier one. A plausible diff and a passing test are two different facts.

The cap ships a still-broken fix

How to notice it
The step or token cap is reached before the tests actually passed, and the run ends with an answer that reads like a normal report unless the forced-stop step is checked.
How to test for it
Force a low cap on a task that needs more attempts than the cap allows (this page's own test suite does exactly this) and confirm the answer's citations are empty rather than claiming success.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

2Model calls, best case (one fix, then stop)
4Model calls, worst case (step cap reached)
~260–470Tokens in, one attempt
~0.7sWall time, one attempt
Compared with a fixed autofix workflow (level 3)A one-line fix costs close to what a single scripted patch-and-test step would; a fix that takes several attempts costs that many times over, for a problem a fixed workflow could not have known the shape of in advance.

How to Evaluate It

coding_agents does not do the site’s own question-answering task (there is no document, no question and no citation to grade), so it is not part of the shared 60-question set, the same way embeddings_search and memory are not (see docs/EVALS.md). What would actually be measured is specific to this task: the share of attempts that reach a passing state at all, the number of attempts it takes, whether a fix that passes its own test also passes every other test in the suite it was not shown, and how closely the size of the diff matches the size of the actual bug. scripts/eval_run.py knows coding_agents and refuses to score it, printing that reason; no runner for the measures above exists yet.

Run it

What to monitor

The share of runs that reach a passing state, the average number of attempts per task, and how often a change that passed its own test later broke something outside it.

Cost at volume

Cost per task varies with how many attempts a fix actually takes, the same as any agent loop; budget from the cap, and watch whether harder tasks are quietly consuming a disproportionate share of it.

How it fails in production

A permission or sandbox setting is looser than the task needed, so a change that should have stayed inside one file reaches further than intended before anyone reviews it.

What to log

Every proposed edit, every test result it produced, and which cap (if any) forced the final answer, so a bad merge traces back to the specific attempt that introduced it rather than to 'the agent did it.'

Try it

  1. Use it

    Find a project's AGENTS.md, CLAUDE.md, or similar instructions file, if it has one. Does it name a test command a coding agent could actually run, or only prose a human would read?

  2. Build it

    Run python -m examples.coding_agents --model stub:scripted from the repo root. The first proposed edit counts the even numbers instead of summing them and the test catches it (3, not 12); the second passes all four cases. For the run that never gets there, read test_step_cap_forces_a_stop_when_the_model_never_fixes_it in tests/test_example_coding_agents.py.

  3. Either lane

    Open examples/coding_agents/run.py and change TEST_CASES to add a case the current BUGGY_SOURCE and the fix in the tests both need to handle differently. Does the existing scripted fix in the tests still pass?

How it connects

Before, after and instead of this

Move up when

  • Lead agent and workersThe change is larger than one agent can hold at once and splits into parts that can be worked separately.

Decoded in

Optional: products, tools, and models

13 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 7 more examples
  • Aider Product · Aider

    Open-source coding agent

    Checked 09/18/2026
  • Devin Product · Cognition

    Coding agent

    Checked 09/19/2026
  • Replit Agent Product · Replit

    Coding agent

    Checked 09/18/2026
  • Devin Desktop Product · Cognition

    Coding agent in an editor

    Checked 09/18/2026
  • Grok Build Product · SpaceXAI

    Coding agent

    Checked 09/18/2026
  • Jules Product · Google

    Coding agent

    Checked 09/18/2026
  • Zed Product · Zed Industries

    Ai code editor

    Checked 09/18/2026
In practice

Fix a failing test

An agent reads the relevant code, makes an edit, runs tests, and uses their output to decide what to change next.

Out there

Named products, tools and models

Products13
  • AiderAider · open-source coding agent
  • Antigravity CLIGoogle · coding agent · formerly Gemini CLI
  • Claude CodeAnthropic · coding agent
  • ClineCline · open-source coding agent
  • CodexOpenAI · coding agent
  • CursorAnysphere · coding agent in an editor
  • DevinCognition · coding agent
  • Devin DesktopCognition · coding agent in an editor · formerly Windsurf
  • GitHub CopilotGitHub · coding agent in an editor
  • Grok BuildSpaceXAI · coding agent
  • JulesGoogle · coding agent
  • Replit AgentReplit · coding agent
  • ZedZed Industries · ai code editor

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. AGENTS.md · agents.md (stewarded by the Agentic AI Foundation, Linux Foundation) (accessed 09/19/2026)
  2. How the agent loop works · Anthropic (Claude Agent SDK documentation) (accessed 09/19/2026)
  3. Sandbox · OpenAI (Codex documentation) (accessed 09/19/2026)
  4. Agent approvals & security · OpenAI (Codex documentation) (accessed 09/19/2026)
  5. Checkpointing · Anthropic (Claude Code documentation) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page