A coding agent, decoded
One kind of product, taken apart into the techniques it is built from. Every claim about a product here is what its maker documents, quoted and linked.
Reaches level 6
Decoded from Claude Code (Anthropic), Codex (OpenAI), Cursor (Anysphere). Names and makers as registered on 09/19/2026; the names index carries each entry's own source.
What you see
Type a request into a terminal or an editor: fix the bug where checkout charges twice, or add a dark-mode toggle. A few seconds later the tool has already read the files that matter, proposed or made a change, and run the project’s tests. If a test fails, it reads the failure and tries again on its own. When it stops, it tells you what it did and shows a diff to review. Claude Code, Codex and Cursor all work this way.
That loop is the whole difference from a chatbot that writes code. A chatbot hands you a code block; you paste it in, run it yourself, and copy any error back into the chat for another try. A coding agent runs the command itself, reads what came back, and decides whether that counts as done. The write-run-read-fix loop that used to live in your head now runs inside the tool, and the tool decides when to stop.
What is happening underneath
Strip away the editor and the terminal, and a coding agent is a single agent whose tools happen to be read, search, edit and run, instead of a research agent’s. Anthropic describes the mechanism plainly, and says the same loop powers Claude Code: “Claude evaluates your prompt, calls tools to take action, receives the results, and repeats until the task is complete”[2]. The loop ends when a turn comes back with no tool call in it. Code fits it unusually well, because a test is a check a computer can run rather than a judgment call, and coding agents is the page on why that matters.
Getting to the right files is its own mechanism, and the two makers who document theirs document different ones. Claude Code’s page says it “maps and explains entire codebases in a few seconds” using “agentic search to understand project structure and dependencies without you having to manually select context files”[1]: it searches its way in, the way a person would. Cursor describes a search engine instead, “Instant Grep, a custom search engine that outperforms ripgrep on large codebases”[9], and says where it runs: “Instant Grep builds and queries its index on your machine.”[9] Cursor is also specific about what that does not cover: “When Agent opens a match, that file content can still be included in the model request.”[9] Search decides what the agent puts in the window, not what the window ends up holding.
Beside the loop sits a place for standing instructions, and Anthropic’s documentation separates two
kinds by when they load. A project file loads once and stays: the SDK documentation lists
CLAUDE.md files as loading at “Session start” with their “Full content in every request (but
prompt-cached, so only the first request pays full cost)”[2]. A skill does not: “Unlike
CLAUDE.md content, a skill’s body loads only when it’s used, so long reference material costs
almost nothing until you need it.”[5] That is the
skills mechanism, a procedure kept on the shelf rather than
carried into every request.
A task too large for one context can spawn smaller agents instead. Anthropic’s documentation puts it plainly: “Each subagent runs in its own context window with a custom system prompt, specific tool access, and independent permissions.”[3] Cursor’s Explore subagent does a version of the same job: it “runs in its own context window with a faster model” and “executes many parallel searches without bloating the main conversation”[9]. Either way this is the shape lead agent and workers names: one model’s output decides what another model is asked to do, while your code still bounds every worker.
What a coding agent may do without asking is a setting, not a personality trait. OpenAI’s documentation names the split for Codex: a sandbox “is the boundary that lets the agent act autonomously without giving it unrestricted access to your machine”[7], and that is separate from permission: “The sandbox defines technical boundaries. The approval policy decides when the agent must stop and ask before crossing them.”[7] OpenAI also states Codex’s starting position: “By default, the agent runs with network access turned off.”[8] Reviewing what ran inside that boundary is safety, privacy and governance’s territory.
Long sessions outgrow the window. Anthropic’s SDK documentation says that when the context window “approaches its limit, the SDK automatically compacts the conversation: it summarizes older history to free space, keeping your most recent exchanges and key decisions intact”[2]. What survives a compaction is a harness decision, covered on the agent harness, not a fact about the model answering that turn.
Which page explains each part
| What you see | What it is | Page |
|---|---|---|
| Finds and edits the right files without being told which | A single agent’s tool loop | Single agent |
| Writes code, then runs the tests itself | A loop with a check a computer can run | Coding agents |
| Follows standing instructions, loads a procedure only when needed | A project file every turn, a skill on demand | Skills |
| Hands part of a task to another instance, keeps only the summary | A lead model delegating to workers | Lead agent and workers |
| Asks before touching the network or a file outside the project | Sandboxing and an approval policy | Safety, privacy and governance |
| Keeps working through a long session without losing the task | Context compaction inside the harness | The agent harness |
What the makers say
Three makers, in their own words, on what sits around the model. Anthropic describes the safety net around a Claude Code session as a habit built into every turn rather than a manual save: checkpointing “automatically captures the state of your code before each prompt you send that starts a turn”[6]. OpenAI describes the sandbox as a trust boundary rather than a judgment about the agent’s intentions: “You aren’t just trusting the agent’s intentions; you are trusting that the agent is operating inside enforced limits.”[7] Cursor describes its agent as components it tunes per model: “Cursor’s agent orchestrates these components for each model we support, tuning instructions and tools specifically for every frontier model”[10]. All three describe a system built around the model, which is the reason this kind of product is worth taking apart.
Where it fails
A fix that passes the test it was pointed at can still break something else. Coding agents’ own failure modes name that one first, and the check they give is running the whole suite after the agent reports success.
A loaded skill can be the failure itself. Anthropic warns that “a malicious Skill can direct Claude to invoke tools or execute code in ways that don’t match the Skill’s stated purpose”[4], so a skill checked into a shared repo needs the same review as the code around it.
An undo is narrower than it looks. Claude Code’s documentation states two limits: “Checkpointing does not track files modified by Bash commands”[6], and “A subagent makes edits with Claude’s file editing tools, but Claude Code usually doesn’t capture those edits in your session’s checkpoints”[6]. A rewind can miss the change a shell command or a delegated worker made.
A sandbox is only as tight as the mode it runs in. Codex’s most permissive setting “removes the filesystem and network boundaries and should be used only when you want the agent to act with full access”[7], worth reading before switching to it for convenience.
If you build one
Start at level 5, not higher. A single agent loop over a small, well-chosen tool set is most of the value here, and coding agents’ build lane shows a working version. Add a project instruction file next: it is the cheapest thing on this list, and Anthropic’s own documentation treats it as the default home for standing facts, with skills for anything that has grown into a procedure[5].
Reach for skills only once there is more than one standing procedure, and for lead agent and workers only once a task genuinely does not fit inside one context window. Both cost real tokens and real complexity.
Decide the permission model and the sandbox boundary before the first real run rather than after a bad one. That is safety, privacy and governance’s territory, and OpenAI gives the reason to settle it early: a boundary you have already approved is what lets the agent “read files, make edits, and run routine project commands” without confirming each one[7]. The closest recipe here is coding assistant on your own repo, which assembles those same three pieces.
6 techniques explain this product
The highest level it reaches is level 6.
The agent harness
SourcedEverything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.
Primary sources
- Claude Code by Anthropic | AI Coding Agent, Terminal, IDE · Anthropic (accessed 09/18/2026)
- How the agent loop works · Anthropic (Claude Agent SDK documentation) (accessed 09/18/2026)
- Create custom subagents · Anthropic (Claude Code documentation) (accessed 09/18/2026)
- Agent Skills · Anthropic (accessed 09/18/2026)
- Extend Claude with skills · Anthropic (Claude Code documentation) (accessed 09/18/2026)
- Checkpointing · Anthropic (Claude Code documentation) (accessed 09/18/2026)
- Sandbox · OpenAI (Codex documentation) (accessed 09/18/2026)
- Agent approvals & security · OpenAI (Codex documentation) (accessed 09/18/2026)
- Search · Cursor (Anysphere) (accessed 09/18/2026)
- Agent overview · Cursor (Anysphere) (accessed 09/18/2026)
Last reviewed 09/18/2026. Products change faster than techniques do, so this teardown expires on 03/17/2027 and is re-reviewed or retired then. Markdown version of this page