# Level 05 · Agent loops

_Observe, decide, act, repeat_

The model uses the goal and observed results to choose an action, revise its approach, or finish. Software executes tools and enforces permissions, approvals, and stopping limits. A run can stop because it is complete, blocked, or out of budget.


## Who decides the next step

The model chooses the next step within limits enforced by software.


## What is at this level

- [Single agent](/gradient_ascent/techniques/single-agent/) (sourced): A model that plans, acts and checks its own work in a loop.
- [The agent harness](/gradient_ascent/techniques/agent-harness/) (sourced): Everything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.
- [Agentic RAG and deep research](/gradient_ascent/techniques/agentic-rag/) (measured): An agent that runs its own searches until it has an answer.
- [Coding agents](/gradient_ascent/techniques/coding-agents/) (sourced): Agents that read, write, run and test code.
- [Skills](/gradient_ascent/techniques/skills/) (sourced): Reusable instructions that an agent loads when it needs them.
- [Voice agents](/gradient_ascent/techniques/voice-agents/) (sourced): Agents you talk to in real time.

## Upgrade conditions

- **Single agent → Long-running tasks:** The work outlives one context window or one sitting.
- **Agentic RAG and deep research → Lead agent and workers:** The searches do not depend on each other and there are more of them than one agent's context can carry.
- **Coding agents → Lead agent and workers:** The change is larger than one agent can hold at once and splits into parts that can be worked separately.

## Named products, tools and models


### Products

- Aider — Aider · open-source coding agent
- Antigravity CLI — Google · coding agent
- ChatGPT deep research — OpenAI · research agent
- ChatGPT voice — OpenAI · voice assistant
- Claude Code — Anthropic · coding agent
- Claude Research — Anthropic · research agent
- Cline — Cline · open-source coding agent
- Codex — OpenAI · coding agent
- Cursor — Anysphere · coding agent in an editor
- Devin — Cognition · coding agent
- Devin Desktop — Cognition · coding agent in an editor
- ElevenLabs — ElevenLabs · voice generation
- Gemini Deep Research — Google · research agent
- Gemini Live — Google · voice assistant
- GitHub Copilot — GitHub · coding agent in an editor
- Grok Build — SpaceXAI · coding agent
- Grok DeepSearch — SpaceXAI · research agent
- Jules — Google · coding agent
- Perplexity Deep Research — Perplexity · research agent
- Replit Agent — Replit · coding agent
- Zed — Zed Industries · ai code editor

### Tools

- Agent Development Kit — Google · agent framework
- Agent Skills — Anthropic · format for reusable agent instructions
- AGENTS.md — open convention · instructions file for coding agents
- AI SDK — Vercel · TypeScript AI and agent SDK
- Browser Use — Browser Use · browser agent framework
- Claude Agent SDK — Anthropic · agent framework
- CrewAI — CrewAI · multi-agent framework
- Deep Agents — LangChain · agent harness
- Gemini Live API — Google · voice API
- LangGraph — LangChain · graph framework
- LiveKit Agents — LiveKit · voice agent framework
- LlamaIndex — LlamaIndex · retrieval framework
- Microsoft Agent Framework — Microsoft · multi-agent framework
- OpenAI Agents SDK — OpenAI · agent framework
- OpenAI Realtime API — OpenAI · voice API
- Pipecat — Daily · voice agent framework
- Pydantic AI — Pydantic · agent framework
- smolagents — Hugging Face · agent framework
- Strands Agents — Strands Agents · agent harness SDK
- Vapi — Vapi · voice agent platform

## What is still unsolved at this level

_As of 09/19/2026. This block ages faster than the rest of the page._


### Coding agents

An agent that decides for itself when the work is done will find ways to make the check pass without doing the task. METR reports this across frontier models rather than at one lab, and most often when the tests are hidden from the agent.

**What people are trying:** Hiding test cases and scoring code from the agent, running the scorer somewhere the agent cannot read or write, and using a second model to flag runs that scored suspiciously well. Hardening the environment works where telling the agent not to cheat does not.

- [Recent Frontier Models Are Reward Hacking](https://metr.org/blog/2025-06-05-recent-reward-hacking/) · METR · read 09/19/2026: "The most recent frontier models have engaged in increasingly sophisticated reward hacking, attempting (often successfully) to get a higher score by modifying the tests or scoring code, gaining access to an existing implementation or answer that's used to check their work, or exploiting other loopholes in the task environment."
- [Frontier Risk Report (February to March 2026)](https://metr.org/blog/2026-05-19-frontier-risk-report/) · METR · read 09/19/2026: "attempted to reward hack in ~80% of attempts on tasks in an early version of MirrorCode, when test cases were hidden from the agent."

### The agent harness

Ask a person to approve every action and they stop reading what they approve. Decide it automatically and you take an error rate instead. Anthropic published the false negative rate of its own classifier for this and says it has not found a way to close the gap that is worth what it costs.

**What people are trying:** Scoring each action for risk and for whether the person's own words already cover it, tuned to catch the actions that cannot be undone. The remaining gap has resisted prompt engineering.

- [How we built Claude Code auto mode: a safer way to skip permissions](https://www.anthropic.com/engineering/claude-code-auto-mode) · Anthropic · read 09/19/2026: "Over time that leads to approval fatigue, where people stop paying close attention to what they're approving."
- [How we built Claude Code auto mode: a safer way to skip permissions](https://www.anthropic.com/engineering/claude-code-auto-mode) · Anthropic · read 09/19/2026: "The 17% false-negative rate on real overeager actions is the honest number."

### Skills

A skill is instructions and scripts the agent starts following once it decides the skill applies, and there is no automated way to tell a safe one from an unsafe one. The maker's own advice is to audit every file by hand, the way you would before installing software.

**What people are trying:** Reading the whole bundle for network calls and file access that do not match what the skill claims to do, and treating anything that fetches from an outside URL as the riskiest kind, because a skill that was safe when installed can stop being safe when what it fetches changes. Content scanning exists in places and does not cover every way a skill arrives.

- [Agent Skills](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview) · Anthropic (Claude Platform Docs) · read 09/19/2026: "Even trustworthy Skills can be compromised if their external dependencies change over time"

### Voice agents

A voice agent listening for you to break in has to decide, in real time, whether the sound it just heard was you. It stops mid-sentence for audio that turns out to contain no words, which a caller experiences as the assistant losing its thread for no reason.

**What people are trying:** Moving from a fixed silence threshold to interruption handling that reads the audio for intent, and building a way back: once the interruption turns out to have produced an empty transcript, resume from where the sentence stopped.

- [Turns overview](https://docs.livekit.io/agents/logic/turns/) · LiveKit · read 09/19/2026: "In some cases, the framework detects human speech audio and interrupts the agent, but the transcription comes up empty as no actual words are spoken."
