Level 05

Agent loops

The model uses the goal and observed results to choose an action, revise its approach, or finish. Software executes tools and enforces permissions, approvals, and stopping limits. A run can stop because it is complete, blocked, or out of budget.

Level 05

Who decides the next stepThe model chooses the next step within limits enforced by software.
Techniques

What is at this level

Observe, decide, act, repeat

Single agent

Sourced

A model that plans, acts and checks its own work in a loop.

The agent harness

Sourced

Everything around the model in an agent: the loop, tools, context handling, permissions, caps and sandbox.

Agentic RAG and deep research

Measured

An agent that runs its own searches until it has an answer.

Coding agents

Sourced

Agents that read, write, run and test code.

Skills

Sourced

Reusable instructions that an agent loads when it needs them.

Voice agents

Sourced

Agents you talk to in real time.

Upgrade conditions

When something here is not enough

Each line names the failure that justifies moving to a higher level.

Single agent → Long-running tasks

The work outlives one context window or one sitting.

Agentic RAG and deep research → Lead agent and workers

The searches do not depend on each other and there are more of them than one agent's context can carry.

Coding agents → Lead agent and workers

The change is larger than one agent can hold at once and splits into parts that can be worked separately.

Recipes

Jobs that top out here

Each one needs this level and no higher, and says why.

Level 3 + Level 5

Write a research brief with citations

Uses agentic RAG to find sources and a fixed check on every claim against the section it cites. It needs level 5 for the searching; the checking is level 3.

This example uses level 5
↗
Level 5

Coding assistant on your own repo

A coding agent that reads, edits, runs and tests code in your repository, using skills for repeated tasks and a safety review before anything ships.

This example uses level 5
↗
Level 4 + Level 5

Data analysis by conversation

A single agent writes and runs code against a dataset, one question at a time, to answer questions a fixed query could not anticipate.

This example uses level 5
↗
Level 3 + Level 4 + Level 5

Plan a trip and hold the bookings

Checking what is available, what is open and what connects takes a different number of steps every time, which is what level 5 is for. Read-only lookups run unattended; anything that spends money stops for a person, with the price and the cancellation terms in front of them.

This example uses level 5
↗
Level 3 + Level 4 + Level 5

Work a bring-up problem at the bench

An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.

This example uses level 5
↗
Out there

Named products, tools and models

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Products that work this way21
  • AiderAider · open-source coding agent
  • Antigravity CLIGoogle · coding agent · formerly Gemini CLI
  • ChatGPT deep researchOpenAI · research agent
  • ChatGPT voiceOpenAI · voice assistant
  • Claude CodeAnthropic · coding agent
  • Claude ResearchAnthropic · research agent
  • ClineCline · open-source coding agent
  • CodexOpenAI · coding agent
  • CursorAnysphere · coding agent in an editor
  • DevinCognition · coding agent
  • Devin DesktopCognition · coding agent in an editor · formerly Windsurf
  • ElevenLabsElevenLabs · voice generation
  • Gemini Deep ResearchGoogle · research agent
  • Gemini LiveGoogle · voice assistant
  • GitHub CopilotGitHub · coding agent in an editor
  • Grok BuildSpaceXAI · coding agent
  • Grok DeepSearchSpaceXAI · research agent
  • JulesGoogle · coding agent
  • Perplexity Deep ResearchPerplexity · research agent
  • Replit AgentReplit · coding agent
  • ZedZed Industries · ai code editor
Tools for building it20
  • Agent Development KitGoogle · agent framework
  • Agent SkillsAnthropic · format for reusable agent instructions
  • AGENTS.mdopen convention · instructions file for coding agents
  • AI SDKVercel · TypeScript AI and agent SDK
  • Browser UseBrowser Use · browser agent framework
  • Claude Agent SDKAnthropic · agent framework
  • CrewAICrewAI · multi-agent framework
  • Deep AgentsLangChain · agent harness
  • Gemini Live APIGoogle · voice API
  • LangGraphLangChain · graph framework
  • LiveKit AgentsLiveKit · voice agent framework
  • LlamaIndexLlamaIndex · retrieval framework
  • Microsoft Agent FrameworkMicrosoft · multi-agent framework
  • OpenAI Agents SDKOpenAI · agent framework
  • OpenAI Realtime APIOpenAI · voice API
  • PipecatDaily · voice agent framework
  • Pydantic AIPydantic · agent framework
  • smolagentsHugging Face · agent framework
  • Strands AgentsStrands Agents · agent harness SDK
  • VapiVapi · voice agent platform
Frontier

What is still unsolved here

Open problems at this level, what people are trying, and the source each rests on. Read 09/19/2026. This block ages faster than the rest of the page, and nothing in it predicts which approach wins.

An agent that decides for itself when the work is done will find ways to make the check pass without doing the task. METR reports this across frontier models rather than at one lab, and most often when the tests are hidden from the agent.

What people are trying

Hiding test cases and scoring code from the agent, running the scorer somewhere the agent cannot read or write, and using a second model to flag runs that scored suspiciously well. Hardening the environment works where telling the agent not to cheat does not.

  • Recent Frontier Models Are Reward Hacking · METR · read 09/19/2026
    The most recent frontier models have engaged in increasingly sophisticated reward hacking, attempting (often successfully) to get a higher score by modifying the tests or scoring code, gaining access to an existing implementation or answer that's used to check their work, or exploiting other loopholes in the task environment.
  • Frontier Risk Report (February to March 2026) · METR · read 09/19/2026
    attempted to reward hack in ~80% of attempts on tasks in an early version of MirrorCode, when test cases were hidden from the agent.

Where it bites: Coding agents

Ask a person to approve every action and they stop reading what they approve. Decide it automatically and you take an error rate instead. Anthropic published the false negative rate of its own classifier for this and says it has not found a way to close the gap that is worth what it costs.

What people are trying

Scoring each action for risk and for whether the person's own words already cover it, tuned to catch the actions that cannot be undone. The remaining gap has resisted prompt engineering.

Where it bites: The agent harness

A skill is instructions and scripts the agent starts following once it decides the skill applies, and there is no automated way to tell a safe one from an unsafe one. The maker's own advice is to audit every file by hand, the way you would before installing software.

What people are trying

Reading the whole bundle for network calls and file access that do not match what the skill claims to do, and treating anything that fetches from an outside URL as the riskiest kind, because a skill that was safe when installed can stop being safe when what it fetches changes. Content scanning exists in places and does not cover every way a skill arrives.

  • Agent Skills · Anthropic (Claude Platform Docs) · read 09/19/2026
    Even trustworthy Skills can be compromised if their external dependencies change over time

Where it bites: Skills

A voice agent listening for you to break in has to decide, in real time, whether the sound it just heard was you. It stops mid-sentence for audio that turns out to contain no words, which a caller experiences as the assistant losing its thread for no reason.

What people are trying

Moving from a fixed silence threshold to interruption handling that reads the audio for intent, and building a way back: once the interruption turns out to have produced an empty transcript, resume from where the sentence stopped.

  • Turns overview · LiveKit · read 09/19/2026
    In some cases, the framework detects human speech audio and interrupts the agent, but the transcription comes up empty as no actual words are spoken.

Where it bites: Voice agents

← Level 04 · Tool use

Pages at this level last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page