# Level 06 · Teams of Agents

_Agents coordinate work_

Agents divide, coordinate, or review work across separate contexts. A coordinator can combine their findings, and the agents may use the same model or different models. Coordination adds overhead, and separate reviewers can still make correlated mistakes.


## Who decides the next step

Several agents coordinate, delegate, or review work; they can use the same underlying model.


## What is at this level

- [Lead agent and workers](/gradient_ascent/techniques/orchestrator-workers/) (sourced): A lead agent splits the task and hands parts to other agents.
- [Agent graphs](/gradient_ascent/techniques/agent-graphs/) (sourced): Describing a team of agents and how work passes between them.
- [Review and debate](/gradient_ascent/techniques/debate-review/) (sourced): Agents that check, or argue with, each other's work.

## Upgrade conditions

- **Lead agent and workers → Long-running tasks:** The task cannot finish inside one bounded team run and has to pick up again across separate sessions.
- **Agent graphs → Organizations of agents:** The roster of agents itself has to change while the work is in progress, not stay fixed for one run.
- **Review and debate → Always-on assistants:** The checking has to run continuously against everything a standing assistant does, not once against one finished draft.

## Named products, tools and models


### Products

- Claude Code subagents — Anthropic · multi-agent feature of a coding agent
- Claude Research — Anthropic · research agent
- Devin Desktop — Cognition · coding agent in an editor
- Grok Build — SpaceXAI · coding agent
- Grok Heavy — SpaceXAI · several agents answering one question
- Muse — Meta · always-on agent

### Tools

- Agent Development Kit — Google · agent framework
- Agent2Agent (A2A) Protocol — open standard · protocol between agents
- AutoGen — Microsoft · multi-agent framework
- Claude Agent SDK — Anthropic · agent framework
- CrewAI — CrewAI · multi-agent framework
- Deep Agents — LangChain · agent harness
- LangGraph — LangChain · graph framework
- Microsoft Agent Framework — Microsoft · multi-agent framework
- OpenAI Agents SDK — OpenAI · agent framework
- Strands Agents — Strands Agents · agent harness SDK

## What is still unsolved at this level

_As of 09/19/2026. This block ages faster than the rest of the page._


### Lead agent and workers

Run the workers one at a time and the lead waits on the slowest one and cannot steer any of them mid-task. Run them at once and you take on keeping their results, their state and their errors consistent. Nobody has both.

**What people are trying:** Anthropic runs subagents one at a time today because it is the easier half to coordinate, and treats concurrent execution as worth having once the coordination problems are handled rather than as something already solved.

- [How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system) · Anthropic · read 09/19/2026: "But this asynchronicity adds challenges in result coordination, state consistency, and error propagation across the subagents."
  - Dated June 2025. It is still the maker's own account of this architecture, and nothing newer from them supersedes it.

### Agent graphs

When one agent hands a task to another, the authorization to act on it often travels with the task. Passed in band, that credential is visible to every agent in the chain and not only to the one it was meant for.

**What people are trying:** The Agent2Agent specification asks that a credential be bound to the agent that originated the request, and encrypted when it carries anything sensitive, so only that agent can use it. Its own preference is to deliver credentials out of band rather than through the chain at all.

- [Agent2Agent (A2A) Protocol Specification](https://a2a-protocol.org/latest/specification/) · Agent2Agent Protocol Project, Linux Foundation · read 09/19/2026: "In-band credential exchange can allow credentials to be passed across chains of multiple A2A agents, exposing those credentials to each agent participating in the chain."

### Review and debate

Asking several models to review the same work looks like several opinions and is not. Their mistakes are correlated, so a panel carries a fraction of the independent judgment its size suggests, and one measurement found nine judges worth about two independent votes.

**What people are trying:** Measuring a panel against what genuinely independent voting would give, rather than assuming more reviewers means more reliability, and testing whether cleverer ways of combining votes recover the difference. In that measurement they mostly did not.

- [Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels](https://arxiv.org/abs/2605.29800) · arXiv · read 09/19/2026: "the 9 judges effectively provide only about 2 independent votes' worth of information."
- [Correlated Errors in Large Language Models](https://arxiv.org/abs/2506.07962) · arXiv · read 09/19/2026: "We find substantial correlation in model errors -- on one leaderboard dataset, models agree 60% of the time when both models err."
