Level 06 · Teams of Agents

Lead agent and workers

A lead agent splits the task and hands parts to other agents.

Sourced

Concept at a glance

A lead delegates pieces and combines the answers.

Parallel branchesConceptual illustration
A lead delegates pieces and combines the answers.Lead agent leads to Worker A. Worker A leads to Lead combines. Lead agent leads to Worker B. Worker B leads to Lead combines. Lead agent leads to Worker C. Worker C leads to Lead combines. The lead chooses the subtasks; workers return their findings to it.Lead agentDecide how to split the taskWorker AInvestigate one subtaskWorker BInvestigate anotherWorker CInvestigate anotherLead combinesCheck and synthesize findingsA lead delegates pieces and combines the answers.Lead agent leads to Worker A. Worker A leads to Lead combines. Lead agent leads to Worker B. Worker B leads to Lead combines. Lead agent leads to Worker C. Worker C leads to Lead combines. The lead chooses the subtasks; workers return their findings to it.Lead agentDecide how to split the taskWorker AInvestigate one subtaskWorker BInvestigate anotherWorker CInvestigate anotherLead combinesCheck and synthesize findings
Read the connections in words
  • Lead agent → Worker A: Investigate one subtask.
  • Worker A → Lead combines: Check and synthesize findings.
  • Lead agent → Worker B: Investigate another.
  • Worker B → Lead combines: Check and synthesize findings.
  • Lead agent → Worker C: Investigate another.
  • Worker C → Lead combines: Check and synthesize findings.
Key idea

The lead chooses the subtasks; workers return their findings to it.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Lead agent and workers: see it in practice.

A lead agent delegating bounded subtasks to workers and combining their results.

What you’ll walk through

Follow a coordinator splitting a task into focused assignments and assembling their results. Inspect whether each worker had enough context and whether the combined answer resolves overlapping or conflicting findings.

The task in this version

Compare venues for cost and accessibility with sources.

What you’ll learn to check

Task briefs, worker findings with sources, conflict resolution, and a consolidated recommendation without automatic purchase.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Compare venues for cost and accessibility with sources.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Budget $500; step-free entry required. Fictional worker evidence: A costs $450 and its venue sheet confirms step-free entry; B costs $400 but access is undocumented. Workers inspect pricing, transport, and accessibility.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Worker outputs are claims requiring integration, not independent proof merely because several agents produced them.

1 / 6

Apply this to your project

Describe your task to your own model and use Lead agent and workers as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Level 6 starts where a single model stops being the whole team. A lead model reads the task, decides how to split it, and hands each piece to a worker (itself either one call or a single agent loop) and a lead call combines what comes back. Anthropic names this shape orchestrator-workers: “a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results”[1], well suited, it says, to work “where you can’t predict the subtasks needed”[1].

What makes this level 6 and not level 5 is not that several models run a loop; a single agent already does that. It is that one model’s own output now decides what other models are asked to do. Your code still runs every worker, moves every message and result between lead and worker, and enforces a cap on how many workers may spawn and how many tokens the whole team may spend: caps the lead cannot see or override, the same way a single agent’s step cap works.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Lead agent and workers

A lead splits the question into sub-questions no code could predict in advance, each worker answers from its own document, and the lead combines the results with citations.

Level 6 · Teams of Agents
QuestionQuestionMODELLead splits the taskLead splitsthe taskMODELWorker 1Worker 1MODELWorker 2Worker 2MODELLead combines resultsLead combinesresultsAnswerAnswer
0of 1 step so far chosen by the model

Both assignments come from one split call, so the recorded trace counts one model-decided step where the diagram plays two: one model output selected both transitions.

your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 06Your code chose

The question arrives

"What is the DW-300's annual energy use, and how
much does the DR-520's drive belt cost?"
0 tokens · 0 ms

Practical guidance

You will meet this shape inside a product rather than switch it on yourself. Claude Code’s subagents are the clearest version a reader outside a research lab has likely used. Each one runs “in its own context window with a custom system prompt, specific tool access, and independent permissions”, and the handoff is described this way: “When Claude encounters a task that matches a subagent’s description, it delegates to that subagent, which works independently and returns results”[3]. Anthropic’s own Research feature works the same way: “the lead agent analyzes it, develops a strategy, and spawns subagents to explore different aspects simultaneously”[2]. Grok Build and Devin Desktop do the same under their own names: the shape is worth recognizing even though none of it is yours to configure.

The reason a product delegates like this instead of answering directly is capacity, not showmanship. Anthropic puts it this way: “Subagents facilitate compression by operating in parallel with their own context windows, exploring different aspects of the question simultaneously before condensing the most important tokens for the lead research agent”[2]. That is the job a team buys: a question too broad for one pass, split into pieces small enough to finish.

Two things are worth asking rather than assuming. Whether the product caps how many workers it spawns: Anthropic’s own team found that without a specific brief, “agents duplicate work, leave gaps, or fail to find necessary information”[2], and in an early version, “agents made errors like spawning 50 subagents for simple queries”[2]. Ask directly: “How many workers did you use for this, and did any of them cover the same ground?” If a multi-part answer looks thinner than the question deserved, an uncapped or duplicated team is the likely reason, not a missing answer.

Then expect the bill. Anthropic states its own system’s price plainly: “In our data, agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats”[2]. That 15× is Anthropic’s own measurement of its own system, not a rate to expect elsewhere, but it sets the trade correctly: a team costs more per answer than one pass, in exchange for covering more ground than one pass could. If your question is narrow enough for a single assistant to answer directly, try that first: cheaper, and nothing here to set up.

Implementation details

One call asks the lead to split the question into independent sub-questions, one per line, at most MAX_WORKERS (3 by default). Code parses that text into a list, drops exact duplicates, and caps it at MAX_WORKERS if the lead asked for more: the lead’s split is a proposal the code is free to cut down, never a command code obeys blindly. Anthropic’s own research system names what a good brief needs: “Each subagent needs an objective, an output format, guidance on the tools and sources to use, and clear task boundaries”[2]; this example’s brief is just the sub-question text, the simplest version of that, since every worker already shares one tool (the corpus) and one output shape (a cited answer).

Each worker is examples.rag.run.run, imported and called unmodified: a level-2 single call, not a loop, which is the cheap end of what a worker can be. A team that needed a worker to search iteratively would hand it examples.agentic_rag.run.run instead; nothing else in this file would change, since both share the same (question, model, embedder, tracer) -> Answer signature.

examples/orchestrator_workers/run.py · lines 76–117
def run(
    question: str,
    model: Model,
    embedder: Embedder,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    max_workers: int = MAX_WORKERS,
    max_team_tokens: int = MAX_TEAM_TOKENS,
) -> Answer:
    subquestions = _split(question, model, tracer)

    if len(subquestions) > max_workers:
        tracer.record(
            kind="code",
            decided_by="code",
            title="Cap the team",
            detail=f"lead asked for {len(subquestions)} workers, capped at {max_workers}",
        )
        subquestions = subquestions[:max_workers]

    worker_answers: list[tuple[str, Answer]] = []
    for i, subq in enumerate(subquestions, start=1):
        team_tokens = tracer.tokens_in_total() + tracer.tokens_out_total()
        if team_tokens >= max_team_tokens:
            tracer.record(
                kind="code",
                decided_by="code",
                title="Team token budget reached",
                detail=f"stopping before worker {i} of {len(subquestions)}: {team_tokens} >= {max_team_tokens}",
            )
            break
        tracer.record(kind="code", decided_by="code", title=f"Spawn worker {i}", detail=subq)
        worker_answer = rag_worker(subq, model, embedder, tracer, corpus_dir=corpus_dir)
        worker_answers.append((subq, worker_answer))

    if not worker_answers:
        return Answer(text="No worker returned an answer.", citations=[])

    combined_text = _combine(question, worker_answers, model, tracer)
    retrieved = sorted({c for _, a in worker_answers for c in a.retrieved_sources})
    return Answer.from_text(combined_text, retrieved_sources=retrieved)

tracer.tokens_in_total() + tracer.tokens_out_total() is the whole team’s running spend, checked before every worker spawns; once it passes MAX_TEAM_TOKENS (6,000 by default) the remaining sub-questions are dropped and the run ends with whatever workers already answered, rather than spawning one more. The split call is the only step in this file with decided_by: "model": the lead’s output is what picks which sub-questions exist and how many workers run. Every worker’s own steps stay decided_by: "code", the same as RAG’s page, because retrieval inside a worker is still fixed. The diagram above draws that one call as two dashed edges, because a reader has to see both assignments happen; its own note says so, and the recorded trace counts the decision once.

Run it yourself:

examples/orchestrator_workers/README.md · lines 16–16
python -m examples.orchestrator_workers --model stub:scripted

CrewAI’s own README describes a similar split. It lists what Crews enable, and the first two entries are “Natural, autonomous decision-making between agents” and “Dynamic task delegation and collaboration”[4]; its optional hierarchical process “automatically assigns a manager to the defined crew to properly coordinate the planning and execution of tasks through delegation and validation of results”[4]. That manager checks the work as well as parceling it out, which this example’s combine step does not. Microsoft’s AutoGen, also registered against this technique, now carries a maintenance notice: “AutoGen is now in maintenance mode. It will not receive new features or enhancements and is community managed going forward.” and “New users should start with Microsoft Agent Framework.”[5] A framework named in a tutorial today may not be the one to build on by the time you read this.

When you do not need this

Try a single agent first if one model, in one loop, can hold the whole task in its own context window: most tasks can. A team only pays for itself once the work genuinely does not fit one window or one line of reasoning.

Try parallel calls instead if you already know, before the question arrives, what the fixed set of subtasks is: sectioning a document into three known parts, for instance. That costs the same every run and needs no lead call to decide anything.

The same test in engineering terms: a characterization sweep over five prototype boards, four input voltages, three load currents and three ambients is a set of conditions written down before the run starts, so it is a nested loop with no model anywhere in it, not a team. Splitting it across agents buys nothing a loop does not already give you, and costs a model call per condition.

Move up to a lead and workers once the split itself cannot be written down in advance: the number and shape of the subtasks depend on what the specific question turns out to need.

Failure modes

Duplicated work

How to notice it
Two or more workers researched the same sub-question from slightly different angles, wasting the tokens of every worker but the first, because the lead's split overlapped instead of dividing the task.
How to test for it
Read every worker's sub-question side by side. Two that would be answered by the same passage of the same document are a duplicate, whatever words the lead used to phrase them.

Runaway spawning

How to notice it
The lead asks for far more workers than the question has independent parts, and the team cost multiplies with every one, whether or not any of them found something the others missed.
How to test for it
Count the sub-questions the split step actually proposed against the worker cap. A simple question that asks for the cap's full width, every time, is asking for more workers than it needs.

The lead drops a worker at combine time

How to notice it
A worker returned a real, cited answer, but the combined final answer never uses it. This is the same failure RAG has when a retrieved passage goes unused, one level up.
How to test for it
Compare every worker's citations against the final answer's citations. A worker's citation that never appears in the combined answer was dropped, not wrong.

The team budget ships a partial answer

How to notice it
The token cap is reached before every worker ran, and the lead combines only the workers that did, silently unless the run is inspected for how many sub-questions the split actually proposed.
How to test for it
Script a split that proposes more sub-questions than a small token budget can afford (this page's own test suite does exactly this) and confirm the run still returns an answer built from whichever workers actually ran.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

4Model calls, 2 workers (split, 2 workers, combine)
5Model calls, worst case (3 workers, all spawn)
~80Tokens in, the split call
~3.5sWall time, 2 workers run one after another
Compared with a single agent answering the same question (level 5)Every worker pays roughly what a RAG call alone costs, on top of the split and combine calls, so a two-worker run costs on the order of three single calls, not one: before counting a bigger team or a worker that is itself a loop.

How to Evaluate It

See a team of agents that improves your project brief. A lead coordinates a writer, parallel receiving agents, and independent reviewers, then decides whether to request a revision, ask the user, or return the result. The example shows role boundaries, shared artifacts, failure handling, and the user-facing outcome.

60 questionslookupmulti-hopnumericunanswerableconflicting sources

orchestrator_workers answers the same question-about-the-documents task rag and single_agent are scored on: every worker cites what it retrieved, and the lead’s combined answer keeps those citations, so it fits the site’s 60-question set the same way (exact or rubric match, citation hit rate) plus the split-specific numbers a team adds: workers spawned per question, and the share of questions where the team token budget cut a worker off before it ran.

scripts/eval_run.py counts orchestrator_workers among the examples the question set can score, alongside agent_graphs and debate_review (see docs/EVALS.md). No result file exists for it yet, so this page cannot say a number for any of it. Run python scripts/eval_run.py --example orchestrator_workers --model <spec> --dry to project the cost of a real run first: a team’s projection is the one to look at before spending, because the ceiling counts every worker the cap allows.

Run it

What to monitor

Workers spawned per question against the cap, the share of splits that propose a duplicate sub-question, and the share of runs where the team token budget cut a worker off before it ran.

Cost at volume

Cost multiplies with team size, not just question count: a split that asks for the full worker cap on every question costs several times what a single-agent answer to the same question would, whether or not the extra workers found anything the first one missed.

How it fails in production

The lead asks for more workers than the question has independent parts, or two workers investigate the same thing from different angles, so the team spends several times a single agent's cost without a proportional gain in the answer.

What to log

The split call's full text, every worker's sub-question and citations, which worker (if any) the team budget cut off, and the lead's combine call, so a bad answer traces back to a bad split, a dropped citation, or a genuine gap no worker covered.

Try it

  1. Use it

    Give a coding agent that has subagents (Claude Code, for one) a task big enough that it might delegate part of it. Does it tell you it spawned a subagent, and if so, what was that subagent asked to do?

  2. Build it

    Run python -m examples.orchestrator_workers --model stub:scripted from the repo root. The lead splits one two-part question into two, spawns a worker for each, and merges both answers with both citations. Run it again with --model stub and the split never happens: the echo comes back as one placeholder sub-question, so one worker is spawned and the merge has one thing to merge. For the caps, run python -m unittest tests.test_example_orchestrator_workers -v, which scripts a lead into the worker cap and the team token budget.

  3. Either lane

    Write a two-part question a single call could not answer well, then write what you would tell two separate people to go find, if you were the lead instead of a model. Compare that split to what the example's stub test scripts the lead to propose.

How it connects

Before, after and instead of this

Read first

Move up when

  • Long-running tasksThe task cannot finish inside one bounded team run and has to pick up again across separate sessions.

Pages that need this one

Decoded in

Optional: products, tools, and models

7 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 1 more examples
  • Deep Agents Tool or framework · LangChain

    Agent harness

    Checked 09/19/2026
In practice

Research a question with several parts

A lead splits the question, assigns subtasks to workers, and checks their findings before combining them.

Out there

Named products, tools and models

Products4
  • Claude Code subagentsAnthropic · multi-agent feature of a coding agent
  • Claude ResearchAnthropic · research agent
  • Devin DesktopCognition · coding agent in an editor · formerly Windsurf
  • Grok BuildSpaceXAI · coding agent
Tools4
  • AutoGenMicrosoft · multi-agent frameworkSuperseded by Microsoft Agent Framework
  • Claude Agent SDKAnthropic · agent framework
  • CrewAICrewAI · multi-agent framework
  • Deep AgentsLangChain · agent harness

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Building effective agents · Anthropic, 12/19/2024 (accessed 09/19/2026)
  2. How we built our multi-agent research system · Anthropic, 06/13/2025 (accessed 09/19/2026)
  3. Subagents · Anthropic (Claude Agent SDK documentation) (accessed 09/19/2026)
  4. crewAI · CrewAI (accessed 09/19/2026)
  5. AutoGen · Microsoft (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page