# Lead agent and workers

_Level 06 · Teams of Agents · sourced_

A lead agent splits the task and hands parts to other agents.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a coordinator splitting a task into focused assignments and assembling their results. Inspect whether each worker had enough context and whether the combined answer resolves overlapping or conflicting findings.

**Assumptions:** Worker outputs are claims requiring integration, not independent proof merely because several agents produced them.

**Design choices:** Delegate separable work with clear deliverables. Keep shared constraints in every assignment and retain synthesis responsibility with the coordinator.

**Request:** Compare venues for cost and accessibility with sources.

**Starting evidence:** Budget $500; step-free entry required. Fictional worker evidence: A costs $450 and its venue sheet confirms step-free entry; B costs $400 but access is undocumented. Workers inspect pricing, transport, and accessibility.

**Action and control:** Lead assigns bounded research tasks and reconciles findings; workers do not book or redefine requirements.

**Stage records (authored, not executed):**

### Input record

Budget $500; step-free entry required. Fictional worker evidence: A costs $450 and its venue sheet confirms step-free entry; B costs $400 but access is undocumented. Workers inspect pricing, transport, and accessibility.

What changed: Establish the facts supplied for this version of the task.

### Design note

Delegate separable work with clear deliverables. Keep shared constraints in every assignment and retain synthesis responsibility with the coordinator.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Lead assigns bounded research tasks and reconciles findings; workers do not book or redefine requirements.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

A meets supplied criteria; B accessibility unverified. Return a sourced comparison, not a purchase.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Task briefs, worker findings with sources, conflict resolution, and a consolidated recommendation without automatic purchase.

If the result falls short:
If a worker fails or results disagree, identify the missing evidence and reassign or investigate only the affected part. Do not average incompatible conclusions.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for comparisons, research, or implementation subtasks. A single agent is often preferable when the work shares too much state to divide cleanly.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** A meets supplied criteria; B accessibility unverified. Return a sourced comparison, not a purchase.

**Change something — Workers use different attendance counts:** Normalize assumptions and redo affected estimates before ranking venues.

**Decision:** Can the lead average incompatible estimates?

**Answer:** No; reconcile assumptions first.

**Why:** Workers may duplicate work or return incompatible assumptions; delegation must have clear scope and evidence.

**Review criteria:** Task briefs, worker findings with sources, conflict resolution, and a consolidated recommendation without automatic purchase.

**Recovery:** If a worker fails or results disagree, identify the missing evidence and reassign or investigate only the affected part. Do not average incompatible conclusions.

**Adapt it:** Use this for comparisons, research, or implementation subtasks. A single agent is often preferable when the work shares too much state to divide cleanly.

Level 6 starts where a single model stops being the whole team. A lead model reads the task,
decides how to split it, and hands each piece to a worker (itself either one call or a
[single agent](/gradient_ascent/techniques/single-agent/) loop) and a lead call combines what
comes back. Anthropic names this shape orchestrator-workers: "a central LLM dynamically breaks
down tasks, delegates them to worker LLMs, and synthesizes their results"[1], well
suited, it says, to work "where you can't predict the subtasks needed"[1].

What makes this level 6 and not level 5 is not that several models run a loop; a single agent
already does that. It is that one model's own output now decides what *other* models are asked to
do. Your code still runs every worker, moves every message and result between lead and worker,
and enforces a cap on how many workers may spawn and how many tokens the whole team may spend:
caps the lead cannot see or override, the same way a single agent's step cap works.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 6 · Lead agent and workers. The same steps are described in the sections below._

## Practical guidance

You will meet this shape inside a product rather than switch it on yourself. Claude Code's
subagents are the clearest version a reader outside a research lab has likely used. Each one runs
"in its own context window with a custom system prompt, specific tool access, and independent
permissions", and the handoff is described this way: "When Claude encounters a task that matches
a subagent's description, it delegates to that subagent, which works independently and returns
results"[3]. Anthropic's own Research feature works the same way: "the lead agent
analyzes it, develops a strategy, and spawns subagents to explore different aspects
simultaneously"[2]. Grok Build and Devin Desktop do the same under their own names: the
shape is worth recognizing even though none of it is yours to configure.

The reason a product delegates like this instead of answering directly is capacity, not
showmanship. Anthropic puts it this way: "Subagents facilitate compression by operating in
parallel with their own context windows, exploring different aspects of the question
simultaneously before condensing the most important tokens for the lead research agent"[2]. That is the job a team buys: a question too broad for one pass, split into pieces small
enough to finish.

Two things are worth asking rather than assuming. Whether the product caps how many workers it
spawns: Anthropic's own team found that without a specific brief, "agents duplicate work, leave
gaps, or fail to find necessary information"[2], and in an early version, "agents made
errors like spawning 50 subagents for simple queries"[2]. Ask directly: "How many
workers did you use for this, and did any of them cover the same ground?" If a multi-part answer
looks thinner than the question deserved, an uncapped or duplicated team is the likely reason, not
a missing answer.

Then expect the bill. Anthropic states its own system's price plainly: "In our data, agents
typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15×
more tokens than chats"[2]. That 15× is Anthropic's own measurement of its own system,
not a rate to expect elsewhere, but it sets the trade correctly: a team costs more per answer than
one pass, in exchange for covering more ground than one pass could. If your question is narrow
enough for a single assistant to answer directly, try that first: cheaper, and nothing here to set
up.

## Implementation details

One call asks the lead to split the question into independent sub-questions, one per line, at
most `MAX_WORKERS` (3 by default). Code parses that text into a list, drops exact duplicates, and
caps it at `MAX_WORKERS` if the lead asked for more: the lead's split is a proposal the code is
free to cut down, never a command code obeys blindly. Anthropic's own research system names what a
good brief needs: "Each subagent needs an objective, an output format, guidance on the tools and
sources to use, and clear task boundaries"[2]; this example's brief is just the
sub-question text, the simplest version of that, since every worker already shares one tool (the
corpus) and one output shape (a cited answer).

Each worker is `examples.rag.run.run`, imported and called unmodified: a level-2 single call, not
a loop, which is the cheap end of what a worker can be. A team that needed a
worker to search iteratively would hand it `examples.agentic_rag.run.run` instead; nothing else in
this file would change, since both share the same `(question, model, embedder, tracer) -> Answer`
signature.

`examples/orchestrator_workers/run.py` (lines 76-117)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    max_workers: int = MAX_WORKERS,
    max_team_tokens: int = MAX_TEAM_TOKENS,
) -> Answer:
    subquestions = _split(question, model, tracer)

    if len(subquestions) > max_workers:
        tracer.record(
            kind="code",
            decided_by="code",
            title="Cap the team",
            detail=f"lead asked for {len(subquestions)} workers, capped at {max_workers}",
        )
        subquestions = subquestions[:max_workers]

    worker_answers: list[tuple[str, Answer]] = []
    for i, subq in enumerate(subquestions, start=1):
        team_tokens = tracer.tokens_in_total() + tracer.tokens_out_total()
        if team_tokens >= max_team_tokens:
            tracer.record(
                kind="code",
                decided_by="code",
                title="Team token budget reached",
                detail=f"stopping before worker {i} of {len(subquestions)}: {team_tokens} >= {max_team_tokens}",
            )
            break
        tracer.record(kind="code", decided_by="code", title=f"Spawn worker {i}", detail=subq)
        worker_answer = rag_worker(subq, model, embedder, tracer, corpus_dir=corpus_dir)
        worker_answers.append((subq, worker_answer))

    if not worker_answers:
        return Answer(text="No worker returned an answer.", citations=[])

    combined_text = _combine(question, worker_answers, model, tracer)
    retrieved = sorted({c for _, a in worker_answers for c in a.retrieved_sources})
    return Answer.from_text(combined_text, retrieved_sources=retrieved)
```

`tracer.tokens_in_total() + tracer.tokens_out_total()` is the whole team's running spend, checked
before every worker spawns; once it passes `MAX_TEAM_TOKENS` (6,000 by default) the remaining
sub-questions are dropped and the run ends with whatever workers already answered, rather than
spawning one more. The split call is the only step in this file with `decided_by: "model"`: the
lead's output is what picks which sub-questions exist and how many workers run. Every worker's own
steps stay `decided_by: "code"`, the same as [RAG](/gradient_ascent/techniques/rag/)'s page,
because retrieval inside a worker is still fixed. The diagram above draws that one call as two
dashed edges, because a reader has to see both assignments happen; its own note says so, and the
recorded trace counts the decision once.

Run it yourself:

`examples/orchestrator_workers/README.md` (lines 16-16)

```text
python -m examples.orchestrator_workers --model stub:scripted
```

CrewAI's own README describes a similar split. It lists what Crews enable, and the first two
entries are "Natural, autonomous decision-making between agents" and "Dynamic task delegation and
collaboration"[4]; its optional hierarchical process "automatically assigns a manager
to the defined crew to properly coordinate the planning and execution of tasks through delegation
and validation of results"[4]. That manager checks the work as well as parceling it out,
which this example's combine step does not. Microsoft's AutoGen, also registered against this
technique, now carries a maintenance notice: "AutoGen is now in maintenance mode. It will not
receive new features or enhancements and is community managed going forward." and "New users
should start with Microsoft Agent Framework."[5] A framework named in a tutorial today
may not be the one to build on by the time you read this.

## When you do not need this

Try a [single agent](/gradient_ascent/techniques/single-agent/) first if one model, in one loop,
can hold the whole task in its own context window: most tasks can. A team only pays for itself
once the work genuinely does not fit one window or one line of reasoning.

Try [parallel calls](/gradient_ascent/techniques/parallelization/) instead if you already know,
before the question arrives, what the fixed set of subtasks is: sectioning a document into three
known parts, for instance. That costs the same every run and needs no lead call to decide anything.

The same test in engineering terms: a characterization sweep over five prototype boards, four
input voltages, three load currents and three ambients is a set of conditions written down before
the run starts, so it is a nested loop with no model anywhere in it, not a team. Splitting it
across agents buys nothing a loop does not already give you, and costs a model call per condition.

Move up to a lead and workers once the split itself cannot be written down in advance: the number
and shape of the subtasks depend on what the specific question turns out to need.

## Failure modes

### Duplicated work

- **How to notice it:** Two or more workers researched the same sub-question from slightly different angles, wasting the tokens of every worker but the first, because the lead's split overlapped instead of dividing the task.
- **How to test for it:** Read every worker's sub-question side by side. Two that would be answered by the same passage of the same document are a duplicate, whatever words the lead used to phrase them.

### Runaway spawning

- **How to notice it:** The lead asks for far more workers than the question has independent parts, and the team cost multiplies with every one, whether or not any of them found something the others missed.
- **How to test for it:** Count the sub-questions the split step actually proposed against the worker cap. A simple question that asks for the cap's full width, every time, is asking for more workers than it needs.

### The lead drops a worker at combine time

- **How to notice it:** A worker returned a real, cited answer, but the combined final answer never uses it. This is the same failure RAG has when a retrieved passage goes unused, one level up.
- **How to test for it:** Compare every worker's citations against the final answer's citations. A worker's citation that never appears in the combined answer was dropped, not wrong.

### The team budget ships a partial answer

- **How to notice it:** The token cap is reached before every worker ran, and the lead combines only the workers that did, silently unless the run is inspected for how many sub-questions the split actually proposed.
- **How to test for it:** Script a split that proposes more sub-questions than a small token budget can afford (this page's own test suite does exactly this) and confirm the run still returns an answer built from whichever workers actually ran.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, 2 workers (split, 2 workers, combine):** 4
- **Model calls, worst case (3 workers, all spawn):** 5
- **Tokens in, the split call:** ~80
- **Wall time, 2 workers run one after another:** ~3.5s

**Compared with a single agent answering the same question (level 5).** Every worker pays roughly what a RAG call alone costs, on top of the split and combine calls, so a two-worker run costs on the order of three single calls, not one: before counting a bigger team or a worker that is itself a loop.

## How to Evaluate It

See [a team of agents that improves your project brief](/gradient_ascent/examples/reviewer-feedback-loop/). A lead coordinates a writer, parallel receiving agents, and independent reviewers, then decides whether to request a revision, ask the user, or return the result. The example shows role boundaries, shared artifacts, failure handling, and the user-facing outcome.

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

`orchestrator_workers` answers the same question-about-the-documents task `rag` and
`single_agent` are scored on: every worker cites what it retrieved, and the lead's combined
answer keeps those citations, so it fits the site's 60-question set the same way (exact or
rubric match, citation hit rate) plus the split-specific numbers a team adds: workers spawned per
question, and the share of questions where the team token budget cut a worker off before it ran.

`scripts/eval_run.py` counts `orchestrator_workers` among the examples the question set can
score, alongside `agent_graphs` and `debate_review` (see `docs/EVALS.md`). No result file exists
for it yet, so this page cannot say a number for any of it. Run
`python scripts/eval_run.py --example orchestrator_workers --model <spec> --dry` to project the
cost of a real run first: a team's projection is the one to look at before spending, because the
ceiling counts every worker the cap allows.

## Run it

**What to monitor.** Workers spawned per question against the cap, the share of splits that propose a duplicate sub-question, and the share of runs where the team token budget cut a worker off before it ran.

**Cost at volume.** Cost multiplies with team size, not just question count: a split that asks for the full worker cap on every question costs several times what a single-agent answer to the same question would, whether or not the extra workers found anything the first one missed.

**How it fails in production.** The lead asks for more workers than the question has independent parts, or two workers investigate the same thing from different angles, so the team spends several times a single agent's cost without a proportional gain in the answer.

**What to log.** The split call's full text, every worker's sub-question and citations, which worker (if any) the team budget cut off, and the lead's combine call, so a bad answer traces back to a bad split, a dropped citation, or a genuine gap no worker covered.

## Try it

1. **Use it.** Give a coding agent that has subagents (Claude Code, for one) a task big enough that it might delegate part of it. Does it tell you it spawned a subagent, and if so, what was that subagent asked to do?
2. **Build it.** Run python -m examples.orchestrator_workers --model stub:scripted from the repo root. The lead splits one two-part question into two, spawns a worker for each, and merges both answers with both citations. Run it again with --model stub and the split never happens: the echo comes back as one placeholder sub-question, so one worker is spawned and the merge has one thing to merge. For the caps, run python -m unittest tests.test_example_orchestrator_workers -v, which scripts a lead into the worker cap and the team token budget.
3. **Either lane.** Write a two-part question a single call could not answer well, then write what you would tell two separate people to go find, if you were the lead instead of a model. Compare that split to what the example's stub test scripts the lead to propose.


## Sources

1. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic, 2024-12-19 (accessed 2026-09-19)
2. [How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system) — Anthropic, 2025-06-13 (accessed 2026-09-19)
3. [Subagents](https://code.claude.com/docs/en/subagents) — Anthropic (Claude Agent SDK documentation) (accessed 2026-09-19)
4. [crewAI](https://github.com/crewAIInc/crewAI) — CrewAI (accessed 2026-09-19)
5. [AutoGen](https://github.com/microsoft/autogen) — Microsoft (accessed 2026-09-19)


Last reviewed 2026-09-19.
