# Organizations of agents

_Level 07 · Always-on agents · sourced_

Large groups of agents with roles and shared goals.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow several groups coordinating related work and shared resources. Inspect how responsibilities, dependencies, and escalation affect the result when one group falls behind.

**Assumptions:** More agents increase coordination demands and can duplicate mistakes. Shared objectives and ownership need to be explicit.

**Design choices:** Choose teams only when specialization or parallel capacity justifies communication overhead. Centralize decisions that affect shared resources or conflicting priorities.

**Request:** Coordinate a simulated launch across documentation, support, and review teams.

**Starting evidence:** Shared fact: feature X delayed. Each team has an owner and budget.

**Action and control:** Assign responsibilities and synchronize shared facts before teams write outputs.

**Stage records (authored, not executed):**

### Input record

Shared fact: feature X delayed. Each team has an owner and budget.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose teams only when specialization or parallel capacity justifies communication overhead. Centralize decisions that affect shared resources or conflicting priorities.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Assign responsibilities and synchronize shared facts before teams write outputs.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Docs and support mark X unavailable. Consolidated readiness records unresolved dependencies; no launch is authorized.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Dependency board, ownership, conflicting updates, shared budget, escalation, and an evidence-based launch readiness report.

If the result falls short:
If ownership is unclear or work is duplicated, reconcile state and assign a single accountable owner for the decision. More delegation is not necessarily the remedy.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to large simulated projects or organizational workflows. Start with the smallest team structure that improves the task and measure coordination cost as well as output.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Docs and support mark X unavailable. Consolidated readiness records unresolved dependencies; no launch is authorized.

**Change something — One team uses the old release brief:** Identify and repair the shared-state mismatch instead of adding agents.

**Decision:** Will more workers fix stale shared assumptions?

**Answer:** No; reconcile state and ownership first.

**Why:** More agents increase coordination costs and can spread bad assumptions; compare a smaller team baseline.

**Review criteria:** Dependency board, ownership, conflicting updates, shared budget, escalation, and an evidence-based launch readiness report.

**Recovery:** If ownership is unclear or work is duplicated, reconcile state and assign a single accountable owner for the decision. More delegation is not necessarily the remedy.

**Adapt it:** Apply this to large simulated projects or organizational workflows. Start with the smallest team structure that improves the task and measure coordination cost as well as output.

An organization of agents is several standing agents with distinct roles and a shared goal, not
one agent calling another for one task and getting an answer back. This page sits at level 7,
not level 6, because the roster itself (who exists, what they are working on, when a new round
of work starts) keeps running and gets decided along the way, rather than being fixed by a
person for one job and torn down after.

This is mostly frontier. What ships today is closer to [orchestrator-workers](/gradient_ascent/techniques/orchestrator-workers/) and
[agent graphs](/gradient_ascent/techniques/agent-graphs/) wearing a bigger name than the standing,
many-role organizations described below and in the research this page cites. Keep that
distinction in mind reading it: a documented, shipping framework assigns roles and runs them
through fixed procedures; a research paper reports what a simulation of many agents actually did
under study conditions; anything past those two is this page's own reasoning about where the
approach runs into trouble, and it says so.

This page is sourced, not measured: what these many-agent systems do comes from their makers' and
researchers' own papers, and none has been run and scored here. It is illustrated.

_The web page for this technique includes an interactive step-through of Level 7 · Organizations of agents. The same steps are described in the sections below._

## Practical guidance

There is nothing here to sign up for. A standing organization of agents is a framework a developer
assembles, not a product you turn on, and this site can name no service selling one today. If what
you want is an assistant that keeps working while you are away, that is
[always-on assistants](/gradient_ascent/techniques/agent-teammates/), and it is the one shape at
this level you can actually buy.

The nearest named thing is a developer's framework. MetaGPT's own README states the idea outright,
as "Assign different roles to GPTs to form a collaborative entity for complex tasks", and lists
"product managers / architects / project managers / engineers" as the roles it
includes[3]. That is a fixed roster running a fixed procedure on one task, closer to a
simulated org chart than to agents that keep existing and decide what to work on next.

One measured finding does travel to any product that passes an answer through several agents
before you see it, whatever the product calls that. A 2026 paper ran 500 cascades across 10
knowledge domains on three models, 1,250 responses in all, passing each answer along a chain of
agents that revise it. In three-agent chains the normalized hallucination score fell (from 0.422
at the first agent to 0.272 at the last) and factual accuracy fell with it, "from 0.789 to 0.769",
which the authors call "a trade-off between hallucination suppression and factual
preservation"[4]. Both movements are small, over one benchmark, at one chain length, so
read it as a direction rather than a rate. What it means in a meeting: when a vendor says several
agents checked the answer, ask what the chain dropped, not only what it caught.

If someone proposes running a roster like this inside your organization, two questions settle it
before any demo. "How is the whole thing stopped at once?" A roster needs one switch that drains
the work and refuses new claims, not a role-by-role hunt while the rest keep working. And "when a
mistake reaches a customer through three roles that each added something, whose name is on it?"
No maker's page or paper this site has read answers the second one. A calibration certificate
carries the name of whoever signed it for exactly this reason, and a roster of agents has no
equivalent unless somebody writes one down. Get it in writing first, or do not start.

## Implementation details

The example is deliberately small: three roles (researcher, writer, reviewer) share one task
board, and a coordinator model decides which open task goes to which role. What each role would
do with an assigned task is out of scope; a role agent would use whatever pattern in this manual
fits its own job. What this example isolates is only the part specific
to an organization: the shared board, and an assignment step the model does not fully control.

`examples/organizations_swarms/run.py` (lines 71-114)

```python
def coordinate(board: Board, model: Model, tracer: Tracer) -> list[dict]:
    """One coordination round. Returns the assignments actually made — which can be fewer than
    the coordinator asked for, since every proposed assignment is checked against the board and
    the budget before it counts.

    The budget is per round and refills here, at the start of each one. Claims are not: a task
    claimed in an earlier round is still claimed, because the board persists and the counters do
    not. Both facts have to be tested, and testing the second one needs a round that still has
    an open task to offer -- otherwise the round returns before the model is ever called and the
    test passes without checking anything.
    """
    board.assigned_count = {r: 0 for r in ROLES}  # the budget is per round, so it starts full
    if not any(t.status == "open" for t in board.tasks):
        return []

    messages = [Message(role="system", content=SYSTEM), Message(role="user", content=_digest(board))]
    completion = model.complete(messages, tools=[ASSIGN_TOOL], max_tokens=200)
    proposed = list(completion.tool_calls)
    desc = ", ".join(f"assign({c.arguments.get('task_id')}, {c.arguments.get('role')})" for c in proposed) or "no assignments proposed"
    tracer.record(
        kind="model", decided_by="model", title="Coordinator assigns open tasks to roles",
        detail=desc, tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )

    made: list[dict] = []
    for call in proposed:
        task_id = str(call.arguments.get("task_id", ""))
        role = str(call.arguments.get("role", ""))
        task = board.get(task_id)
        if task is None or task.status != "open":
            tracer.record(kind="code", decided_by="code", title="Refuse: task is not open", detail=f"{task_id} (already claimed, or does not exist)")
            continue
        if role not in ROLES:
            tracer.record(kind="code", decided_by="code", title="Refuse: no such role", detail=f"{role} for {task_id} (roles are fixed in code, not named by the coordinator)")
            continue
        if board.assigned_count[role] >= BUDGET_PER_ROLE:
            tracer.record(kind="code", decided_by="code", title="Refuse: role is over budget", detail=f"{role} for {task_id}")
            continue
        task.status = "claimed"
        task.assigned_to = role
        board.assigned_count[role] += 1
        tracer.record(kind="code", decided_by="code", title="Claim the task for the role", detail=f"{task_id} -> {role}")
        made.append({"task": task_id, "role": role})
    return made
```

One coordinator call runs per round, and that single call is the round's only
`decided_by: "model"` step however many assignments it proposes; the diagram above draws each
proposal as its own dashed edge and says so under its tally.

`BUDGET_PER_ROLE` is a fixed cap the coordinator cannot raise by asking; a role that has already
been assigned two tasks this round gets no more, whatever the coordinator's output requests next.
Claiming is checked against the board itself, not against what the coordinator believes is true:
`task.status != "open"` refuses an assignment outright, so if the coordinator's own single call
proposes the same task to two different roles (nothing stops a model from repeating itself), only
the first proposal is ever honored. The second reads the board
fresh, sees the task is no longer open, and is refused, not silently overwritten and not queued
to run later. That check is what the shared board is actually for: a coordinator's plan is a
proposal against it, not an instruction the board has to accept.

MetaGPT's own paper describes a fuller version of the same idea, fixed procedure rather than a
per-round budget: it "utilizes an assembly line paradigm to assign diverse roles to various
agents, efficiently breaking down complex tasks into subtasks involving many agents working
together" and encodes "Standardized Operating Procedures (SOPs) into prompt sequences for more
streamlined workflows"[2]. The paper's claim for that design is scoped: "On
collaborative software engineering benchmarks, MetaGPT generates more coherent solutions than
previous chat-based multi-agent systems"[2]. That is the authors' result on those
benchmarks, against the systems they chose to compare with, and not a result this site has
measured.

A different line of research reaches for coordination without any fixed roster at all. Park et
al.'s Generative Agents paper populated a sandbox (twenty-five agents in a small town, no shared
board and no fixed roles, just memory, planning and reflection) and reports
that "starting with only a single user-specified notion that one agent wants to throw a
Valentine's Day party, the agents autonomously spread invitations to the party over the next two
days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up
for the party together at the right time"[1]. The paper's subject is believable human
behavior in a simulation, not work getting done: read it for the mechanism, which is that
coordination can emerge without a coordinator role, and not as evidence that a roster of agents
will organize itself around a real task.

The two rules look alike and behave differently, which is worth testing separately: the budget is
spent per round and refills at the start of the next one, while a claim is permanent until the
task is done. `tests/test_example_organizations_swarms.py` scripts a coordinator that over-assigns
one role, one that proposes the same task twice in a single call, and one that tries to re-claim
a task in a later round: that last test keeps a second task open on purpose, because a round with
nothing open returns before the model is ever called and would otherwise pass without checking
anything. Run it yourself:

`examples/organizations_swarms/README.md` (lines 12-12)

```text
python -m examples.organizations_swarms --model stub:scripted
```

## When you do not need this

Try [orchestrator-workers](/gradient_ascent/techniques/orchestrator-workers/) or
[agent graphs](/gradient_ascent/techniques/agent-graphs/) first for almost anything you are
actually building today. A fixed roster assembled for one job, run once, and torn down
afterward is level 6, cheaper to reason about, and is what the one shipping framework this page
names actually does.

Move up to a standing organization only once the roster itself needs to persist and change
between tasks (new roles added, old ones retired, work assigned across many separate requests
rather than one) and once you have an answer, in writing, for who is accountable when it acts.

## Failure modes

### A task is claimed by two roles at once

- **How to notice it:** Two roles both believe they own the same task and either duplicate the work or step on each other's output.
- **How to test for it:** Script a coordinator that proposes the same task to two roles in one call (this page's own test suite does exactly this) and confirm only the first proposal is honored, not that both fail, not that both succeed.

### A role quietly exceeds its budget

- **How to notice it:** One role ends up doing most of the work for a round because nothing capped how much it could take on, defeating the point of having separate roles at all.
- **How to test for it:** Script a coordinator that tries to assign every open task to the same role and confirm the count assigned never exceeds the fixed budget, regardless of how many the coordinator asked for.

### An error compounds instead of getting caught

- **How to notice it:** A mistake made by one role passes through a second and third role that were supposed to check it, and comes out the other end looking more confident than it started, not less.
- **How to test for it:** Trace one output back through every role that touched it and check whether each one actually verified something or just passed the previous role's claim along unchanged.

### Nobody can say which role is responsible for a bad result

- **How to notice it:** A wrong or harmful output reaches a person and the postmortem cannot identify which role introduced the error, because the board and the logs record what was assigned, not what each role actually checked before acting.
- **How to test for it:** Pick a finished task and try to reconstruct, from the logs alone, which role's decision the final output actually depended on. If you cannot, the logging is not enough for this level, whatever it is enough for at level 6.

### The roster grows because adding a role feels free

- **How to notice it:** Each new role seems to add a capability, but the system as a whole gets slower and harder to predict, and nobody can say what the fourth or fifth role actually improved.
- **How to test for it:** Remove one role and rerun the same tasks. If the outcome does not measurably change, that role was not paying for its share of the coordination cost.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one coordination round:** 1
- **Roles sharing the board:** 3
- **Tokens in, one round:** ~180
- **Wall time, one round:** ~0.5s

**Compared with a single agent (level 5) doing all three roles’ work itself.** The coordination call itself is cheap: one small model call per round. What this strip does not show, and what a real system pays for, is each role’s own work once assigned: three separate agents’ worth of calls this example does not run, on top of the coordination shown here.

## How to Evaluate It

This example coordinates a task board; it does not answer questions about the document corpus, so
the site's shared 60-question set does not apply, the same reason it does not apply to
[always-on assistants](/gradient_ascent/techniques/agent-teammates/)' policy example. What would
be measured here: the share of coordination rounds that violate the budget or double-claim a task
under adversarial scripting (this should be exactly zero, and
`tests/test_example_organizations_swarms.py` checks it on the stub every run), how evenly work
actually spreads across roles against how evenly the coordinator intended it to, and (the harder
measurement research has not settled) whether a final output's accuracy degrades as it passes
through more roles, the way the hallucination-cascade paper measures for its own three-agent
chain[4].

## Run it

**What to monitor.** Assignments refused for budget or for a task no longer being open: both should be rare relative to the volume of tasks, and a rising rate of either is worth reading as a sign the roster or the budgets no longer match the workload, not just logging and moving on.

**Cost at volume.** One coordination call per round regardless of roster size, plus each role's own work once assigned; coordination cost itself stays close to flat as the roster grows, but the total system cost does not, since every added role is added work somewhere downstream of this example.

**How it fails in production.** A role silently falls behind its budget every round because the coordinator keeps proposing more than it can take, and nobody notices until a backlog of unclaimed tasks is large enough to see on a dashboard nobody was watching.

**What to log.** Every proposed assignment, whether the board accepted or refused it and why, which role actually touched a task before it was marked done, and the full chain of roles a given output passed through, so an accountability question has an answer that does not require asking any of the agents.

## Try it

1. **Use it.** Read MetaGPT's own documentation for the roles it assigns to a task. For one of those roles, write down what you would need to see in a log to trust that role's output without re-doing its work yourself.
2. **Build it.** Run python -m examples.organizations_swarms --model stub:scripted from the repo root. The coordinator proposes five assignments in one call; the board claims four and refuses the fifth, which hands T1 to a second role after an earlier proposal took it.
3. **Either lane.** Cause the over-budget failure on purpose: lower BUDGET_PER_ROLE in examples/organizations_swarms/run.py to 1 and run the same command. The researcher's second task is refused for budget, and T4 is still open at the end.


## Sources

1. [Generative Agents: Interactive Simulacra of Human Behavior](https://arxiv.org/abs/2304.03442) — arXiv (Stanford University, Google Research), 2023-04-07 (accessed 2026-09-19)
2. [MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework](https://arxiv.org/abs/2308.00352) — arXiv, 2023-08-01 (accessed 2026-09-19)
3. [FoundationAgents/MetaGPT](https://github.com/FoundationAgents/MetaGPT) — MetaGPT (GitHub README) (accessed 2026-09-19)
4. [Hallucination Cascade: Analyzing Error Propagation in Multi-Agent LLM Systems](https://arxiv.org/abs/2606.07937) — arXiv, 2026-06-06 (accessed 2026-09-19)


Last reviewed 2026-09-19.
