Level 07 · Always-on agents

Organizations of agents

Large groups of agents with roles and shared goals.

Sourced

Concept at a glance

Coordinate roles around a shared goal.

Parallel branchesConceptual illustration
Coordinate roles around a shared goal.Shared goal leads to Research role. Research role leads to Shared state. Shared goal leads to Execution role. Execution role leads to Shared state. Shared goal leads to Review role. Review role leads to Shared state. More agents add coordination work: ownership, shared state, budgets, and review.Shared goalAllocate roles and budgetsResearch roleOwn a bounded taskExecution roleOwn a bounded taskReview roleOwn a bounded taskShared stateTrack handoffs and progressCoordinate roles around a shared goal.Shared goal leads to Research role. Research role leads to Shared state. Shared goal leads to Execution role. Execution role leads to Shared state. Shared goal leads to Review role. Review role leads to Shared state. More agents add coordination work: ownership, shared state, budgets, and review.Shared goalAllocate roles and budgetsResearch roleOwn a bounded taskExecution roleOwn a bounded taskReview roleOwn a bounded taskShared stateTrack handoffs and progress
Read the connections in words
  • Shared goal → Research role: Own a bounded task.
  • Research role → Shared state: Track handoffs and progress.
  • Shared goal → Execution role: Own a bounded task.
  • Execution role → Shared state: Track handoffs and progress.
  • Shared goal → Review role: Own a bounded task.
  • Review role → Shared state: Track handoffs and progress.
Key idea

More agents add coordination work: ownership, shared state, budgets, and review.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Organizations of agents: see it in practice.

Coordinating many agents with distinct responsibilities, shared resources, and organizational constraints.

What you’ll walk through

Follow several groups coordinating related work and shared resources. Inspect how responsibilities, dependencies, and escalation affect the result when one group falls behind.

The task in this version

Coordinate a simulated launch across documentation, support, and review teams.

What you’ll learn to check

Dependency board, ownership, conflicting updates, shared budget, escalation, and an evidence-based launch readiness report.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Coordinate a simulated launch across documentation, support, and review teams.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Shared fact: feature X delayed. Each team has an owner and budget.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

More agents increase coordination demands and can duplicate mistakes. Shared objectives and ownership need to be explicit.

1 / 6

Apply this to your project

Describe your task to your own model and use Organizations of agents as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

An organization of agents is several standing agents with distinct roles and a shared goal, not one agent calling another for one task and getting an answer back. This page sits at level 7, not level 6, because the roster itself (who exists, what they are working on, when a new round of work starts) keeps running and gets decided along the way, rather than being fixed by a person for one job and torn down after.

This is mostly frontier. What ships today is closer to orchestrator-workers and agent graphs wearing a bigger name than the standing, many-role organizations described below and in the research this page cites. Keep that distinction in mind reading it: a documented, shipping framework assigns roles and runs them through fixed procedures; a research paper reports what a simulation of many agents actually did under study conditions; anything past those two is this page’s own reasoning about where the approach runs into trouble, and it says so.

This page is sourced, not measured: what these many-agent systems do comes from their makers’ and researchers’ own papers, and none has been run and scored here. It is illustrated.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Organizations of agents

A coordinator assigns a shared task board; code enforces the budget and stops a double claim.

Level 7 · Always-on agents
Task boardTask boardMODELproposes assignmentsproposesassignmentschecks board + budgetchecks board+ budgetClaimedClaimedRefused: not openRefused: not open
0of 1 step so far chosen by the model

Both proposals come from one coordinator call per round, so the recorded trace counts one model-decided step per round where the diagram plays two: one model output selected both transitions.

your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

Three tasks are open; the coordinator is asked to assign this round

T1 look up a warranty term · T2 draft a bulletin summary
T3 check the draft's citations
budget this round: researcher 2, writer 2, reviewer 2
0 tokens · 0 ms

Practical guidance

There is nothing here to sign up for. A standing organization of agents is a framework a developer assembles, not a product you turn on, and this site can name no service selling one today. If what you want is an assistant that keeps working while you are away, that is always-on assistants, and it is the one shape at this level you can actually buy.

The nearest named thing is a developer’s framework. MetaGPT’s own README states the idea outright, as “Assign different roles to GPTs to form a collaborative entity for complex tasks”, and lists “product managers / architects / project managers / engineers” as the roles it includes[3]. That is a fixed roster running a fixed procedure on one task, closer to a simulated org chart than to agents that keep existing and decide what to work on next.

One measured finding does travel to any product that passes an answer through several agents before you see it, whatever the product calls that. A 2026 paper ran 500 cascades across 10 knowledge domains on three models, 1,250 responses in all, passing each answer along a chain of agents that revise it. In three-agent chains the normalized hallucination score fell (from 0.422 at the first agent to 0.272 at the last) and factual accuracy fell with it, “from 0.789 to 0.769”, which the authors call “a trade-off between hallucination suppression and factual preservation”[4]. Both movements are small, over one benchmark, at one chain length, so read it as a direction rather than a rate. What it means in a meeting: when a vendor says several agents checked the answer, ask what the chain dropped, not only what it caught.

If someone proposes running a roster like this inside your organization, two questions settle it before any demo. “How is the whole thing stopped at once?” A roster needs one switch that drains the work and refuses new claims, not a role-by-role hunt while the rest keep working. And “when a mistake reaches a customer through three roles that each added something, whose name is on it?” No maker’s page or paper this site has read answers the second one. A calibration certificate carries the name of whoever signed it for exactly this reason, and a roster of agents has no equivalent unless somebody writes one down. Get it in writing first, or do not start.

Implementation details

The example is deliberately small: three roles (researcher, writer, reviewer) share one task board, and a coordinator model decides which open task goes to which role. What each role would do with an assigned task is out of scope; a role agent would use whatever pattern in this manual fits its own job. What this example isolates is only the part specific to an organization: the shared board, and an assignment step the model does not fully control.

examples/organizations_swarms/run.py · lines 71–114
def coordinate(board: Board, model: Model, tracer: Tracer) -> list[dict]:
    """One coordination round. Returns the assignments actually made — which can be fewer than
    the coordinator asked for, since every proposed assignment is checked against the board and
    the budget before it counts.

    The budget is per round and refills here, at the start of each one. Claims are not: a task
    claimed in an earlier round is still claimed, because the board persists and the counters do
    not. Both facts have to be tested, and testing the second one needs a round that still has
    an open task to offer -- otherwise the round returns before the model is ever called and the
    test passes without checking anything.
    """
    board.assigned_count = {r: 0 for r in ROLES}  # the budget is per round, so it starts full
    if not any(t.status == "open" for t in board.tasks):
        return []

    messages = [Message(role="system", content=SYSTEM), Message(role="user", content=_digest(board))]
    completion = model.complete(messages, tools=[ASSIGN_TOOL], max_tokens=200)
    proposed = list(completion.tool_calls)
    desc = ", ".join(f"assign({c.arguments.get('task_id')}, {c.arguments.get('role')})" for c in proposed) or "no assignments proposed"
    tracer.record(
        kind="model", decided_by="model", title="Coordinator assigns open tasks to roles",
        detail=desc, tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )

    made: list[dict] = []
    for call in proposed:
        task_id = str(call.arguments.get("task_id", ""))
        role = str(call.arguments.get("role", ""))
        task = board.get(task_id)
        if task is None or task.status != "open":
            tracer.record(kind="code", decided_by="code", title="Refuse: task is not open", detail=f"{task_id} (already claimed, or does not exist)")
            continue
        if role not in ROLES:
            tracer.record(kind="code", decided_by="code", title="Refuse: no such role", detail=f"{role} for {task_id} (roles are fixed in code, not named by the coordinator)")
            continue
        if board.assigned_count[role] >= BUDGET_PER_ROLE:
            tracer.record(kind="code", decided_by="code", title="Refuse: role is over budget", detail=f"{role} for {task_id}")
            continue
        task.status = "claimed"
        task.assigned_to = role
        board.assigned_count[role] += 1
        tracer.record(kind="code", decided_by="code", title="Claim the task for the role", detail=f"{task_id} -> {role}")
        made.append({"task": task_id, "role": role})
    return made

One coordinator call runs per round, and that single call is the round’s only decided_by: "model" step however many assignments it proposes; the diagram above draws each proposal as its own dashed edge and says so under its tally.

BUDGET_PER_ROLE is a fixed cap the coordinator cannot raise by asking; a role that has already been assigned two tasks this round gets no more, whatever the coordinator’s output requests next. Claiming is checked against the board itself, not against what the coordinator believes is true: task.status != "open" refuses an assignment outright, so if the coordinator’s own single call proposes the same task to two different roles (nothing stops a model from repeating itself), only the first proposal is ever honored. The second reads the board fresh, sees the task is no longer open, and is refused, not silently overwritten and not queued to run later. That check is what the shared board is actually for: a coordinator’s plan is a proposal against it, not an instruction the board has to accept.

MetaGPT’s own paper describes a fuller version of the same idea, fixed procedure rather than a per-round budget: it “utilizes an assembly line paradigm to assign diverse roles to various agents, efficiently breaking down complex tasks into subtasks involving many agents working together” and encodes “Standardized Operating Procedures (SOPs) into prompt sequences for more streamlined workflows”[2]. The paper’s claim for that design is scoped: “On collaborative software engineering benchmarks, MetaGPT generates more coherent solutions than previous chat-based multi-agent systems”[2]. That is the authors’ result on those benchmarks, against the systems they chose to compare with, and not a result this site has measured.

A different line of research reaches for coordination without any fixed roster at all. Park et al.’s Generative Agents paper populated a sandbox (twenty-five agents in a small town, no shared board and no fixed roles, just memory, planning and reflection) and reports that “starting with only a single user-specified notion that one agent wants to throw a Valentine’s Day party, the agents autonomously spread invitations to the party over the next two days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up for the party together at the right time”[1]. The paper’s subject is believable human behavior in a simulation, not work getting done: read it for the mechanism, which is that coordination can emerge without a coordinator role, and not as evidence that a roster of agents will organize itself around a real task.

The two rules look alike and behave differently, which is worth testing separately: the budget is spent per round and refills at the start of the next one, while a claim is permanent until the task is done. tests/test_example_organizations_swarms.py scripts a coordinator that over-assigns one role, one that proposes the same task twice in a single call, and one that tries to re-claim a task in a later round: that last test keeps a second task open on purpose, because a round with nothing open returns before the model is ever called and would otherwise pass without checking anything. Run it yourself:

examples/organizations_swarms/README.md · lines 12–12
python -m examples.organizations_swarms --model stub:scripted
When you do not need this

Try orchestrator-workers or agent graphs first for almost anything you are actually building today. A fixed roster assembled for one job, run once, and torn down afterward is level 6, cheaper to reason about, and is what the one shipping framework this page names actually does.

Move up to a standing organization only once the roster itself needs to persist and change between tasks (new roles added, old ones retired, work assigned across many separate requests rather than one) and once you have an answer, in writing, for who is accountable when it acts.

Failure modes

A task is claimed by two roles at once

How to notice it
Two roles both believe they own the same task and either duplicate the work or step on each other's output.
How to test for it
Script a coordinator that proposes the same task to two roles in one call (this page's own test suite does exactly this) and confirm only the first proposal is honored, not that both fail, not that both succeed.

A role quietly exceeds its budget

How to notice it
One role ends up doing most of the work for a round because nothing capped how much it could take on, defeating the point of having separate roles at all.
How to test for it
Script a coordinator that tries to assign every open task to the same role and confirm the count assigned never exceeds the fixed budget, regardless of how many the coordinator asked for.

An error compounds instead of getting caught

How to notice it
A mistake made by one role passes through a second and third role that were supposed to check it, and comes out the other end looking more confident than it started, not less.
How to test for it
Trace one output back through every role that touched it and check whether each one actually verified something or just passed the previous role's claim along unchanged.

Nobody can say which role is responsible for a bad result

How to notice it
A wrong or harmful output reaches a person and the postmortem cannot identify which role introduced the error, because the board and the logs record what was assigned, not what each role actually checked before acting.
How to test for it
Pick a finished task and try to reconstruct, from the logs alone, which role's decision the final output actually depended on. If you cannot, the logging is not enough for this level, whatever it is enough for at level 6.

The roster grows because adding a role feels free

How to notice it
Each new role seems to add a capability, but the system as a whole gets slower and harder to predict, and nobody can say what the fourth or fifth role actually improved.
How to test for it
Remove one role and rerun the same tasks. If the outcome does not measurably change, that role was not paying for its share of the coordination cost.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls, one coordination round
3Roles sharing the board
~180Tokens in, one round
~0.5sWall time, one round
Compared with a single agent (level 5) doing all three roles’ work itselfThe coordination call itself is cheap: one small model call per round. What this strip does not show, and what a real system pays for, is each role’s own work once assigned: three separate agents’ worth of calls this example does not run, on top of the coordination shown here.

How to Evaluate It

This example coordinates a task board; it does not answer questions about the document corpus, so the site’s shared 60-question set does not apply, the same reason it does not apply to always-on assistants’ policy example. What would be measured here: the share of coordination rounds that violate the budget or double-claim a task under adversarial scripting (this should be exactly zero, and tests/test_example_organizations_swarms.py checks it on the stub every run), how evenly work actually spreads across roles against how evenly the coordinator intended it to, and (the harder measurement research has not settled) whether a final output’s accuracy degrades as it passes through more roles, the way the hallucination-cascade paper measures for its own three-agent chain[4].

Run it

What to monitor

Assignments refused for budget or for a task no longer being open: both should be rare relative to the volume of tasks, and a rising rate of either is worth reading as a sign the roster or the budgets no longer match the workload, not just logging and moving on.

Cost at volume

One coordination call per round regardless of roster size, plus each role's own work once assigned; coordination cost itself stays close to flat as the roster grows, but the total system cost does not, since every added role is added work somewhere downstream of this example.

How it fails in production

A role silently falls behind its budget every round because the coordinator keeps proposing more than it can take, and nobody notices until a backlog of unclaimed tasks is large enough to see on a dashboard nobody was watching.

What to log

Every proposed assignment, whether the board accepted or refused it and why, which role actually touched a task before it was marked done, and the full chain of roles a given output passed through, so an accountability question has an answer that does not require asking any of the agents.

Try it

  1. Use it

    Read MetaGPT's own documentation for the roles it assigns to a task. For one of those roles, write down what you would need to see in a log to trust that role's output without re-doing its work yourself.

  2. Build it

    Run python -m examples.organizations_swarms --model stub:scripted from the repo root. The coordinator proposes five assignments in one call; the board claims four and refuses the fifth, which hands T1 to a second role after an earlier proposal took it.

  3. Either lane

    Cause the over-budget failure on purpose: lower BUDGET_PER_ROLE in examples/organizations_swarms/run.py to 1 and run the same command. The researcher's second task is refused for budget, and T4 is still open at the end.

How it connects

Before, after and instead of this

Optional: products, tools, and models

1 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Coordinate several specialist roles

Research, execution, and review agents share a task board with clear ownership and per-role limits.

Out there

Named products, tools and models

Tools1
  • MetaGPTopen source · multi-agent framework

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Generative Agents: Interactive Simulacra of Human Behavior · arXiv (Stanford University, Google Research), 04/07/2023 (accessed 09/19/2026)
  2. MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework · arXiv, 08/01/2023 (accessed 09/19/2026)
  3. FoundationAgents/MetaGPT · MetaGPT (GitHub README) (accessed 09/19/2026)
  4. Hallucination Cascade: Analyzing Error Propagation in Multi-Agent LLM Systems · arXiv, 06/06/2026 (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page