Teams of Agents · application example

A team of agents that improves your project brief

Describe what you want to accomplish. A lead agent coordinates a brief writer, independent receiving agents, and reviewers to uncover misunderstandings and return a better brief, candidate approaches, and the decisions that still need you.

The application’s team
You: describe the goal · answer consequential questions · review the result

Lead agent

Owns the task record. Chooses which workers to involve, resolves evidence-backed findings, and decides whether to revise, ask you, or stop.

↓ Delegates tasks and context · receives proposals and findings ↑

Brief writer

Turns your goal and files into a versioned brief. Revises it when a review exposes missing or misleading instructions.

Receiving agents · parallel

Independently attempt the requested proposal using the same brief and files. Different models or providers can reveal different interpretations.

Reviewer agents · independent

Compare proposals with your original intent and evidence. Return cited findings and uncertainty without seeing provider labels.

Brief → parallel proposals → independent findings → lead → targeted revision or result for you

How one round unfolds
  1. 01

    You describe the outcome

    Provide your goal, current inputs and outputs, intended automation, and boundaries. The lead asks only about consequential gaps.

  2. 02

    Writer drafts version 1

    The lead passes confirmed facts and unknowns to the writer, then freezes a brief version for this round.

  3. 03

    Workers try the brief

    Fresh receiving agents independently propose how to accomplish the task. The application collects their outputs in parallel.

  4. 04

    Reviewers find mismatches

    Independent reviewers identify omitted facts, extra human work, unsupported claims, and conflicting interpretations.

  5. 05

    Lead coordinates revision

    The lead sends supported findings to the writer, requests another bounded round if useful, or returns the brief and open decisions to you.

The application: a brief improvement team

This is a worked application of a team of agents: a person describes a project, and the team tests how other agents interpret that description before returning an improved handoff. The person uses one interface; the lead handles delegation, context, reviews, and revisions behind it. This page illustrates the design with scripted material. It does not launch agents or call model providers.

The input is your goal, representative files, current workflow, intended human involvement, and boundaries. The output is a revised brief, candidate implementation approaches, a concise account of what changed and why, and unresolved decisions. You should not have to relay messages between agents or study a benchmark report to use the application.

An existing brief or the project brief builder can supply the starting point. If success is vague, the team can help draft observable acceptance criteria, as in the definition-of-done builder. Proposed thresholds remain proposals until accepted.

The lead can choose an additional specialist, ask a targeted question, or stop when another round offers little value. A simpler version can run a fixed writer–receiver–reviewer sequence. Use the adaptive team when meaningful disagreement or task-specific investigation justifies the extra cost.

What the lead controls—and what it cannot decide

  • The application stores a versioned task record, briefs, designated attachments, proposals, findings, and decision log. Each finding names its brief version and supporting source so a late review cannot accidentally revise the wrong draft.
  • Receiving workers get the same brief version and designated files, not each other’s answers. Reviewers get the original requirements and evidence. Private scoring cards and provider mappings stay out of workers’ contexts and accessible filesystems.
  • The lead routes facts and questions between roles, consolidates duplicate findings, and requests focused revisions. It may resolve a factual dispute from available evidence; it must ask you about missing preferences or consequential tradeoffs rather than invent your intent.
  • The runtime enforces provider access, tool permissions, concurrency, cost, time, and revision limits. Prompts describe these boundaries but do not enforce them. Share only files permitted for each provider; a multi-provider design does not require sending every file everywhere.
  • On timeout or provider failure, preserve completed work and label missing contributions. Retry within the budget or return a partial result; do not claim independent review happened when no reviewer completed. Unresolved serious findings remain visible.

A proposed starting budget is two receiving workers, one independent reviewer, and at most two revision rounds. This is a design choice to test, not a measured optimum. Stop on budget exhaustion, lack of meaningful improvement, or a decision only the user can make. Extra reviewer models are optional; majority agreement cannot establish truth.

What you see in the interface

YOUR REQUEST → “Prepare a weekly status report from my existing sources. I want to review it, not reconcile everything by hand.”

TEAM ACTIVITY → Writer drafted v1 · 2 receiving agents completed · Reviewer found 3 mismatches · Lead requested v2

READY FOR YOU
• Revised brief, with changes explained
• Candidate approach and remaining human work
• Evidence-backed concerns and unresolved disagreements
• One question: Where is the current draft, or should the system create a new one?

ACTIONS → Inspect evidence · Answer question · Download brief · Request revision

This is an illustrative interface, not a completed model run. Reviewing a brief does not approve implementation, messages, publication, or equipment operation. The team’s deliverable here is a useful proposal and handoff; a separate authorized implementation phase would build and test the proposed system.

One task through the team: weekly reporting

Scripted teaching case: you want a review-ready weekly report across three projects. The system should read previous reports and current issue exports, identify changes and missing updates, and prepare a draft with links to evidence. You review the draft; it must not send anything. You do not want to rebuild a status spreadsheet each week.

  • Supplied facts: last week’s report, current issue exports, and a note that two sections depend on updates from Rosa and Tom. The current draft location is unknown.
  • Private review criteria: preserve both contributor dependencies; do not turn missing updates into green status; do not invent a draft location or a completed check; preserve review-only human effort and the no-send boundary. The receiver gets the original task facts through the tested handoff, not the private scoring card.

1. A plausible handoff goes wrong

Draft brief: Prepare this week’s report from the shared-drive draft and issue exports.
Receiving proposal: Ask the user to reconcile every project in a spreadsheet, then produce the report. All sources verified.

In this scripted round, receiving agent A proposes that manual spreadsheet step. Receiving agent B instead proposes automatic reconciliation, but assumes every missing issue update means “no change.” Both outputs go to the reviewer independently. Two fluent proposals expose different problems; neither is accepted just because it completed.

2. The reviewer explains the failure

  • Unsupported location: “shared-drive draft” is not in the source facts. Trace the problem to the draft before blaming the receiver.
  • Lost dependencies: Rosa and Tom disappeared from the handoff. The receiving agent cannot reliably recover facts it never received.
  • Extra recurring labor: the spreadsheet reconciliation contradicts the intended workflow. The system should reconcile available evidence and present exceptions.
  • False verification: “All sources verified” has no inspection record. A proposal cannot claim executed checks.

The reviewer also flags B’s missing-update assumption. The lead checks these findings against the task record, routes the omissions and unsupported assumptions to the writer, and asks the user where the current draft is only if reusing it is necessary. It does not ask the user to choose a winning model. The revised design retains B’s automatic reconciliation while requiring explicit unknown status for missing evidence.

3. Revise the handoff and try again

Use the supplied previous report and issue exports to draft the next report. Preserve source links and mark missing or conflicting updates. Rosa and Tom still owe two sections. Current draft location: UNRESOLVED. Automate reconciliation; ask me only about consequential exceptions and final review. Do not send. Distinguish checks proposed from checks actually executed.

Expected behavior to test, not a result observed here: the fresh receiver proposes a sourced draft, keeps missing sections visible, and asks about the draft location only when it matters. Check that the actual reviewable report is the deliverable, not merely a manifest saying processing finished. Replay the original case and a held-out case with conflicting dates; fixing this wording alone does not establish general reliability.

How this connects to the concepts

The main pattern is lead agent and workers: the lead decides which tasks to delegate, which findings need follow-up, and when to involve the user. Review and debate supplies independent critique. Parallel calls let receiving agents attempt the same brief concurrently; adding providers changes the perspectives, not automatically the autonomy level.

A predetermined draft → receive → review → revise sequence fits workflows and write and check at level 3. Multiple model calls or providers do not by themselves make it level 6. When separate agents choose investigations, tool calls, and follow-up checks, review and debate becomes relevant.

Context engineering determines what each role can see. Evaluations supply cases and criteria. Human approval governs consequential action; guardrails and execution permissions enforce boundaries. In this proposal-only example, no external action is authorized.

Reviewed 2026-09-20. Primary background: Anthropic’s workflow and evaluator–optimizer discussion distinguishes predefined workflows from adaptive agents. Judging LLM-as-a-Judge documents judge limitations including position, verbosity, and self-enhancement biases. The concrete protocol and example here are our design recommendations, not guarantees from those sources.

Supporting evaluation method

The team is the application. These supporting notes explain how to check its quality, use multiple providers, and avoid the failures identified in our, Muse’s, and Grok Bot’s evaluations.

Keep the evidence chain intact

Save the original user facts → form answers → exact exported brief → clarification questions and answers → receiving output → review → human adjudication. Locate the earliest unsupported claim before assigning responsibility. Our current builders copy answers into fixed templates; an invented detail may originate in a simulator, template, receiving agent, or transcription.

  • Label user-reported claims, inspected files, executed checks with results, proposed defaults, and unresolved questions separately. “The user says a check passed” is not “I ran the check.”
  • Mentioned attachments are not attached or inspected evidence. Provide actual designated fixtures, record access failures, and keep private cards outside the receiving agent’s accessible workspace. A fresh chat alone is not filesystem isolation.
  • Preserve exact paths, names, dependencies, deadlines, exceptions, and permissions. Do not silently resolve halt-versus-continue policies or missing locations.
  • Extract known facts from the whole brief before asking short, consequential questions. Do not repeatedly ask for unavailable details. Label suggested thresholds as proposals, not agreed requirements.
  • Validate objective specifications with ordinary checks: units, counts, totals, and aspect ratios. For example, 1080 × 1350 is 4:5; 1080 × 1440 is 3:4. Record whether those checks were actually run.
  • State the final output, destination, delivery check, and recurring human actions. Agent-assisted tool development and autonomous recurring operation are separate decisions.
Run different models and providers in parallel

Improve one artifact: parallel reviews

Give two reviewers the same brief, proposal, requirements, and evidence in separate sessions. Collect their findings before showing them one another’s reviews. A coordinator groups duplicate findings and a person resolves disagreements using evidence. Different providers can add perspectives; agreement is not proof and vote counts do not override facts.

Test portability: parallel receiving trials

Same frozen brief + same files
  → Receiver model A → anonymous proposal X
  → Receiver model B → anonymous proposal Y
  → Receiver model C → anonymous proposal Z

Independent reviewers receive X / Y / Z + the frozen rubric and evidence.
Coordinator retains the provider mapping; human adjudicates findings.

Keep task instructions, accessible tools, fixtures, and clarification rules comparable. Answer from frozen facts; record question-and-answer differences. Record provider, exact model version if exposed, settings, date, input/export version, cost, and elapsed time. Mark unavailable metadata as unknown. Do not call a same-stack run a multi-provider test.

Anonymize provider and condition labels, including filenames and Q&A headings. Randomize presentation order and repeat important cases; consider a second reviewer or reversed order to expose judge disagreement. Avoid having a model be the sole judge of its own output. Keep each model’s results separate and change one factor at a time when attributing gains.

Does the builder actually help?

For a stronger evaluation, run three conditions within each receiving model: the raw user description, the plain completed field answers, and the full builder export. Use the same underlying case and fixtures. Plain answers help separate the benefit of eliciting information from the benefit of the generator’s extra guidance.

  • Include thin, messy descriptions as well as complete ones. Rich briefs can create a score ceiling; report these groups separately instead of changing the rubric until a difference appears.
  • Give every condition the same clarification budget and access rules. A practical proposed cap is three rounds, but missing answers must remain unknown.
  • Freeze exact artifacts and rubric versions. Preserve original failed trials; record excluded contaminated trials and reasons. Use held-out cases after revisions to check for overfitting.
  • Repeat trials before claiming a reliable gain. Report case counts, score distributions, serious failures, reviewer disagreement, cost, and human effort. Small single-run averages are exploratory, not a provider leaderboard.
  • Test real people’s ability to understand and complete the forms separately. An agent successfully filling a form does not establish human usability.
Score evidence, not confidence or length

Use eight dimensions: outcome fidelity; automation and human effort; clarification; artifact usefulness; evidence and files; authority and boundaries; verification and failure handling; proportionality. Define case-specific evidence before scoring. Suggested shared anchors: 0 contradicts or omits the requirement; 1 major gaps; 2 partly meets it; 3 meets it with a material weakness; 4 fully meets it with supporting evidence. Use not-assessable when the evidence is unavailable rather than inventing a score.

Requirement | Output passage | Supporting evidence | Finding | Severity | Score / not assessable
Check proposed | Test input | Expected outcome | Judge | Actual evidence | Status: Not run

Keep acceptance matrices and Not run status. Measure human effort without inventing an approved time target. Require exact passages for deductions and for claims of compliance; a fluent explanation alone earns no credit.

Define critical failures before the trial: fabricated execution or access, unauthorized actions, and hard-boundary violations must be recorded separately and override an average-based pass. “Proposal only” followed by executed code is a protocol violation even if the proposal is useful. Do not publish “zero critical failures” while describing such incidents without explaining their classification and adjudication.

The reviewer can be wrong. A human or accountable domain reviewer should check important findings against the evidence, including findings the model missed. Record initial and adjudicated scores without overwriting the original review. A proposal evaluation does not validate a working implementation.

Prompts for the separate roles

These are starter instructions for your own conversations, subordinate to your actual permissions. Attach only each role’s designated material. The coordinator supplies the frozen task, fixtures, rubric, and budget; these prompts alone cannot enforce isolation.

Drafting agent

Using only my supplied task facts and designated files, draft a handoff for a fresh agent. Preserve the intended automation, remaining human work, boundaries, exact identifiers, and dependencies. Attribute factual claims to their sources. Label proposed defaults and unknowns. Do not invent file access, checks, locations, or decisions. Ask only consequential missing questions. Return the brief plus a coverage check against my requirements. Do not implement or send anything.

Receiving agent

Use this handoff and the designated attachments to propose the requested approach. Ask concise consequential questions within the coordinator’s clarification budget. Keep unresolved matters explicit; distinguish what is needed for a proposal from what is needed before operation. Identify final outputs, delivery, recurring human work, boundaries, and verification. Attribute reported claims; list checks not run. Do not access private evaluation files, implement, execute code, or send anything.

Independent reviewer

Compare this anonymous proposal with the frozen requirements, rubric, designated evidence, and Q&A. Cite exact passages for each finding. Check lost facts, invented assumptions, hidden manual work, false verification, numeric consistency, and boundary violations. Score each dimension with evidence or mark it not assessable. Report critical failures separately from averages. Do not infer the author or reward length. Identify uncertainty and what a person must verify. Do not execute the proposal.

Coordinator and human review

Trace each finding through source facts, draft/export, Q&A, and receiving output. Confirm or reject it with evidence, preserving the original review. Revise the responsible stage, not the ground truth to fit the answer. Run a fresh matched trial and a held-out case within the preset budget. Record versions, exclusions, remaining failures, cost, and stop reason. Do not claim improvement from an untested revision.

Download the full evaluation protocol for case cards, controls, and run records.

What our three evaluations contributed

These lessons informed the proposed method above. They do not establish universal builder benefit. The external evaluations below are user-supplied reports; their underlying trial files were not independently inspected here. Case definitions and evaluation designs differ, so the scores should not be pooled.

  • Our executed evaluation: exposed misleading unknown fields, altered paths, missed delivery requirements, and reviewer errors. Carry literal evidence through the handoff and retain human adjudication. Four raw-description comparisons did not establish a robust advantage.
  • Muse feedback, attributed summary: reported 13 paired cases with blind review and tied overall averages. Its detailed evidence-chain failures motivate source attribution and preserving open judgments. Its zero-critical-failure conclusion needs reconciliation with the reported false verification and protocol violations.
  • Grok Bot feedback, attributed summary: added raw-description and plain-field controls across six cases. Its clearest reported benefit was definition-of-done guidance for a vague task. Preserve acceptance matrices; do not turn a single-case signal into a claim of reliable benefit. Exact underlying model identity was unavailable.

All three motivate a better next test: matched frozen inputs, three conditions, blind review, repeated trials, explicit evidence, and a separate human usability check. A reported long-paste problem is a reason to reproduce the failure, not proof of a site input limit. Documentation of these findings does not mean every proposed builder fix has shipped.