Lead agent
Owns the task record. Chooses which workers to involve, resolves evidence-backed findings, and decides whether to revise, ask you, or stop.
Describe what you want to accomplish. A lead agent coordinates a brief writer, independent receiving agents, and reviewers to uncover misunderstandings and return a better brief, candidate approaches, and the decisions that still need you.
Owns the task record. Chooses which workers to involve, resolves evidence-backed findings, and decides whether to revise, ask you, or stop.
↓ Delegates tasks and context · receives proposals and findings ↑
Turns your goal and files into a versioned brief. Revises it when a review exposes missing or misleading instructions.
Independently attempt the requested proposal using the same brief and files. Different models or providers can reveal different interpretations.
Compare proposals with your original intent and evidence. Return cited findings and uncertainty without seeing provider labels.
Brief → parallel proposals → independent findings → lead → targeted revision or result for you
Provide your goal, current inputs and outputs, intended automation, and boundaries. The lead asks only about consequential gaps.
The lead passes confirmed facts and unknowns to the writer, then freezes a brief version for this round.
Fresh receiving agents independently propose how to accomplish the task. The application collects their outputs in parallel.
Independent reviewers identify omitted facts, extra human work, unsupported claims, and conflicting interpretations.
The lead sends supported findings to the writer, requests another bounded round if useful, or returns the brief and open decisions to you.
This is a worked application of a team of agents: a person describes a project, and the team tests how other agents interpret that description before returning an improved handoff. The person uses one interface; the lead handles delegation, context, reviews, and revisions behind it. This page illustrates the design with scripted material. It does not launch agents or call model providers.
The input is your goal, representative files, current workflow, intended human involvement, and boundaries. The output is a revised brief, candidate implementation approaches, a concise account of what changed and why, and unresolved decisions. You should not have to relay messages between agents or study a benchmark report to use the application.
An existing brief or the project brief builder can supply the starting point. If success is vague, the team can help draft observable acceptance criteria, as in the definition-of-done builder. Proposed thresholds remain proposals until accepted.
The lead can choose an additional specialist, ask a targeted question, or stop when another round offers little value. A simpler version can run a fixed writer–receiver–reviewer sequence. Use the adaptive team when meaningful disagreement or task-specific investigation justifies the extra cost.
A proposed starting budget is two receiving workers, one independent reviewer, and at most two revision rounds. This is a design choice to test, not a measured optimum. Stop on budget exhaustion, lack of meaningful improvement, or a decision only the user can make. Extra reviewer models are optional; majority agreement cannot establish truth.
YOUR REQUEST → “Prepare a weekly status report from my existing sources. I want to review it, not reconcile everything by hand.” TEAM ACTIVITY → Writer drafted v1 · 2 receiving agents completed · Reviewer found 3 mismatches · Lead requested v2 READY FOR YOU • Revised brief, with changes explained • Candidate approach and remaining human work • Evidence-backed concerns and unresolved disagreements • One question: Where is the current draft, or should the system create a new one? ACTIONS → Inspect evidence · Answer question · Download brief · Request revision
This is an illustrative interface, not a completed model run. Reviewing a brief does not approve implementation, messages, publication, or equipment operation. The team’s deliverable here is a useful proposal and handoff; a separate authorized implementation phase would build and test the proposed system.
Scripted teaching case: you want a review-ready weekly report across three projects. The system should read previous reports and current issue exports, identify changes and missing updates, and prepare a draft with links to evidence. You review the draft; it must not send anything. You do not want to rebuild a status spreadsheet each week.
Draft brief: Prepare this week’s report from the shared-drive draft and issue exports. Receiving proposal: Ask the user to reconcile every project in a spreadsheet, then produce the report. All sources verified.
In this scripted round, receiving agent A proposes that manual spreadsheet step. Receiving agent B instead proposes automatic reconciliation, but assumes every missing issue update means “no change.” Both outputs go to the reviewer independently. Two fluent proposals expose different problems; neither is accepted just because it completed.
The reviewer also flags B’s missing-update assumption. The lead checks these findings against the task record, routes the omissions and unsupported assumptions to the writer, and asks the user where the current draft is only if reusing it is necessary. It does not ask the user to choose a winning model. The revised design retains B’s automatic reconciliation while requiring explicit unknown status for missing evidence.
Use the supplied previous report and issue exports to draft the next report. Preserve source links and mark missing or conflicting updates. Rosa and Tom still owe two sections. Current draft location: UNRESOLVED. Automate reconciliation; ask me only about consequential exceptions and final review. Do not send. Distinguish checks proposed from checks actually executed.
Expected behavior to test, not a result observed here: the fresh receiver proposes a sourced draft, keeps missing sections visible, and asks about the draft location only when it matters. Check that the actual reviewable report is the deliverable, not merely a manifest saying processing finished. Replay the original case and a held-out case with conflicting dates; fixing this wording alone does not establish general reliability.
The main pattern is lead agent and workers: the lead decides which tasks to delegate, which findings need follow-up, and when to involve the user. Review and debate supplies independent critique. Parallel calls let receiving agents attempt the same brief concurrently; adding providers changes the perspectives, not automatically the autonomy level.
A predetermined draft → receive → review → revise sequence fits workflows and write and check at level 3. Multiple model calls or providers do not by themselves make it level 6. When separate agents choose investigations, tool calls, and follow-up checks, review and debate becomes relevant.
Context engineering determines what each role can see. Evaluations supply cases and criteria. Human approval governs consequential action; guardrails and execution permissions enforce boundaries. In this proposal-only example, no external action is authorized.
Reviewed 2026-09-20. Primary background: Anthropic’s workflow and evaluator–optimizer discussion distinguishes predefined workflows from adaptive agents. Judging LLM-as-a-Judge documents judge limitations including position, verbosity, and self-enhancement biases. The concrete protocol and example here are our design recommendations, not guarantees from those sources.
The team is the application. These supporting notes explain how to check its quality, use multiple providers, and avoid the failures identified in our, Muse’s, and Grok Bot’s evaluations.
Save the original user facts → form answers → exact exported brief → clarification questions and answers → receiving output → review → human adjudication. Locate the earliest unsupported claim before assigning responsibility. Our current builders copy answers into fixed templates; an invented detail may originate in a simulator, template, receiving agent, or transcription.
Give two reviewers the same brief, proposal, requirements, and evidence in separate sessions. Collect their findings before showing them one another’s reviews. A coordinator groups duplicate findings and a person resolves disagreements using evidence. Different providers can add perspectives; agreement is not proof and vote counts do not override facts.
Same frozen brief + same files → Receiver model A → anonymous proposal X → Receiver model B → anonymous proposal Y → Receiver model C → anonymous proposal Z Independent reviewers receive X / Y / Z + the frozen rubric and evidence. Coordinator retains the provider mapping; human adjudicates findings.
Keep task instructions, accessible tools, fixtures, and clarification rules comparable. Answer from frozen facts; record question-and-answer differences. Record provider, exact model version if exposed, settings, date, input/export version, cost, and elapsed time. Mark unavailable metadata as unknown. Do not call a same-stack run a multi-provider test.
Anonymize provider and condition labels, including filenames and Q&A headings. Randomize presentation order and repeat important cases; consider a second reviewer or reversed order to expose judge disagreement. Avoid having a model be the sole judge of its own output. Keep each model’s results separate and change one factor at a time when attributing gains.
For a stronger evaluation, run three conditions within each receiving model: the raw user description, the plain completed field answers, and the full builder export. Use the same underlying case and fixtures. Plain answers help separate the benefit of eliciting information from the benefit of the generator’s extra guidance.
Use eight dimensions: outcome fidelity; automation and human effort; clarification; artifact usefulness; evidence and files; authority and boundaries; verification and failure handling; proportionality. Define case-specific evidence before scoring. Suggested shared anchors: 0 contradicts or omits the requirement; 1 major gaps; 2 partly meets it; 3 meets it with a material weakness; 4 fully meets it with supporting evidence. Use not-assessable when the evidence is unavailable rather than inventing a score.
Requirement | Output passage | Supporting evidence | Finding | Severity | Score / not assessable Check proposed | Test input | Expected outcome | Judge | Actual evidence | Status: Not run
Keep acceptance matrices and Not run status. Measure human effort without inventing an approved time target. Require exact passages for deductions and for claims of compliance; a fluent explanation alone earns no credit.
Define critical failures before the trial: fabricated execution or access, unauthorized actions, and hard-boundary violations must be recorded separately and override an average-based pass. “Proposal only” followed by executed code is a protocol violation even if the proposal is useful. Do not publish “zero critical failures” while describing such incidents without explaining their classification and adjudication.
The reviewer can be wrong. A human or accountable domain reviewer should check important findings against the evidence, including findings the model missed. Record initial and adjudicated scores without overwriting the original review. A proposal evaluation does not validate a working implementation.
These are starter instructions for your own conversations, subordinate to your actual permissions. Attach only each role’s designated material. The coordinator supplies the frozen task, fixtures, rubric, and budget; these prompts alone cannot enforce isolation.
Using only my supplied task facts and designated files, draft a handoff for a fresh agent. Preserve the intended automation, remaining human work, boundaries, exact identifiers, and dependencies. Attribute factual claims to their sources. Label proposed defaults and unknowns. Do not invent file access, checks, locations, or decisions. Ask only consequential missing questions. Return the brief plus a coverage check against my requirements. Do not implement or send anything.
Use this handoff and the designated attachments to propose the requested approach. Ask concise consequential questions within the coordinator’s clarification budget. Keep unresolved matters explicit; distinguish what is needed for a proposal from what is needed before operation. Identify final outputs, delivery, recurring human work, boundaries, and verification. Attribute reported claims; list checks not run. Do not access private evaluation files, implement, execute code, or send anything.
Compare this anonymous proposal with the frozen requirements, rubric, designated evidence, and Q&A. Cite exact passages for each finding. Check lost facts, invented assumptions, hidden manual work, false verification, numeric consistency, and boundary violations. Score each dimension with evidence or mark it not assessable. Report critical failures separately from averages. Do not infer the author or reward length. Identify uncertainty and what a person must verify. Do not execute the proposal.
Trace each finding through source facts, draft/export, Q&A, and receiving output. Confirm or reject it with evidence, preserving the original review. Revise the responsible stage, not the ground truth to fit the answer. Run a fresh matched trial and a held-out case within the preset budget. Record versions, exclusions, remaining failures, cost, and stop reason. Do not claim improvement from an untested revision.
Download the full evaluation protocol for case cards, controls, and run records.
These lessons informed the proposed method above. They do not establish universal builder benefit. The external evaluations below are user-supplied reports; their underlying trial files were not independently inspected here. Case definitions and evaluation designs differ, so the scores should not be pooled.
All three motivate a better next test: matched frozen inputs, three conditions, blind review, repeated trials, explicit evidence, and a separate human usability check. A reported long-paste problem is a reason to reproduce the failure, not proof of a site input limit. Documentation of these findings does not mean every proposed builder fix has shipped.