Level 03 · Workflows

Parallel calls

Running several prompts at once and combining the results.

Sourced

Concept at a glance

Work on independent pieces at the same time.

Parallel branchesConceptual illustration
Work on independent pieces at the same time.Split the work leads to Prompt A. Prompt A leads to Combine. Split the work leads to Prompt B. Prompt B leads to Combine. Split the work leads to Prompt C. Prompt C leads to Combine. Code starts the branches and combines their results; one branch does not direct another.Split the workA fixed set of branchesPrompt AOne independent piecePrompt BAnother independent piecePrompt CAnother independent pieceCombineCollect the useful resultsWork on independent pieces at the same time.Split the work leads to Prompt A. Prompt A leads to Combine. Split the work leads to Prompt B. Prompt B leads to Combine. Split the work leads to Prompt C. Prompt C leads to Combine. Code starts the branches and combines their results; one branch does not direct another.Split the workA fixed set of branchesPrompt AOne independent piecePrompt BAnother independent piecePrompt CAnother independent pieceCombineCollect the useful results
Read the connections in words
  • Split the work → Prompt A: One independent piece.
  • Prompt A → Combine: Collect the useful results.
  • Split the work → Prompt B: Another independent piece.
  • Prompt B → Combine: Collect the useful results.
  • Split the work → Prompt C: Another independent piece.
  • Prompt C → Combine: Collect the useful results.
Key idea

Code starts the branches and combines their results; one branch does not direct another.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Parallel calls: see it in practice.

Running independent tasks concurrently and combining their outputs.

What you’ll walk through

Follow independent pieces of a task running alongside one another and then being combined. Watch how a missing branch or inconsistent time window affects the final result.

The task in this version

Collect current updates for Atlas, Beacon, and Cedar in parallel.

What you’ll learn to check

Per-project evidence cards, timestamps, conflicting-source flags, and a combined report with no invented update for silent projects.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Collect current updates for Atlas, Beacon, and Cedar in parallel.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Atlas tracker updated Thursday; Beacon notes Friday; Cedar has no current-week source.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

The branches must be sufficiently independent, and their results must refer to compatible versions or periods.

1 / 6

Apply this to your project

Describe your task to your own model and use Parallel calls as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Parallel calls run more than one model call at the same time instead of one after another, and combine the results in code. Anthropic describes two variations[1]. Sectioning breaks one task into independent subtasks that each run in parallel: one call screens a request for policy violations while another handles it, or several calls each evaluate a different aspect of the same output. Voting runs the same task several times and combines the outputs: several prompts review the same code for vulnerabilities, or several prompts judge the same content against different thresholds and the results are combined into one verdict.

Both stay at level 3 as long as your code decides how many calls to make, what each one gets, and how to combine what comes back, before any call goes out. The model fills in each call’s answer; it does not decide how many branches exist or how they are merged. The line runs through those two decisions: a system where the model reads the task and works out what the subtasks are, or reads the branches and writes the merged answer itself, is lead agent and workers at level 6, not this.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Parallel calls (sectioning)

Answer from each candidate section alone, at the same time, then combine what came back.

Level 3 · Workflows
QuestionQuestionPick sectionsPick sectionsMODELcare-guide#1care-guide#1MODELdw480-manual#3dw480-manual#3MODELdw480-manual#6dw480-manual#6CombineCombineAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 08Your code chose

The question arrives

"What is the DW-480's Normal cycle water use, and how
often should its filter be cleaned?"
0 tokens · 0 ms

Practical guidance

Run several steps of a chain at once instead of one after another, in a tool whose canvas lets you branch into parallel paths and merge them back. Split a task into independent branches only when each branch’s step genuinely does not need another branch’s answer to run: screening a request for policy problems while a separate step drafts a reply to it is independent work; drafting a reply and then checking that same reply is not, because the check needs the draft first. The benefit you are paying for is speed, not a better answer: running three calls at once finishes in about the time of the slowest one instead of the sum of all three, which is worth confirming actually happened before you trust the setup.

Coding agents show this on the canvas directly. Replit, announcing the fourth version of Replit Agent, says “Independent tasks can run in parallel, with progress visible and coordinated,” and for larger jobs it “can split a single task into smaller pieces, work on them simultaneously with sub-agents, and recombine the results”[2]. Claude Code’s subagents feature documents the same pattern for research: “For independent investigations, spawn multiple subagents to work simultaneously”. It says “Each subagent explores its area independently, then Claude synthesizes the findings”, adding that “This works best when the research paths don’t depend on each other”[3].

Open each branch after a parallel run finishes and read what it actually saw and returned, the same way you would check one step of a plain chain: a tool that only shows the merged final answer is hiding exactly the place two branches disagreed or repeated each other. Watch for repetition specifically. Two branches that never saw each other’s work can both answer the same sub-question, so the combined result states one fact twice with nothing that noticed the overlap.

Reach for several independent opinions on the same question, rather than several different sub-tasks, only where a wrong answer costs more than the extra run: a security review, a policy call, a number somebody is about to act on. For anything routine, one pass is enough, and running several is just several times the cost for no benefit anyone will notice.

Implementation details

The example is sectioning: it retrieves a fixed set of candidate document sections, asks the model to answer from each one alone (never seeing the other sections or the other calls) and combines whichever sections actually answered part of the question. Because each call already knows which single section it saw, citations are exact by construction; nothing has to be parsed back out of free text the way RAG and prompt chaining do.

ThreadPoolExecutor.map is what makes this parallel rather than sequential: it submits every call to the pool at once, and the calls run concurrently, but the returned iterator still yields results in the order the candidates were given, regardless of which call actually finishes first. That is what keeps the example deterministic without an explicit sort: order comes from retrieval, never from a race between threads.

Every tracer.record call happens on the main thread, after list(pool.map(...)) has already collected every result: the trace itself is never written to from more than one thread at once. A StubModel built from a fixed list of canned responses is not safe to call from several threads concurrently, since it advances a shared counter with no lock; the example’s tests build their stub from a function that reads the prompt instead, which has no shared state to race on.

The trace above shows a real cost of naive sectioning: two of the three candidate sections happened to answer the same sub-question, so the combined text states the filter fact twice. Nothing in a fixed combine step notices the overlap, because each section answered without seeing what the others said.

examples/parallelization/run.py · lines 22–75
LEVEL = 3
CANDIDATES_K = 3
NO_ANSWER = "NOT IN THIS SECTION"
PER_SECTION_SYSTEM = (
    "You are given exactly one source passage about Halvorsen appliances, and a question that "
    "may have more than one part. If this passage answers all or part of the question, answer "
    "briefly using only this passage. If it answers none of the question, reply with exactly "
    f"'{NO_ANSWER}' and nothing else."
)


def _answer_from_one_section(question: str, section: Section, model: Model) -> Completion:
    prompt = f"Passage [{section.cite}] {section.title}:\n{section.text}\n\nQuestion: {question}"
    return model.complete([Message(role="system", content=PER_SECTION_SYSTEM), Message(role="user", content=prompt)], max_tokens=200)


def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    k: int = CANDIDATES_K,
) -> Answer:
    del embedder  # candidates come from keyword search, not a vector index
    sections = load_sections(corpus_dir)
    candidates = [s for s, score in bm25_search(sections, question, k=k) if score > 0]
    tracer.record(
        kind="code", decided_by="code", title="Pick sections to answer in parallel",
        detail=", ".join(s.cite for s in candidates) or "none",
    )

    # .map submits every call to the pool at once and yields results back in candidate order,
    # so the calls run concurrently but the code below never has to sort them: determinism comes
    # from retrieval order, not from whichever call happens to finish first.
    with ThreadPoolExecutor(max_workers=max(1, len(candidates))) as pool:
        completions = list(pool.map(lambda s: _answer_from_one_section(question, s, model), candidates))

    for section, completion in zip(candidates, completions):
        tracer.record(
            kind="model", decided_by="code", title=f"Answer from {section.cite} alone", detail=completion.text[:200],
            tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
        )

    used = [(s, c) for s, c in zip(candidates, completions) if NO_ANSWER not in c.text.upper()]
    tracer.record(
        kind="code", decided_by="code", title="Combine the sections that answered",
        detail=f"{len(used)} of {len(candidates)} sections answered part of the question",
    )
    if not used:
        return Answer(text="None of the retrieved sections answered the question.", citations=[], retrieved_sources=[s.cite for s in candidates])
    combined = " ".join(c.text.strip() for _, c in used)
    return Answer(text=combined, citations=sorted({s.cite for s, _ in used}), retrieved_sources=[s.cite for s in candidates])

Run it yourself:

examples/parallelization/README.md · lines 16–16
python -m examples.parallelization --model stub:scripted

Real-time parallel calls like these are for when the answer is needed now. When it is not (a nightly re-score of every open ticket, a one-time pass over a large document set), the batch APIs three model makers publish do the same many-calls-one-submission idea asynchronously and cheaper: Anthropic describes its Message Batches API as suited to tasks that do not need an immediate response, “with most batches finishing in less than 1 hour while reducing costs by 50% and increasing throughput”[4]; OpenAI’s Batch API gives a “50% cost discount compared to synchronous APIs” with each batch completing “within 24 hours (and often more quickly)”[5]; Google states that its Gemini Batch API processes requests at “50% of the standard cost” and says of the wait: “The target turnaround time is 24 hours, but in majority of cases, it is much quicker”[6]. Each of those is the maker’s own published figure, checked on the date in the source list below, and each is the same trade: give up the immediate response, halve the price.

When you do not need this

Try prompt chaining or a single call first if the task’s parts actually depend on each other: a later part needs an earlier part’s answer, or the sections would overlap and need to be reconciled against each other. Running dependent work in parallel does not make it independent; it just hides the dependency until the combine step produces a contradiction.

Move up to parallelization once the task genuinely splits into parts that do not need each other’s answers, and the parts are already known before any call runs: a section list you retrieved, a fixed set of checks to run, a fixed number of independent opinions to gather.

Failure modes

Overlapping sections restate the same fact

How to notice it
The combined answer repeats itself, or states the same fact in two slightly different ways, because two sections happened to cover the same ground and neither call could see the other's answer.
How to test for it
Retrieve candidates for a question you know has redundant coverage across sections (the DW-480's filter is described in both its own manual and the shared care-and-cleaning guide) and check whether the combined text repeats the fact.

A fixed combine step cannot resolve a disagreement

How to notice it
Two sections answer the same question differently (an old figure and a superseding one) and a plain concatenation states both without saying which is current, because nothing in the combine step compares them against each other.
How to test for it
Run a question over sections you know conflict (an original spec and a later correction) and check whether the combined answer states both values with no indication of which one is authoritative.

A shared, mutable stub races under real concurrency

How to notice it
A test or a manual run using a list-based StubModel raises IndexError or returns answers in the wrong order under a thread pool, because the stub's internal counter is not safe to advance from more than one thread.
How to test for it
Run the example's own test suite; it is deliberately built on a callable-based stub for exactly this reason, and a regression toward a list-based stub under the thread pool would surface as an intermittent failure, not a consistent one.

Rate limits under real load

How to notice it
Firing many calls at once against a live API returns 429 rate-limit errors once concurrency crosses the provider’s per-minute limit, which a small stub run never exercises.
How to test for it
Check the provider’s published rate limits against the number of parallel calls one request triggers, before running the example against a live model at any real question volume.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

3Model calls, one question
~520Tokens in (summed)
~47Tokens out (summed)
~0.65sWall time
Compared with prompt chaining (level 3, sequential)Cost sums across the three calls, the same as a sequential chain would. Wall time does not: it tracks the slowest single call, not their sum, which is the entire latency argument for running independent work in parallel instead of one call after another.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

The same 60 questions, the same corpus, plus one number of its own: what share of a question’s must_cite sections were actually covered by some section’s answer. Multi-hop questions test that directly, since they need more than one section’s fact combined into a single answer.

Conflicting-source questions are the interesting case to watch: sectioning retrieves both sides of a deliberate contradiction as readily as RAG does, but its fixed combine step has no way to compare them, only to concatenate whatever each section said. Whether that scores better or worse than RAG’s single stuffed-context prompt is exactly the kind of question this site can only answer once a result file exists (see docs/EVALS.md); none does yet. Run python scripts/eval_run.py --example parallelization --model <spec> --dry to project the cost of a real run first.

Run it

What to monitor

Per-branch latency and error rate, not just the overall run's. One slow or failing branch in a thread pool can dominate wall time even though the others finished quickly; averaging across branches hides exactly the branch worth investigating.

Cost at volume

Cost is the number of parallel calls times the number of questions, same as a sequential chain of the same length: parallelism buys latency, not a lower bill. For volume that does not need an immediate answer, a maker's batch API halves the per-call cost in exchange for asynchronous delivery.

How it fails in production

Concurrency crosses a provider's per-minute rate limit once real question volume arrives, producing errors a low-volume stub or manual test never triggers. Separately, a thread pool sized for a fixed number of sections silently under-uses itself if fewer candidates come back than expected, or queues up if more do.

What to log

Each branch's input, output, token counts and wall time individually, plus which branches were kept versus dropped by the combine step, so a bad or missing final answer traces back to one specific branch rather than to 'the parallel step' as a whole.

Try it

  1. Use it

    Ask a coding agent that advertises parallel subagents to work on three independent parts. Does it run them at once, and does the result repeat itself?

  2. Build it

    Run python -m examples.parallelization --model stub:scripted from the repo root. Three branches answer from one section each, one says NOT IN THIS SECTION, and the combine keeps the two that did. Now change CANDIDATES_K from 3 to 5 in examples/parallelization/run.py: the run stops with ScriptExhausted, naming the passage no reply matches.

  3. Either lane

    Cause the overlap failure on purpose: find a question evals/corpus/ answers in two places, run it, and count how often the combined answer repeats itself.

How it connects

Before, after and instead of this

Read first

Move up when

Decoded in

Optional: products, tools, and models

8 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 2 more examples
  • Message Batches API Tool or framework · Anthropic

    Asynchronous batch processing api

    Checked 09/18/2026
  • Jev Model · TypeSafe AI

    System one decision model

    Checked 09/18/2026
In practice

Review independent document sections

Run a fixed set of prompts at the same time, one per section, and combine their findings.

Out there

Named products, tools and models

Products3
  • Claude Code subagentsAnthropic · multi-agent feature of a coding agent
  • Grok HeavySpaceXAI · several agents answering one question
  • Replit AgentReplit · coding agent
Tools4
  • Batch APIOpenAI · asynchronous batch processing api
  • Batch APIGoogle · asynchronous batch processing api · formerly Batch Mode
  • LangChainLangChain · application framework
  • Message Batches APIAnthropic · asynchronous batch processing api
Models1
  • JevTypeSafe AI · system one decision model

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Building effective agents · Anthropic, 12/19/2024 (accessed 09/19/2026)
  2. Introducing Replit Agent 4: Built for Creativity · Replit (accessed 09/19/2026)
  3. Subagents · Anthropic (Claude Code documentation) (accessed 09/19/2026)
  4. Batch processing · Anthropic (accessed 09/19/2026)
  5. Batch API · OpenAI (accessed 09/19/2026)
  6. Batch API · Google (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page