# Parallel calls

_Level 03 · Workflows · sourced_

Running several prompts at once and combining the results.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow independent pieces of a task running alongside one another and then being combined. Watch how a missing branch or inconsistent time window affects the final result.

**Assumptions:** The branches must be sufficiently independent, and their results must refer to compatible versions or periods.

**Design choices:** Parallelize work that can be combined without hidden ordering dependencies. Balance latency benefits against source load, duplication, and aggregation effort.

**Request:** Collect current updates for Atlas, Beacon, and Cedar in parallel.

**Starting evidence:** Atlas tracker updated Thursday; Beacon notes Friday; Cedar has no current-week source.

**Action and control:** Read independent sources concurrently, then reconcile freshness before combining.

**Stage records (authored, not executed):**

### Input record

Atlas tracker updated Thursday; Beacon notes Friday; Cedar has no current-week source.

What changed: Establish the facts supplied for this version of the task.

### Design note

Parallelize work that can be combined without hidden ordering dependencies. Balance latency benefits against source load, duplication, and aggregation effort.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Read independent sources concurrently, then reconcile freshness before combining.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Atlas and Beacon have updates. Cedar: no fresh update found; current status unknown.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Per-project evidence cards, timestamps, conflicting-source flags, and a combined report with no invented update for silent projects.

If the result falls short:
Retry only the failed branch when safe. Report partial coverage or wait for required inputs according to the task, rather than silently substituting stale information.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for comparisons, checks, or collecting project updates. Define what all branches must share and whether an incomplete result is still useful.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Atlas and Beacon have updates. Cedar: no fresh update found; current status unknown.

**Change something — Beacon tracker and notes disagree:** Keep both dated claims visible and reconcile with the owner. The fastest result does not automatically win.

**Decision:** Does finishing all calls mean every project is verified?

**Answer:** No; freshness and conflicts still need review.

**Why:** Project-specific work can run concurrently, but stale sources and inconsistent milestone dates require reconciliation before aggregation.

**Review criteria:** Per-project evidence cards, timestamps, conflicting-source flags, and a combined report with no invented update for silent projects.

**Recovery:** Retry only the failed branch when safe. Report partial coverage or wait for required inputs according to the task, rather than silently substituting stale information.

**Adapt it:** Use this for comparisons, checks, or collecting project updates. Define what all branches must share and whether an incomplete result is still useful.

Parallel calls run more than one model call at the same time instead of one after another, and
combine the results in code. Anthropic describes two variations[1]. **Sectioning**
breaks one task into independent subtasks that each run in parallel: one call screens a request
for policy violations while another handles it, or several calls each evaluate a different
aspect of the same output. **Voting** runs the *same* task several times and combines the
outputs: several prompts review the same code for vulnerabilities, or several prompts judge the
same content against different thresholds and the results are combined into one verdict.

Both stay at level 3 as long as your code decides how many calls to make, what each one gets,
and how to combine what comes back, before any call goes out. The model fills in each call's
answer; it does not decide how many branches exist or how they are merged. The line runs through
those two decisions: a system where the model reads the task and works out what the subtasks are,
or reads the branches and writes the merged answer itself, is [lead agent and workers](/gradient_ascent/techniques/orchestrator-workers/) at level 6, not this.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 3 · Parallel calls. The same steps are described in the sections below._

## Practical guidance

Run several steps of a chain at once instead of one after another, in a tool whose canvas lets
you branch into parallel paths and merge them back. Split a task into independent branches only
when each branch's step genuinely does not need another branch's answer to run: screening a
request for policy problems while a separate step drafts a reply to it is independent work;
drafting a reply and then checking that same reply is not, because the check needs the draft
first. The benefit you are paying for is speed, not a better answer: running three calls at once
finishes in about the time of the slowest one instead of the sum of all three, which is worth
confirming actually happened before you trust the setup.

Coding agents show this on the canvas directly. Replit, announcing the fourth version of Replit
Agent, says "Independent tasks can run in parallel, with progress visible and coordinated," and
for larger jobs it "can split a single task into smaller pieces, work on them simultaneously with
sub-agents, and recombine the results"[2]. Claude Code's subagents feature documents
the same pattern for research: "For independent investigations, spawn multiple subagents to work
simultaneously". It says "Each subagent explores its area independently, then Claude synthesizes
the findings", adding that "This works best when the research paths don't depend on each
other"[3].

Open each branch after a parallel run finishes and read what it actually saw and returned, the
same way you would check one step of a plain chain: a tool that only shows the merged final
answer is hiding exactly the place two branches disagreed or repeated each other. Watch for
repetition specifically. Two branches that never saw each other's work can both answer the same
sub-question, so the combined result states one fact twice with nothing that noticed the overlap.

Reach for several independent opinions on the same question, rather than several different
sub-tasks, only where a wrong answer costs more than the extra run: a security review, a policy
call, a number somebody is about to act on. For anything routine, one pass is enough, and running
several is just several times the cost for no benefit anyone will notice.

## Implementation details

The example is sectioning: it retrieves a fixed set of candidate document sections, asks the
model to answer from each one *alone* (never seeing the other sections or the other calls) and
combines whichever sections actually answered part of the question. Because each call already
knows which single section it saw, citations are exact by construction; nothing has to be parsed
back out of free text the way RAG and prompt chaining do.

`ThreadPoolExecutor.map` is what makes this parallel rather than sequential: it submits every
call to the pool at once, and the calls run concurrently, but the returned iterator still yields
results in the order the candidates were given, regardless of which call actually finishes
first. That is what keeps the example deterministic without an explicit sort: order comes from
retrieval, never from a race between threads.

Every `tracer.record` call happens on the main thread, after `list(pool.map(...))` has already
collected every result: the trace itself is never written to from more than one thread at once.
A `StubModel` built from a fixed list of canned responses is not safe to call from several
threads concurrently, since it advances a shared counter with no lock; the example's tests build
their stub from a function that reads the prompt instead, which has no shared state to race on.

The trace above shows a real cost of naive sectioning: two of the three candidate sections
happened to answer the same sub-question, so the combined text states the filter fact twice.
Nothing in a fixed combine step notices the overlap, because each section answered without
seeing what the others said.

`examples/parallelization/run.py` (lines 22-75)

```python
LEVEL = 3
CANDIDATES_K = 3
NO_ANSWER = "NOT IN THIS SECTION"
PER_SECTION_SYSTEM = (
    "You are given exactly one source passage about Halvorsen appliances, and a question that "
    "may have more than one part. If this passage answers all or part of the question, answer "
    "briefly using only this passage. If it answers none of the question, reply with exactly "
    f"'{NO_ANSWER}' and nothing else."
)

def _answer_from_one_section(question: str, section: Section, model: Model) -> Completion:
    prompt = f"Passage [{section.cite}] {section.title}:\n{section.text}\n\nQuestion: {question}"
    return model.complete([Message(role="system", content=PER_SECTION_SYSTEM), Message(role="user", content=prompt)], max_tokens=200)

def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    k: int = CANDIDATES_K,
) -> Answer:
    del embedder  # candidates come from keyword search, not a vector index
    sections = load_sections(corpus_dir)
    candidates = [s for s, score in bm25_search(sections, question, k=k) if score > 0]
    tracer.record(
        kind="code", decided_by="code", title="Pick sections to answer in parallel",
        detail=", ".join(s.cite for s in candidates) or "none",
    )

    # .map submits every call to the pool at once and yields results back in candidate order,
    # so the calls run concurrently but the code below never has to sort them: determinism comes
    # from retrieval order, not from whichever call happens to finish first.
    with ThreadPoolExecutor(max_workers=max(1, len(candidates))) as pool:
        completions = list(pool.map(lambda s: _answer_from_one_section(question, s, model), candidates))

    for section, completion in zip(candidates, completions):
        tracer.record(
            kind="model", decided_by="code", title=f"Answer from {section.cite} alone", detail=completion.text[:200],
            tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
        )

    used = [(s, c) for s, c in zip(candidates, completions) if NO_ANSWER not in c.text.upper()]
    tracer.record(
        kind="code", decided_by="code", title="Combine the sections that answered",
        detail=f"{len(used)} of {len(candidates)} sections answered part of the question",
    )
    if not used:
        return Answer(text="None of the retrieved sections answered the question.", citations=[], retrieved_sources=[s.cite for s in candidates])
    combined = " ".join(c.text.strip() for _, c in used)
    return Answer(text=combined, citations=sorted({s.cite for s, _ in used}), retrieved_sources=[s.cite for s in candidates])
```

Run it yourself:

`examples/parallelization/README.md` (lines 16-16)

```text
python -m examples.parallelization --model stub:scripted
```

Real-time parallel calls like these are for when the answer is needed now. When it is not (a
nightly re-score of every open ticket, a one-time pass over a large document set), the batch
APIs three model makers publish do the same many-calls-one-submission idea asynchronously and
cheaper: Anthropic describes its Message Batches API as suited to tasks that do not need an
immediate response, "with most batches finishing in less than 1 hour while reducing costs by 50%
and increasing throughput"[4]; OpenAI's Batch API gives a "50% cost discount compared to
synchronous APIs" with each batch completing "within 24 hours (and often more quickly)"[5]; Google states that its Gemini Batch API processes requests at "50% of the standard cost"
and says of the wait: "The target turnaround time is 24 hours, but in majority of cases, it is
much quicker"[6]. Each of those is the maker's own published figure, checked on the
date in the source list below, and each is the same trade: give up the immediate response, halve
the price.

## When you do not need this

Try [prompt chaining](/gradient_ascent/techniques/prompt-chaining/) or a single call first if
the task's parts actually depend on each other: a later part needs an earlier part's answer,
or the sections would overlap and need to be reconciled against each other. Running dependent
work in parallel does not make it independent; it just hides the dependency until the combine
step produces a contradiction.

Move up to parallelization once the task genuinely splits into parts that do not need each
other's answers, and the parts are already known before any call runs: a section list you
retrieved, a fixed set of checks to run, a fixed number of independent opinions to gather.

## Failure modes

### Overlapping sections restate the same fact

- **How to notice it:** The combined answer repeats itself, or states the same fact in two slightly different ways, because two sections happened to cover the same ground and neither call could see the other's answer.
- **How to test for it:** Retrieve candidates for a question you know has redundant coverage across sections (the DW-480's filter is described in both its own manual and the shared care-and-cleaning guide) and check whether the combined text repeats the fact.

### A fixed combine step cannot resolve a disagreement

- **How to notice it:** Two sections answer the same question differently (an old figure and a superseding one) and a plain concatenation states both without saying which is current, because nothing in the combine step compares them against each other.
- **How to test for it:** Run a question over sections you know conflict (an original spec and a later correction) and check whether the combined answer states both values with no indication of which one is authoritative.

### A shared, mutable stub races under real concurrency

- **How to notice it:** A test or a manual run using a list-based StubModel raises IndexError or returns answers in the wrong order under a thread pool, because the stub's internal counter is not safe to advance from more than one thread.
- **How to test for it:** Run the example's own test suite; it is deliberately built on a callable-based stub for exactly this reason, and a regression toward a list-based stub under the thread pool would surface as an intermittent failure, not a consistent one.

### Rate limits under real load

- **How to notice it:** Firing many calls at once against a live API returns 429 rate-limit errors once concurrency crosses the provider’s per-minute limit, which a small stub run never exercises.
- **How to test for it:** Check the provider’s published rate limits against the number of parallel calls one request triggers, before running the example against a live model at any real question volume.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 3
- **Tokens in (summed):** ~520
- **Tokens out (summed):** ~47
- **Wall time:** ~0.65s

**Compared with prompt chaining (level 3, sequential).** Cost sums across the three calls, the same as a sequential chain would. Wall time does not: it tracks the slowest single call, not their sum, which is the entire latency argument for running independent work in parallel instead of one call after another.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

The same 60 questions, the same corpus, plus one number of its own: what share of a question's
`must_cite` sections were actually covered by *some* section's answer. Multi-hop questions test
that directly, since they need more than one section's fact combined into a single answer.

Conflicting-source questions are the interesting case to watch: sectioning retrieves both sides
of a deliberate contradiction as readily as RAG does, but its fixed combine step has no way to
compare them, only to concatenate whatever each section said. Whether that scores better or worse
than RAG's single stuffed-context prompt is exactly the kind of question this site can only
answer once a result file exists (see `docs/EVALS.md`); none does yet. Run `python
scripts/eval_run.py --example parallelization --model <spec> --dry` to project the cost of a real
run first.

## Run it

**What to monitor.** Per-branch latency and error rate, not just the overall run's. One slow or failing branch in a thread pool can dominate wall time even though the others finished quickly; averaging across branches hides exactly the branch worth investigating.

**Cost at volume.** Cost is the number of parallel calls times the number of questions, same as a sequential chain of the same length: parallelism buys latency, not a lower bill. For volume that does not need an immediate answer, a maker's batch API halves the per-call cost in exchange for asynchronous delivery.

**How it fails in production.** Concurrency crosses a provider's per-minute rate limit once real question volume arrives, producing errors a low-volume stub or manual test never triggers. Separately, a thread pool sized for a fixed number of sections silently under-uses itself if fewer candidates come back than expected, or queues up if more do.

**What to log.** Each branch's input, output, token counts and wall time individually, plus which branches were kept versus dropped by the combine step, so a bad or missing final answer traces back to one specific branch rather than to 'the parallel step' as a whole.

## Try it

1. **Use it.** Ask a coding agent that advertises parallel subagents to work on three independent parts. Does it run them at once, and does the result repeat itself?
2. **Build it.** Run python -m examples.parallelization --model stub:scripted from the repo root. Three branches answer from one section each, one says NOT IN THIS SECTION, and the combine keeps the two that did. Now change CANDIDATES_K from 3 to 5 in examples/parallelization/run.py: the run stops with ScriptExhausted, naming the passage no reply matches.
3. **Either lane.** Cause the overlap failure on purpose: find a question evals/corpus/ answers in two places, run it, and count how often the combined answer repeats itself.


## Sources

1. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic, 2024-12-19 (accessed 2026-09-19)
2. [Introducing Replit Agent 4: Built for Creativity](https://replit.com/blog/introducing-agent-4-built-for-creativity) — Replit (accessed 2026-09-19)
3. [Subagents](https://code.claude.com/docs/en/subagents) — Anthropic (Claude Code documentation) (accessed 2026-09-19)
4. [Batch processing](https://platform.claude.com/docs/en/build-with-claude/batch-processing) — Anthropic (accessed 2026-09-19)
5. [Batch API](https://developers.openai.com/api/docs/guides/batch) — OpenAI (accessed 2026-09-19)
6. [Batch API](https://ai.google.dev/gemini-api/docs/batch-api) — Google (accessed 2026-09-19)


Last reviewed 2026-09-19.
