# Prompt chaining

_Level 03 · Workflows · sourced_

Splitting a task into steps, each with its own prompt.


## Try this in a recipe
- [Build a weekly update without invented progress](/gradient_ascent/recipes/weekly-status-report.md): Extract evidence into a checked table, then draft an update from that table in a fixed two-call workflow.

## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow work through dependent stages, where each stage hands a specific result to the next. Inspect how an early evidence error can survive into a polished final draft.

**Assumptions:** Later stages depend on the quality and completeness of earlier outputs. A successful model response is not necessarily a successful handoff.

**Design choices:** Use separate stages when they have distinct responsibilities or checks. Keep a single call when splitting adds overhead without improving control or quality.

**Request:** Prepare our weekly report through evidence, project summaries, and a final draft.

**Starting evidence:** Previous report: Atlas on track. Current tracker: milestone delayed to Friday. Notes: cause under investigation.

**Action and control:** Fixed stage 1 extracts dated facts; stage 2 writes the project summary; stage 3 assembles the draft. No stage sends it.

**Stage records (authored, not executed):**

### Input record

Previous report: Atlas on track. Current tracker: milestone delayed to Friday. Notes: cause under investigation.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use separate stages when they have distinct responsibilities or checks. Keep a single call when splitting adds overhead without improving control or quality.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Fixed stage 1 extracts dated facts; stage 2 writes the project summary; stage 3 assembles the draft. No stage sends it.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Evidence: milestone moved. Summary: schedule risk. Draft: delayed to Friday, cause under investigation; review pending.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A source-linked evidence sheet, intermediate project summaries, report diff, and a corrected unsupported claim.

If the result falls short:
Stop or repair the affected stage when a required handoff is incomplete. Preserve accepted upstream work rather than rerunning every stage blindly.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Adapt the sequence to research, writing, analysis, or reporting. Define each stage's inputs, output contract, and useful checks; the number of stages is not fixed.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Evidence: milestone moved. Summary: schedule risk. Draft: delayed to Friday, cause under investigation; review pending.

**Change something — Misread the milestone during extraction:** The wrong date propagates into polished prose. Correct the intermediate evidence sheet before regenerating downstream work.

**Decision:** Where should you first correct a propagated factual error?

**Answer:** At extraction, then regenerate downstream work.

**Why:** An extraction error can become a polished false claim downstream; previous reports supply continuity, not proof of current status.

**Review criteria:** A source-linked evidence sheet, intermediate project summaries, report diff, and a corrected unsupported claim.

**Recovery:** Stop or repair the affected stage when a required handoff is incomplete. Preserve accepted upstream work rather than rerunning every stage blindly.

**Adapt it:** Adapt the sequence to research, writing, analysis, or reporting. Define each stage's inputs, output contract, and useful checks; the number of stages is not fixed.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow work through dependent stages, where each stage hands a specific result to the next. Inspect how an early evidence error can survive into a polished final draft.

**Assumptions:** Later stages depend on the quality and completeness of earlier outputs. A successful model response is not necessarily a successful handoff.

**Design choices:** Use separate stages when they have distinct responsibilities or checks. Keep a single call when splitting adds overhead without improving control or quality.

**Request:** Turn a school newsletter into a family action list and calendar draft.

**Starting evidence:** Newsletter: costume day Friday; permission form due Wednesday. No event times supplied.

**Action and control:** First extract facts, then group actions, then prepare calendar drafts without invented times.

**Stage records (authored, not executed):**

### Input record

Newsletter: costume day Friday; permission form due Wednesday. No event times supplied.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use separate stages when they have distinct responsibilities or checks. Keep a single call when splitting adds overhead without improving control or quality.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

First extract facts, then group actions, then prepare calendar drafts without invented times.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Action: return form Wednesday. Reminder draft: costume day Friday, time unspecified.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Compare extracted dates with the newsletter before accepting calendar drafts.

If the result falls short:
Stop or repair the affected stage when a required handoff is incomplete. Preserve accepted upstream work rather than rerunning every stage blindly.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Adapt the sequence to research, writing, analysis, or reporting. Define each stage's inputs, output contract, and useful checks; the number of stages is not fixed.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Action: return form Wednesday. Reminder draft: costume day Friday, time unspecified.

**Change something — Extraction swaps Wednesday and Friday:** Wrong deadlines propagate through every later step. Fix extraction and regenerate the downstream drafts.

**Decision:** Where should a propagated deadline error be repaired?

**Answer:** At extraction, then redo dependent outputs.

**Why:** Fixed chains make intermediate artifacts useful checkpoints.

**Review criteria:** Compare extracted dates with the newsletter before accepting calendar drafts.

**Recovery:** Stop or repair the affected stage when a required handoff is incomplete. Preserve accepted upstream work rather than rerunning every stage blindly.

**Adapt it:** Adapt the sequence to research, writing, analysis, or reporting. Define each stage's inputs, output contract, and useful checks; the number of stages is not fixed.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow work through dependent stages, where each stage hands a specific result to the next. Inspect how an early evidence error can survive into a polished final draft.

**Assumptions:** Later stages depend on the quality and completeness of earlier outputs. A successful model response is not necessarily a successful handoff.

**Design choices:** Use separate stages when they have distinct responsibilities or checks. Keep a single call when splitting adds overhead without improving control or quality.

**Request:** Turn a requirement into a test outline and review checklist.

**Starting evidence:** Requirement: verify output remains within supplied bounds after a 20 ms settling interval.

**Action and control:** Extract parameters, map framework functions, then draft sequence and checks in fixed stages.

**Stage records (authored, not executed):**

### Input record

Requirement: verify output remains within supplied bounds after a 20 ms settling interval.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use separate stages when they have distinct responsibilities or checks. Keep a single call when splitting adds overhead without improving control or quality.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Extract parameters, map framework functions, then draft sequence and checks in fixed stages.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Outline preserves the settling interval, cites approved bounds, and calls existing measurement functions.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Review the extracted requirement, API mapping, and expected measured quantity separately.

If the result falls short:
Stop or repair the affected stage when a required handoff is incomplete. Preserve accepted upstream work rather than rerunning every stage blindly.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Adapt the sequence to research, writing, analysis, or reporting. Define each stage's inputs, output contract, and useful checks; the number of stages is not fixed.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Outline preserves the settling interval, cites approved bounds, and calls existing measurement functions.

**Change something — Function-mapping stage selects a different measurement mode:** Later code can look polished but measure the wrong quantity. Correct mapping before generation.

**Decision:** Does a valid final script prove the intermediate mapping was right?

**Answer:** No; inspect requirements-to-function mapping.

**Why:** A chain can faithfully propagate an early semantic error.

**Review criteria:** Review the extracted requirement, API mapping, and expected measured quantity separately.

**Recovery:** Stop or repair the affected stage when a required handoff is incomplete. Preserve accepted upstream work rather than rerunning every stage blindly.

**Adapt it:** Adapt the sequence to research, writing, analysis, or reporting. Define each stage's inputs, output contract, and useful checks; the number of stages is not fixed.

Prompt chaining splits one task into a fixed sequence of steps, and hands each step's output to
the next. Anthropic's own description is direct: it "decomposes a task into a sequence of steps,
where each LLM call processes the output of the previous one"[1]. A model call can sit
inside any step, but the sequence itself, and what happens between steps, is fixed by your code
before the chain ever runs.

The step between two model calls is usually a gate: ordinary code that checks the output so far
before letting the chain continue. Anthropic gives two examples: generate marketing copy, then
translate it, or write a document outline, check it against a rule, then write the document from
that outline[1].

Prompt chaining sits at level 3, workflows. The model fills in the content of each step; your
code decides how many steps there are, what order they run in, and what gate sits between them.
The line to level 4 falls where the model's output starts selecting what runs next: where it
is offered a tool and can choose to call it.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 3 · Prompt chaining. The same steps are described in the sections below._

## Practical guidance

Build a chain in an automation tool with a visual canvas: a trigger, then an ordered list of
steps, each one able to use what came before it. Start with the trigger, such as when a form is
submitted or when an email arrives, add one step that does one clear job (summarize the message,
draft a reply from the summary), and connect them in order. Automation services such as Zapier,
Make, n8n and Power Automate are all built around this shape.

Add a gate between two steps rather than trusting the chain straight through: an ordinary
condition, checked in the tool itself, before the next step is allowed to run. "Only continue if
the drafted reply names a dollar figure" is a gate; it does not ask a model whether the draft
looks fine, it tests something specific in the output and stops the chain when that is missing.

Before you trust a chain, open each step on the canvas and check what it was actually given and
what it actually returned, not just the final result at the end. n8n advertises this directly:
"Every step of your agents' reasoning, traceable on the canvas"[2], and that is the
thing to check for in any of them, whatever the tool. If a step's input does not include
something you assumed it would, the original request, an earlier step's full output, that is
usually where a chain silently goes wrong: a translated document can read fluently while being a
fluent translation of a document an earlier step got wrong, and nothing at the end is checking it
against the original request, only against the step before it.

Build only as many steps as the task needs. A chain that always runs four fixed steps costs more
than one that runs one, on a question a single step could already answer. If nothing between the
first step and the last is worth checking on its own, that is a sign the chain is more machinery
than the job needs, and a single step will do. When a step fails outright rather than just
answering badly, check whether the tool retries that one step alone or reruns the whole chain
from the trigger; the second is a slower, more expensive habit worth knowing about before it
happens on something time-sensitive.

## Implementation details

The example runs the same four steps on every question, around a keyword search: rewrite the
question into up to three short search queries, retrieve for each query separately, draft an
answer from everything retrieved, then check the draft's citations against what retrieval
actually found. Two of the four steps call the model (the rewrite and the draft), but the code
decides that sequence before either call happens, and always runs all four steps regardless of
what either call returns. That is the whole difference from level 4: here the model fills in
step *content*; it never picks the next step.

The fourth step is the gate. `_check_citations` takes the set of citations the draft actually
claims and the set of sections retrieval actually found, and keeps only the intersection: a
citation the model invented, to a section nothing ever retrieved, is silently dropped rather
than trusted. This is exactly the gate Anthropic describes: "You can add programmatic checks (see
'gate' in the diagram below) on any intermediate steps to ensure that the process is still on
track"[1]. The check does not ask the model whether it did well; it tests the output
against a fact the code can verify on its own.

Because each step is an ordinary function that takes plain values and returns plain values, each
one is testable without the others and without a real model. `_check_citations` takes a draft
string and a list of sections and returns the grounded set: a test can hand it a draft that
invents a citation and assert it gets dropped, with no model call anywhere in the test. The
site's own test suite does exactly this for the chain as a whole, against a scripted stub model.

A chain fails most often at its weakest single step, and the failure travels forward invisibly.
If the rewrite step turns "is the vent length still 35 feet" into a query that only matches the
original manual, retrieval never sees the correcting service bulletin, and the draft answers
confidently from stale text: nothing downstream can tell that the search itself was incomplete.
Frameworks built to run fixed multi-step processes at production scale (Temporal, Prefect,
Apache Airflow, Inngest) exist mainly to make that first kind of failure recoverable rather than
silent: Temporal's own description is that its workflows "automatically capture state at every
step, and in the event of failure, can pick up exactly where they left off"[3], which
is a durability guarantee this example's plain function calls do not have.

The same shape shows up turning a requirements list into a test plan: read each requirement (say,
the SRB-5030's datasheet limits), propose a test for it, build a traceability table linking tests
to requirements, then check that every requirement has a test and every test names a requirement.
All four steps stay in this order regardless of the model's answers, the way the walkthrough
above does, and a person still approves the table before a production test sequence or an
engineering characterization plan is built from it. Prompt chaining is one of the techniques
behind the [turn a goal or a set of requirements into a
structured plan](/gradient_ascent/shapes/#plan-and-decompose) job shape.

`examples/prompt_chaining/run.py` (lines 20-97)

```python
LEVEL = 3
MAX_QUERIES = 3
PER_QUERY_K = 2
REWRITE_SYSTEM = (
    "Break the user's question into 1 to 3 short search queries over Halvorsen appliance "
    "documents, one per line, plain text, no numbering."
)
DRAFT_SYSTEM = (
    "You answer questions about Halvorsen appliances using only the numbered sources below. "
    "If the sources do not contain the answer, say so instead of guessing. End your answer with "
    "a line starting 'Sources:' listing the citations, like 'dw300-manual#3', that you used."
)

def _rewrite_queries(question: str, model: Model, tracer: Tracer) -> list[str]:
    completion = model.complete([Message(role="system", content=REWRITE_SYSTEM), Message(role="user", content=question)], max_tokens=150)
    queries = [line.strip() for line in completion.text.splitlines() if line.strip()][:MAX_QUERIES] or [question]
    tracer.record(
        kind="model",
        decided_by="code",
        title="Rewrite into search queries",
        detail="; ".join(queries),
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    return queries

def _retrieve(queries: list[str], sections: dict[str, Section], tracer: Tracer) -> list[Section]:
    seen: dict[str, Section] = {}
    for query in queries:
        for section, score in bm25_search(sections, query, k=PER_QUERY_K):
            if score > 0:
                seen[section.cite] = section
    sources = list(seen.values())
    tracer.record(kind="code", decided_by="code", title="Retrieve for each query", detail=", ".join(seen.keys()) or "none")
    return sources

def _draft(question: str, sources: list[Section], model: Model, tracer: Tracer):
    blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
    prompt = f"Sources:\n\n{blocks}\n\nQuestion: {question}"
    completion = model.complete([Message(role="system", content=DRAFT_SYSTEM), Message(role="user", content=prompt)], max_tokens=500)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Draft answer from sources",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    return completion

def _check_citations(draft_text: str, sources: list[Section], tracer: Tracer) -> list[str]:
    retrieved = {s.cite for s in sources}
    claimed = set(cited_sources(draft_text))
    grounded = sorted(claimed & retrieved)
    dropped = sorted(claimed - retrieved)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check citations against retrieval",
        detail=f"kept {grounded}" + (f", dropped ungrounded {dropped}" if dropped else ""),
    )
    return grounded

def run(question: str, model: Model, embedder: Embedder | None, tracer: Tracer, *, corpus_dir: Path = DEFAULT_CORPUS_DIR) -> Answer:
    del embedder  # level 3 retrieves by keyword, not by vector
    sections = load_sections(corpus_dir)
    queries = _rewrite_queries(question, model, tracer)
    sources = _retrieve(queries, sections, tracer)
    completion = _draft(question, sources, model, tracer)
    citations = _check_citations(completion.text, sources, tracer)
    return Answer(text=completion.text, citations=citations, retrieved_sources=[s.cite for s in sources])
```

Run it yourself:

`examples/prompt_chaining/README.md` (lines 16-16)

```text
python -m examples.prompt_chaining --model stub:scripted
```

## When you do not need this

Try [RAG](/gradient_ascent/techniques/rag/) or a single call first if one retrieval pass and one
answer already handles the question: a chain that always runs four fixed steps costs more than
one that runs one, for no benefit on a question a single pass could already answer.

Move up to prompt chaining once a task genuinely needs more than one model-filled step in a
known order, with something worth checking in between: rewriting a query before searching,
outlining before writing, drafting before verifying.

## Failure modes

### A bad step early in the chain travels forward unnoticed

- **How to notice it:** A later step's output looks fine on its own, but is built from a wrong or incomplete result earlier in the chain that nothing re-checked against the original request.
- **How to test for it:** Feed a deliberately bad output into the middle of the chain (call a later step directly with it) and see whether anything downstream catches it, or only whether the final text reads smoothly.

### A gate that never fails

- **How to notice it:** The programmatic check between two steps always passes, on every input, including ones it should catch: usually because the check tests something the step can never actually get wrong, rather than the thing that matters.
- **How to test for it:** Deliberately produce the exact failure the gate exists to catch (an invented citation, an outline missing a required section) and confirm the gate rejects it, not just that it accepts good input.

### The chain runs every step, even when the question did not need them

- **How to notice it:** A question a single call could answer still pays for all N steps and all N model calls, because the chain has no way to skip ahead.
- **How to test for it:** Time and cost a batch of easy, single-fact questions through the chain and compare against a single call; the gap is the fixed cost of running every step unconditionally.

### Step boundaries lose information

- **How to notice it:** A step is designed to pass forward only its stated output (a list of queries, a draft), so a detail the next step actually needed, but that was not part of the handoff, is gone by the time it would matter.
- **How to test for it:** Compare what the first step could see (the full question) against what the last step can see (only what earlier steps decided to pass on) for a question with a qualifying detail buried in its middle.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 2
- **Tokens in:** ~430
- **Tokens out:** ~66
- **Wall time:** ~2.0s

**Compared with RAG (level 2).** One extra model call to rewrite the question into queries, in the illustrated run above. The citation check itself costs nothing extra: it runs in code, not as a model call.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Scored on the same 60-question set as every other technique, over the appliance documents in
`evals/corpus/`. The citation check gives prompt chaining an extra number RAG does not have on
its own: how often a citation the draft claims was actually something retrieval found, tracked
separately from whether the final answer was correct.

Conflicting-source and multi-hop questions are where the extra query rewrite step is expected to
earn its cost: a single retrieval pass over "is the vent length still 35 feet" can miss the
correcting service bulletin entirely, where a second, differently worded query aimed at it has a
chance to find it. No result file exists yet (see `docs/EVALS.md`), so this page cannot say
whether that expectation holds. Run `python scripts/eval_run.py --example prompt_chaining --model
<spec> --dry` to project the cost of a real run before spending anything on one.

## Run it

**What to monitor.** How many steps a run actually completes versus how many it was supposed to; a chain that silently short-circuits is worse than one that errors loudly. Also track the citation-check drop rate: a rising share of invented citations is a sign the draft step is drifting.

**Cost at volume.** Cost scales with the number of steps times the number of questions, not with question difficulty, since every question runs every step. Two model calls per question here means roughly twice the language-model spend of a single-call or RAG pipeline at the same volume.

**How it fails in production.** An early step's prompt or the document set it depends on changes, and the step keeps returning plausible-looking output that is now subtly wrong; nothing downstream is positioned to notice, because each step only checks against the step before it, never against the original request.

**What to log.** Every step's input and output, not just the final answer, with the gate's verdict at each check. A bad final answer is only debuggable if you can see which of the N steps actually introduced the problem.

## Try it

1. **Use it.** Find a multi-step automation you use: an email rule, a form that files a ticket, a scheduled report. Write its steps down in order. Which one, if it silently got something wrong, would nobody downstream catch?
2. **Build it.** Run python -m examples.prompt_chaining --model stub:scripted from the repo root. The rewrite turns one question into three queries, retrieval brings back five sections, and the citation check keeps service-bulletin#2, the bulletin correcting the manual. Now set MAX_QUERIES in examples/prompt_chaining/run.py from 3 to 1: one query returns two sections, neither the bulletin, and the check drops the citation the draft still claims.
3. **Either lane.** Take the gate that never fails, above, and cause it on purpose: find a check in a process you run that has never once rejected anything, and work out whether that is because nothing bad has arrived or because the check cannot see the thing that would be bad.
4. **Either lane.** Take a short requirements list you actually have and chain it by hand: propose a test for each requirement, build the traceability table, then check that every requirement has a test and every test names a requirement. Where does the chain, not the requirements, turn out to be the hard part?


## Sources

1. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic, 2024-12-19 (accessed 2026-09-19)
2. [n8n](https://n8n.io) — n8n (accessed 2026-09-19)
3. [Temporal](https://temporal.io) — Temporal (accessed 2026-09-19)


Last reviewed 2026-09-19.
