Level 03 · Workflows

Prompt chaining

Splitting a task into steps, each with its own prompt.

Sourced

Concept at a glance

Hand each step’s output to the next.

SequenceConceptual illustration
Hand each step’s output to the next.Rewrite the question leads to Draft from sources. Draft from sources leads to Check the draft. The code fixes the steps and their order before the first model call.Rewrite the questionFirst promptDraft from sourcesSecond promptCheck the draftPass along a checked resultHand each step’s output to the next.Rewrite the question leads to Draft from sources. Draft from sources leads to Check the draft. The code fixes the steps and their order before the first model call.Rewrite the questionFirst promptDraft from sourcesSecond promptCheck the draftPass along a checked result
Read the connections in words
  • Rewrite the question → Draft from sources: Second prompt.
  • Draft from sources → Check the draft: Pass along a checked result.
Key idea

The code fixes the steps and their order before the first model call.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Prompt chaining: see it in practice.

Running model calls in a predetermined sequence where one step's output feeds the next.

What you’ll walk through

Follow work through dependent stages, where each stage hands a specific result to the next. Inspect how an early evidence error can survive into a polished final draft.

The task in this version

Prepare our weekly report through evidence, project summaries, and a final draft.

What you’ll learn to check

A source-linked evidence sheet, intermediate project summaries, report diff, and a corrected unsupported claim.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Prepare our weekly report through evidence, project summaries, and a final draft.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Previous report: Atlas on track. Current tracker: milestone delayed to Friday. Notes: cause under investigation.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Later stages depend on the quality and completeness of earlier outputs. A successful model response is not necessarily a successful handoff.

1 / 6

Apply this to your project

Describe your task to your own model and use Prompt chaining as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Prompt chaining splits one task into a fixed sequence of steps, and hands each step’s output to the next. Anthropic’s own description is direct: it “decomposes a task into a sequence of steps, where each LLM call processes the output of the previous one”[1]. A model call can sit inside any step, but the sequence itself, and what happens between steps, is fixed by your code before the chain ever runs.

The step between two model calls is usually a gate: ordinary code that checks the output so far before letting the chain continue. Anthropic gives two examples: generate marketing copy, then translate it, or write a document outline, check it against a rule, then write the document from that outline[1].

Prompt chaining sits at level 3, workflows. The model fills in the content of each step; your code decides how many steps there are, what order they run in, and what gate sits between them. The line to level 4 falls where the model’s output starts selecting what runs next: where it is offered a tool and can choose to call it.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Prompt chaining

Break the task into fixed steps. The model fills each one in; your code decides the order.

Level 3 · Workflows
QuestionQuestionMODELRewrite into queriesRewrite into queriesRetrieve per queryRetrieve per queryMODELDraft an answerDraft an answerCheck citationsCheck citationsAnswerAnswerQuestionQuestionMODELRewrite into queriesRewrite into queriesRetrieve per queryRetrieve per queryMODELDraft an answerDraft an answerCheck citationsCheck citationsAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

The question arrives

"Is the DR-520 vent length still 35 feet?"
0 tokens · 0 ms

Practical guidance

Build a chain in an automation tool with a visual canvas: a trigger, then an ordered list of steps, each one able to use what came before it. Start with the trigger, such as when a form is submitted or when an email arrives, add one step that does one clear job (summarize the message, draft a reply from the summary), and connect them in order. Automation services such as Zapier, Make, n8n and Power Automate are all built around this shape.

Add a gate between two steps rather than trusting the chain straight through: an ordinary condition, checked in the tool itself, before the next step is allowed to run. “Only continue if the drafted reply names a dollar figure” is a gate; it does not ask a model whether the draft looks fine, it tests something specific in the output and stops the chain when that is missing.

Before you trust a chain, open each step on the canvas and check what it was actually given and what it actually returned, not just the final result at the end. n8n advertises this directly: “Every step of your agents’ reasoning, traceable on the canvas”[2], and that is the thing to check for in any of them, whatever the tool. If a step’s input does not include something you assumed it would, the original request, an earlier step’s full output, that is usually where a chain silently goes wrong: a translated document can read fluently while being a fluent translation of a document an earlier step got wrong, and nothing at the end is checking it against the original request, only against the step before it.

Build only as many steps as the task needs. A chain that always runs four fixed steps costs more than one that runs one, on a question a single step could already answer. If nothing between the first step and the last is worth checking on its own, that is a sign the chain is more machinery than the job needs, and a single step will do. When a step fails outright rather than just answering badly, check whether the tool retries that one step alone or reruns the whole chain from the trigger; the second is a slower, more expensive habit worth knowing about before it happens on something time-sensitive.

Implementation details

The example runs the same four steps on every question, around a keyword search: rewrite the question into up to three short search queries, retrieve for each query separately, draft an answer from everything retrieved, then check the draft’s citations against what retrieval actually found. Two of the four steps call the model (the rewrite and the draft), but the code decides that sequence before either call happens, and always runs all four steps regardless of what either call returns. That is the whole difference from level 4: here the model fills in step content; it never picks the next step.

The fourth step is the gate. _check_citations takes the set of citations the draft actually claims and the set of sections retrieval actually found, and keeps only the intersection: a citation the model invented, to a section nothing ever retrieved, is silently dropped rather than trusted. This is exactly the gate Anthropic describes: “You can add programmatic checks (see ‘gate’ in the diagram below) on any intermediate steps to ensure that the process is still on track”[1]. The check does not ask the model whether it did well; it tests the output against a fact the code can verify on its own.

Because each step is an ordinary function that takes plain values and returns plain values, each one is testable without the others and without a real model. _check_citations takes a draft string and a list of sections and returns the grounded set: a test can hand it a draft that invents a citation and assert it gets dropped, with no model call anywhere in the test. The site’s own test suite does exactly this for the chain as a whole, against a scripted stub model.

A chain fails most often at its weakest single step, and the failure travels forward invisibly. If the rewrite step turns “is the vent length still 35 feet” into a query that only matches the original manual, retrieval never sees the correcting service bulletin, and the draft answers confidently from stale text: nothing downstream can tell that the search itself was incomplete. Frameworks built to run fixed multi-step processes at production scale (Temporal, Prefect, Apache Airflow, Inngest) exist mainly to make that first kind of failure recoverable rather than silent: Temporal’s own description is that its workflows “automatically capture state at every step, and in the event of failure, can pick up exactly where they left off”[3], which is a durability guarantee this example’s plain function calls do not have.

The same shape shows up turning a requirements list into a test plan: read each requirement (say, the SRB-5030’s datasheet limits), propose a test for it, build a traceability table linking tests to requirements, then check that every requirement has a test and every test names a requirement. All four steps stay in this order regardless of the model’s answers, the way the walkthrough above does, and a person still approves the table before a production test sequence or an engineering characterization plan is built from it. Prompt chaining is one of the techniques behind the turn a goal or a set of requirements into a structured plan job shape.

examples/prompt_chaining/run.py · lines 20–97
LEVEL = 3
MAX_QUERIES = 3
PER_QUERY_K = 2
REWRITE_SYSTEM = (
    "Break the user's question into 1 to 3 short search queries over Halvorsen appliance "
    "documents, one per line, plain text, no numbering."
)
DRAFT_SYSTEM = (
    "You answer questions about Halvorsen appliances using only the numbered sources below. "
    "If the sources do not contain the answer, say so instead of guessing. End your answer with "
    "a line starting 'Sources:' listing the citations, like 'dw300-manual#3', that you used."
)


def _rewrite_queries(question: str, model: Model, tracer: Tracer) -> list[str]:
    completion = model.complete([Message(role="system", content=REWRITE_SYSTEM), Message(role="user", content=question)], max_tokens=150)
    queries = [line.strip() for line in completion.text.splitlines() if line.strip()][:MAX_QUERIES] or [question]
    tracer.record(
        kind="model",
        decided_by="code",
        title="Rewrite into search queries",
        detail="; ".join(queries),
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    return queries


def _retrieve(queries: list[str], sections: dict[str, Section], tracer: Tracer) -> list[Section]:
    seen: dict[str, Section] = {}
    for query in queries:
        for section, score in bm25_search(sections, query, k=PER_QUERY_K):
            if score > 0:
                seen[section.cite] = section
    sources = list(seen.values())
    tracer.record(kind="code", decided_by="code", title="Retrieve for each query", detail=", ".join(seen.keys()) or "none")
    return sources


def _draft(question: str, sources: list[Section], model: Model, tracer: Tracer):
    blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
    prompt = f"Sources:\n\n{blocks}\n\nQuestion: {question}"
    completion = model.complete([Message(role="system", content=DRAFT_SYSTEM), Message(role="user", content=prompt)], max_tokens=500)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Draft answer from sources",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    return completion


def _check_citations(draft_text: str, sources: list[Section], tracer: Tracer) -> list[str]:
    retrieved = {s.cite for s in sources}
    claimed = set(cited_sources(draft_text))
    grounded = sorted(claimed & retrieved)
    dropped = sorted(claimed - retrieved)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check citations against retrieval",
        detail=f"kept {grounded}" + (f", dropped ungrounded {dropped}" if dropped else ""),
    )
    return grounded


def run(question: str, model: Model, embedder: Embedder | None, tracer: Tracer, *, corpus_dir: Path = DEFAULT_CORPUS_DIR) -> Answer:
    del embedder  # level 3 retrieves by keyword, not by vector
    sections = load_sections(corpus_dir)
    queries = _rewrite_queries(question, model, tracer)
    sources = _retrieve(queries, sections, tracer)
    completion = _draft(question, sources, model, tracer)
    citations = _check_citations(completion.text, sources, tracer)
    return Answer(text=completion.text, citations=citations, retrieved_sources=[s.cite for s in sources])

Run it yourself:

examples/prompt_chaining/README.md · lines 16–16
python -m examples.prompt_chaining --model stub:scripted
When you do not need this

Try RAG or a single call first if one retrieval pass and one answer already handles the question: a chain that always runs four fixed steps costs more than one that runs one, for no benefit on a question a single pass could already answer.

Move up to prompt chaining once a task genuinely needs more than one model-filled step in a known order, with something worth checking in between: rewriting a query before searching, outlining before writing, drafting before verifying.

Failure modes

A bad step early in the chain travels forward unnoticed

How to notice it
A later step's output looks fine on its own, but is built from a wrong or incomplete result earlier in the chain that nothing re-checked against the original request.
How to test for it
Feed a deliberately bad output into the middle of the chain (call a later step directly with it) and see whether anything downstream catches it, or only whether the final text reads smoothly.

A gate that never fails

How to notice it
The programmatic check between two steps always passes, on every input, including ones it should catch: usually because the check tests something the step can never actually get wrong, rather than the thing that matters.
How to test for it
Deliberately produce the exact failure the gate exists to catch (an invented citation, an outline missing a required section) and confirm the gate rejects it, not just that it accepts good input.

The chain runs every step, even when the question did not need them

How to notice it
A question a single call could answer still pays for all N steps and all N model calls, because the chain has no way to skip ahead.
How to test for it
Time and cost a batch of easy, single-fact questions through the chain and compare against a single call; the gap is the fixed cost of running every step unconditionally.

Step boundaries lose information

How to notice it
A step is designed to pass forward only its stated output (a list of queries, a draft), so a detail the next step actually needed, but that was not part of the handoff, is gone by the time it would matter.
How to test for it
Compare what the first step could see (the full question) against what the last step can see (only what earlier steps decided to pass on) for a question with a qualifying detail buried in its middle.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

2Model calls, one question
~430Tokens in
~66Tokens out
~2.0sWall time
Compared with RAG (level 2)One extra model call to rewrite the question into queries, in the illustrated run above. The citation check itself costs nothing extra: it runs in code, not as a model call.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

Scored on the same 60-question set as every other technique, over the appliance documents in evals/corpus/. The citation check gives prompt chaining an extra number RAG does not have on its own: how often a citation the draft claims was actually something retrieval found, tracked separately from whether the final answer was correct.

Conflicting-source and multi-hop questions are where the extra query rewrite step is expected to earn its cost: a single retrieval pass over “is the vent length still 35 feet” can miss the correcting service bulletin entirely, where a second, differently worded query aimed at it has a chance to find it. No result file exists yet (see docs/EVALS.md), so this page cannot say whether that expectation holds. Run python scripts/eval_run.py --example prompt_chaining --model <spec> --dry to project the cost of a real run before spending anything on one.

Run it

What to monitor

How many steps a run actually completes versus how many it was supposed to; a chain that silently short-circuits is worse than one that errors loudly. Also track the citation-check drop rate: a rising share of invented citations is a sign the draft step is drifting.

Cost at volume

Cost scales with the number of steps times the number of questions, not with question difficulty, since every question runs every step. Two model calls per question here means roughly twice the language-model spend of a single-call or RAG pipeline at the same volume.

How it fails in production

An early step's prompt or the document set it depends on changes, and the step keeps returning plausible-looking output that is now subtly wrong; nothing downstream is positioned to notice, because each step only checks against the step before it, never against the original request.

What to log

Every step's input and output, not just the final answer, with the gate's verdict at each check. A bad final answer is only debuggable if you can see which of the N steps actually introduced the problem.

Try it

  1. Use it

    Find a multi-step automation you use: an email rule, a form that files a ticket, a scheduled report. Write its steps down in order. Which one, if it silently got something wrong, would nobody downstream catch?

  2. Build it

    Run python -m examples.prompt_chaining --model stub:scripted from the repo root. The rewrite turns one question into three queries, retrieval brings back five sections, and the citation check keeps service-bulletin#2, the bulletin correcting the manual. Now set MAX_QUERIES in examples/prompt_chaining/run.py from 3 to 1: one query returns two sections, neither the bulletin, and the check drops the citation the draft still claims.

  3. Either lane

    Take the gate that never fails, above, and cause it on purpose: find a check in a process you run that has never once rejected anything, and work out whether that is because nothing bad has arrived or because the check cannot see the thing that would be bad.

  4. Either lane

    Take a short requirements list you actually have and chain it by hand: propose a test for each requirement, build the traceability table, then check that every requirement has a test and every test names a requirement. Where does the chain, not the requirements, turn out to be the hard part?

How it connects

Before, after and instead of this

Move up when

Optional: products, tools, and models

10 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 4 more examples
  • Inngest Tool or framework · Inngest

    Durable workflow engine

    Checked 09/18/2026
  • LangChain Tool or framework · LangChain

    Application framework

    Checked 09/18/2026
  • Prefect Tool or framework · Prefect

    Workflow engine

    Checked 09/18/2026
  • Temporal Tool or framework · Temporal

    Durable workflow engine

    Checked 09/18/2026
In practice

Turn notes into a checked report

One prompt extracts facts, another drafts the report, and a final step checks its claims against the notes.

Out there

Named products, tools and models

Products4
  • MakeCelonis · automation service
  • n8nn8n · automation service, self-hostable
  • Power AutomateMicrosoft · automation service
  • ZapierZapier · automation service
Tools6
  • Apache Airflowopen source · workflow engine
  • Haystackdeepset · retrieval framework
  • InngestInngest · durable workflow engine
  • LangChainLangChain · application framework
  • PrefectPrefect · workflow engine
  • TemporalTemporal · durable workflow engine

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Building effective agents · Anthropic, 12/19/2024 (accessed 09/19/2026)
  2. n8n · n8n (accessed 09/19/2026)
  3. Temporal · Temporal (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page