Primary sources
- Building effective agents · Anthropic, 12/19/2024 (accessed 09/19/2026)
- n8n · n8n (accessed 09/19/2026)
- Temporal · Temporal (accessed 09/19/2026)
Splitting a task into steps, each with its own prompt.
Sourced
Concept at a glance
The code fixes the steps and their order before the first model call.
Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.
Running model calls in a predetermined sequence where one step's output feeds the next.
Follow work through dependent stages, where each stage hands a specific result to the next. Inspect how an early evidence error can survive into a polished final draft.
Prepare our weekly report through evidence, project summaries, and a final draft.
A source-linked evidence sheet, intermediate project summaries, report diff, and a corrected unsupported claim.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Prepare our weekly report through evidence, project summaries, and a final draft.
Authored case. Select any record below; nothing is sent to a model.What changed: Establish the facts supplied for this version of the task.
Later stages depend on the quality and completeness of earlier outputs. A successful model response is not necessarily a successful handoff.
Describe your task to your own model and use Prompt chaining as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Prompt chaining splits one task into a fixed sequence of steps, and hands each step’s output to the next. Anthropic’s own description is direct: it “decomposes a task into a sequence of steps, where each LLM call processes the output of the previous one”[1]. A model call can sit inside any step, but the sequence itself, and what happens between steps, is fixed by your code before the chain ever runs.
The step between two model calls is usually a gate: ordinary code that checks the output so far before letting the chain continue. Anthropic gives two examples: generate marketing copy, then translate it, or write a document outline, check it against a rule, then write the document from that outline[1].
Prompt chaining sits at level 3, workflows. The model fills in the content of each step; your code decides how many steps there are, what order they run in, and what gate sits between them. The line to level 4 falls where the model’s output starts selecting what runs next: where it is offered a tool and can choose to call it.
This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Break the task into fixed steps. The model fills each one in; your code decides the order.
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
"Is the DR-520 vent length still 35 feet?"
Build a chain in an automation tool with a visual canvas: a trigger, then an ordered list of steps, each one able to use what came before it. Start with the trigger, such as when a form is submitted or when an email arrives, add one step that does one clear job (summarize the message, draft a reply from the summary), and connect them in order. Automation services such as Zapier, Make, n8n and Power Automate are all built around this shape.
Add a gate between two steps rather than trusting the chain straight through: an ordinary condition, checked in the tool itself, before the next step is allowed to run. “Only continue if the drafted reply names a dollar figure” is a gate; it does not ask a model whether the draft looks fine, it tests something specific in the output and stops the chain when that is missing.
Before you trust a chain, open each step on the canvas and check what it was actually given and what it actually returned, not just the final result at the end. n8n advertises this directly: “Every step of your agents’ reasoning, traceable on the canvas”[2], and that is the thing to check for in any of them, whatever the tool. If a step’s input does not include something you assumed it would, the original request, an earlier step’s full output, that is usually where a chain silently goes wrong: a translated document can read fluently while being a fluent translation of a document an earlier step got wrong, and nothing at the end is checking it against the original request, only against the step before it.
Build only as many steps as the task needs. A chain that always runs four fixed steps costs more than one that runs one, on a question a single step could already answer. If nothing between the first step and the last is worth checking on its own, that is a sign the chain is more machinery than the job needs, and a single step will do. When a step fails outright rather than just answering badly, check whether the tool retries that one step alone or reruns the whole chain from the trigger; the second is a slower, more expensive habit worth knowing about before it happens on something time-sensitive.
The example runs the same four steps on every question, around a keyword search: rewrite the question into up to three short search queries, retrieve for each query separately, draft an answer from everything retrieved, then check the draft’s citations against what retrieval actually found. Two of the four steps call the model (the rewrite and the draft), but the code decides that sequence before either call happens, and always runs all four steps regardless of what either call returns. That is the whole difference from level 4: here the model fills in step content; it never picks the next step.
The fourth step is the gate. _check_citations takes the set of citations the draft actually
claims and the set of sections retrieval actually found, and keeps only the intersection: a
citation the model invented, to a section nothing ever retrieved, is silently dropped rather
than trusted. This is exactly the gate Anthropic describes: “You can add programmatic checks (see
‘gate’ in the diagram below) on any intermediate steps to ensure that the process is still on
track”[1]. The check does not ask the model whether it did well; it tests the output
against a fact the code can verify on its own.
Because each step is an ordinary function that takes plain values and returns plain values, each
one is testable without the others and without a real model. _check_citations takes a draft
string and a list of sections and returns the grounded set: a test can hand it a draft that
invents a citation and assert it gets dropped, with no model call anywhere in the test. The
site’s own test suite does exactly this for the chain as a whole, against a scripted stub model.
A chain fails most often at its weakest single step, and the failure travels forward invisibly. If the rewrite step turns “is the vent length still 35 feet” into a query that only matches the original manual, retrieval never sees the correcting service bulletin, and the draft answers confidently from stale text: nothing downstream can tell that the search itself was incomplete. Frameworks built to run fixed multi-step processes at production scale (Temporal, Prefect, Apache Airflow, Inngest) exist mainly to make that first kind of failure recoverable rather than silent: Temporal’s own description is that its workflows “automatically capture state at every step, and in the event of failure, can pick up exactly where they left off”[3], which is a durability guarantee this example’s plain function calls do not have.
The same shape shows up turning a requirements list into a test plan: read each requirement (say, the SRB-5030’s datasheet limits), propose a test for it, build a traceability table linking tests to requirements, then check that every requirement has a test and every test names a requirement. All four steps stay in this order regardless of the model’s answers, the way the walkthrough above does, and a person still approves the table before a production test sequence or an engineering characterization plan is built from it. Prompt chaining is one of the techniques behind the turn a goal or a set of requirements into a structured plan job shape.
LEVEL = 3
MAX_QUERIES = 3
PER_QUERY_K = 2
REWRITE_SYSTEM = (
"Break the user's question into 1 to 3 short search queries over Halvorsen appliance "
"documents, one per line, plain text, no numbering."
)
DRAFT_SYSTEM = (
"You answer questions about Halvorsen appliances using only the numbered sources below. "
"If the sources do not contain the answer, say so instead of guessing. End your answer with "
"a line starting 'Sources:' listing the citations, like 'dw300-manual#3', that you used."
)
def _rewrite_queries(question: str, model: Model, tracer: Tracer) -> list[str]:
completion = model.complete([Message(role="system", content=REWRITE_SYSTEM), Message(role="user", content=question)], max_tokens=150)
queries = [line.strip() for line in completion.text.splitlines() if line.strip()][:MAX_QUERIES] or [question]
tracer.record(
kind="model",
decided_by="code",
title="Rewrite into search queries",
detail="; ".join(queries),
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
return queries
def _retrieve(queries: list[str], sections: dict[str, Section], tracer: Tracer) -> list[Section]:
seen: dict[str, Section] = {}
for query in queries:
for section, score in bm25_search(sections, query, k=PER_QUERY_K):
if score > 0:
seen[section.cite] = section
sources = list(seen.values())
tracer.record(kind="code", decided_by="code", title="Retrieve for each query", detail=", ".join(seen.keys()) or "none")
return sources
def _draft(question: str, sources: list[Section], model: Model, tracer: Tracer):
blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
prompt = f"Sources:\n\n{blocks}\n\nQuestion: {question}"
completion = model.complete([Message(role="system", content=DRAFT_SYSTEM), Message(role="user", content=prompt)], max_tokens=500)
tracer.record(
kind="model",
decided_by="code",
title="Draft answer from sources",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
return completion
def _check_citations(draft_text: str, sources: list[Section], tracer: Tracer) -> list[str]:
retrieved = {s.cite for s in sources}
claimed = set(cited_sources(draft_text))
grounded = sorted(claimed & retrieved)
dropped = sorted(claimed - retrieved)
tracer.record(
kind="code",
decided_by="code",
title="Check citations against retrieval",
detail=f"kept {grounded}" + (f", dropped ungrounded {dropped}" if dropped else ""),
)
return grounded
def run(question: str, model: Model, embedder: Embedder | None, tracer: Tracer, *, corpus_dir: Path = DEFAULT_CORPUS_DIR) -> Answer:
del embedder # level 3 retrieves by keyword, not by vector
sections = load_sections(corpus_dir)
queries = _rewrite_queries(question, model, tracer)
sources = _retrieve(queries, sections, tracer)
completion = _draft(question, sources, model, tracer)
citations = _check_citations(completion.text, sources, tracer)
return Answer(text=completion.text, citations=citations, retrieved_sources=[s.cite for s in sources])Run it yourself:
python -m examples.prompt_chaining --model stub:scriptedTry RAG or a single call first if one retrieval pass and one answer already handles the question: a chain that always runs four fixed steps costs more than one that runs one, for no benefit on a question a single pass could already answer.
Move up to prompt chaining once a task genuinely needs more than one model-filled step in a known order, with something worth checking in between: rewriting a query before searching, outlining before writing, drafting before verifying.
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
Scored on the same 60-question set as every other technique, over the appliance documents in
evals/corpus/. The citation check gives prompt chaining an extra number RAG does not have on
its own: how often a citation the draft claims was actually something retrieval found, tracked
separately from whether the final answer was correct.
Conflicting-source and multi-hop questions are where the extra query rewrite step is expected to
earn its cost: a single retrieval pass over “is the vent length still 35 feet” can miss the
correcting service bulletin entirely, where a second, differently worded query aimed at it has a
chance to find it. No result file exists yet (see docs/EVALS.md), so this page cannot say
whether that expectation holds. Run python scripts/eval_run.py --example prompt_chaining --model <spec> --dry to project the cost of a real run before spending anything on one.
How many steps a run actually completes versus how many it was supposed to; a chain that silently short-circuits is worse than one that errors loudly. Also track the citation-check drop rate: a rising share of invented citations is a sign the draft step is drifting.
Cost scales with the number of steps times the number of questions, not with question difficulty, since every question runs every step. Two model calls per question here means roughly twice the language-model spend of a single-call or RAG pipeline at the same volume.
An early step's prompt or the document set it depends on changes, and the step keeps returning plausible-looking output that is now subtly wrong; nothing downstream is positioned to notice, because each step only checks against the step before it, never against the original request.
Every step's input and output, not just the final answer, with the gate's verdict at each check. A bad final answer is only debuggable if you can see which of the N steps actually introduced the problem.
Find a multi-step automation you use: an email rule, a form that files a ticket, a scheduled report. Write its steps down in order. Which one, if it silently got something wrong, would nobody downstream catch?
Run python -m examples.prompt_chaining --model stub:scripted from the repo root. The rewrite turns one question into three queries, retrieval brings back five sections, and the citation check keeps service-bulletin#2, the bulletin correcting the manual. Now set MAX_QUERIES in examples/prompt_chaining/run.py from 3 to 1: one query returns two sections, neither the bulletin, and the check drops the citation the draft still claims.
Take the gate that never fails, above, and cause it on purpose: find a check in a process you run that has never once rejected anything, and work out whether that is because nothing bad has arrived or because the check cannot see the thing that would be bad.
Take a short requirements list you actually have and chain it by hand: propose a test for each requirement, build the traceability table, then check that every requirement has a test and every test names a requirement. Where does the chain, not the requirements, turn out to be the hard part?
10 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.
Automation service
Maker’s documentation Checked 09/18/2026Automation service, self-hostable
Maker’s documentation Checked 09/18/2026Automation service
Maker’s documentation Checked 09/18/2026Automation service
Maker’s documentation Checked 09/18/2026Workflow engine
Maker’s documentation Checked 09/18/2026Retrieval framework
Maker’s documentation Checked 09/18/2026Durable workflow engine
Checked 09/18/2026Application framework
Checked 09/18/2026Workflow engine
Checked 09/18/2026Durable workflow engine
Checked 09/18/2026One prompt extracts facts, another drafts the report, and a final step checks its claims against the notes.
Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page