Primary sources
- Building effective agents · Anthropic, 12/19/2024 (accessed 09/19/2026)
- Optimizers: choosing one · DSPy (Stanford NLP) (accessed 09/19/2026)
One prompt writes, another checks, and the loop repeats until the check passes.
Sourced
Concept at a glance
Ending or continuingStop when the check passes or the revision limit is reached.
A failed check feeds a revision. A passing check or a fixed cap ends the loop.
A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.
Iterating between generating a candidate and evaluating it against criteria until it passes or a limit stops the loop.
Follow a draft through feedback and revision. Inspect whether the revision improves a stated criterion without losing facts or satisfying a weak reviewer through superficial changes.
Improve this onboarding article without unsupported policy claims.
Draft diffs, criterion-level feedback, a maximum-attempt stop, and an independent source check.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Improve this onboarding article without unsupported policy claims.
Authored case. Select any record below; nothing is sent to a model.What changed: Establish the facts supplied for this version of the task.
The evaluator needs a meaningful rubric and enough evidence to judge it. Model-generated feedback can itself be mistaken.
Describe your task to your own model and use Write and check as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Write and check runs two prompts against each other: one writes, a separate one checks the result against explicit criteria, and if it fails, the first revises and the check runs again. Anthropic describes it as “one LLM call generates a response while another provides evaluation and feedback in a loop”, and calls the workflow “particularly effective when we have clear evaluation criteria, and when iterative refinement provides measurable value”: its examples are literary translation and “Complex search tasks that require multiple rounds of searching and analysis”[1].
“Clear evaluation criteria” is the load-bearing phrase: a checker asked whether an answer is “good” just drafts again with extra steps, since a vague verdict drifts between calls, while one asked a specific checkable question, like whether every citation appears in its sources, answers the same way every time.
The checker’s verdict does change what happens next, and the code branches on it. It is still
level 3 because the code owns the loop, not the model: the while condition and the cap are
written in advance. Hand the model the loop itself and the stop becomes a model decision: level 5.
This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
One call drafts, another checks against one explicit criterion, and the loop repeats until it passes or hits a cap.
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
"How often should the DW-300's filter be cleaned?"
You can run this loop by hand in any chat app, and it is the move to reach for when a draft is nearly right and “make it better” has stopped changing anything. The one rule is that you write the criteria down before you read the draft. A criterion invented while looking at a draft is an opinion about that draft.
Three messages, in the same conversation.
Stop after two rounds. If the same rule fails a third time, the rule is the problem rather than the draft: either nobody could tell yes from no by reading it, or what you actually want is something you have not written down yet.
The check on the exercise is whether the verdict ever changes anything. If every rule comes back yes on the first pass, add a rule you expect the draft to break and see whether it catches it. A checker that never says no is not a check, it is a delay.
Ask the same thing of any product advertising this. Does its checking step test something specific, or does it ask whether the draft is good? The second is common and mostly cosmetic: a pause and a second bill, not a check, because “is this good” is not something a second pass of the same kind of model answers more reliably than the first pass did. No product in this site’s registry is documented well enough to name here as a ready-made version of this loop, so treat a visible “reviewing” step as an unverified claim until the product’s documentation says what it tests.
When one clear instruction gets it right first time, skip all of this. That is prompt engineering, and it costs one message instead of three.
The example drafts an answer, asks a second, separate prompt whether every citation the draft
claims actually appears among the sources it was given, and if not, revises using that verdict
verbatim. PASS_TOKEN is the entire contract between the two prompts: the checker either returns
it exactly, or returns the specific list of what is missing, and the code never has to interpret
anything softer than a string match to know which case it got.
The loop’s shape is a while with two conditions the code owns completely: keep going while the
last check failed and the cap has not been reached. max_revisions (default 2) is a plain
function argument, not something the model can see or influence. When the cap is hit before a
pass, the code records that explicitly and still returns the last draft: shipping an answer that
is known to still fail its own check is a real, visible outcome here, not a bug hidden by the
loop quietly running forever.
The trace above shows why the checker needs a narrow criterion. It catches the draft citing
dw300-manual#6, a real section of the real corpus, that simply was not one of the four sources
this particular retrieval handed to the draft step: a citation the checker can verify
mechanically, with no judgment call. What it would not catch: every one of those four retrieved
sources being wrong for the question, or the drafter and the checker sharing a blind spot,
because they are typically driven by the same kind of model and can fail on the same kind of
question the same way. That is the case for review and
debate instead, where the second opinion is built to differ from the first on purpose.
LEVEL = 3
RETRIEVE_K = 4
MAX_REVISIONS = 2
PASS_TOKEN = "ALL CITATIONS SUPPORTED"
DRAFT_SYSTEM = (
"You answer questions about Halvorsen appliances using only the numbered sources below. End "
"your answer with a line starting 'Sources:' listing the citations, like 'dw300-manual#3', "
"that you used."
)
CHECK_SYSTEM = (
"You check a draft answer against the source passages it was given. List every citation the "
"draft claims that does NOT actually appear among the sources below, one per line, as "
f"'MISSING: <citation>'. If every citation the draft claims is one of the sources, reply with "
f"exactly '{PASS_TOKEN}' and nothing else."
)
REVISE_SYSTEM = (
"Revise your previous answer to fix the citation problems named below. Use only the sources "
"given. Keep the same 'Sources:' line format."
)
def _sources_block(sources: list[Section]) -> str:
return "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
def _draft(question: str, sources: list[Section], model: Model, tracer: Tracer) -> str:
prompt = f"Sources:\n\n{_sources_block(sources)}\n\nQuestion: {question}"
completion = model.complete([Message(role="system", content=DRAFT_SYSTEM), Message(role="user", content=prompt)], max_tokens=400)
tracer.record(
kind="model", decided_by="code", title="Draft an answer", detail=completion.text[:200],
tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
)
return completion.text
def _check(draft_text: str, sources: list[Section], model: Model, tracer: Tracer) -> str | None:
"""None means the draft passed. Otherwise, the checker's own feedback text."""
prompt = f"Sources:\n\n{_sources_block(sources)}\n\nDraft answer:\n{draft_text}"
completion = model.complete([Message(role="system", content=CHECK_SYSTEM), Message(role="user", content=prompt)], max_tokens=200)
verdict = completion.text.strip()
tracer.record(
kind="model", decided_by="code", title="Check citations against the sources", detail=verdict[:200],
tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
)
return None if PASS_TOKEN in verdict.upper() else verdict
def _revise(question: str, draft_text: str, feedback: str, sources: list[Section], model: Model, tracer: Tracer) -> str:
prompt = f"Sources:\n\n{_sources_block(sources)}\n\nQuestion: {question}\n\nPrevious answer:\n{draft_text}\n\nProblems to fix:\n{feedback}"
completion = model.complete([Message(role="system", content=REVISE_SYSTEM), Message(role="user", content=prompt)], max_tokens=400)
tracer.record(
kind="model", decided_by="code", title="Revise using the checker's feedback", detail=completion.text[:200],
tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
)
return completion.text
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
max_revisions: int = MAX_REVISIONS,
) -> Answer:
del embedder # retrieval here is keyword search, not a vector index
sections = load_sections(corpus_dir)
sources = [s for s, score in bm25_search(sections, question, k=RETRIEVE_K) if score > 0]
tracer.record(kind="code", decided_by="code", title="Retrieve sources", detail=", ".join(s.cite for s in sources) or "none")
draft_text = _draft(question, sources, model, tracer)
feedback = _check(draft_text, sources, model, tracer)
revisions = 0
while feedback is not None and revisions < max_revisions:
draft_text = _revise(question, draft_text, feedback, sources, model, tracer)
revisions += 1
feedback = _check(draft_text, sources, model, tracer)
if feedback is not None:
tracer.record(
kind="code", decided_by="code", title="Stop: revision cap reached",
detail=f"shipping a draft that still fails its own check after {revisions} revision(s)",
)
citations = cited_sources(draft_text)
return Answer(text=draft_text, citations=citations, retrieved_sources=[s.cite for s in sources])Run it yourself:
python -m examples.evaluator_optimizer --model stub:scriptedThe same shape (a metric-driven loop instead of a vibe-driven one) is what DSPy’s optimizers do to a prompt itself, at build time rather than at answer time: “All optimizers read a numeric score per example,” and each one “tunes one or more of: instructions, demos, or weights”[2] to raise that score. DSPy loops over many training examples to improve the prompt before it ever answers a real question; this page’s loop runs once, at answer time, to improve one answer. Both need the same thing to work at all: a criterion specific enough that two runs of the check agree.
The design review checklist recipe is the worked version of this loop for engineering test and precise measurement alike: one pass drafts a finding against a design-review rule’s own text, a second checks that finding against the rule it cites and drops any that name none. Neither pass reports a measurement or a margin; where a rule is a number against a threshold, code computes it, and a person still decides whether an unmet rule ships or gets fixed.
Try a single call, or a fixed, code-only check like the one prompt chaining’s example uses (comparing citations by set intersection, no second model call), first if the thing you would check for is something plain code can already test: that is cheaper, always consistent, and does not need a second prompt at all.
Move up to write and check once the failure you are trying to catch needs judgment against a written rule that plain code cannot express directly, but that a second prompt, told the rule in so many words, can apply consistently. This is the site’s draft and check shape.
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
Follow the reviewer feedback loop to test a brief with a fresh receiving agent, review the proposal against original requirements, and revise within a fixed budget. It includes blind comparisons, parallel providers, and prompts for each role.
Scored on the same 60-question set as every other technique, with two numbers specific to this
loop: the average number of revisions a question used, and the share of questions that hit
max_revisions still failing their own check. A high cap-hit rate on real traffic is a sign the
checker’s criterion is too strict for what the drafter can realistically satisfy, or that the
drafter has a systematic problem the checker keeps finding but the model cannot fix from feedback
alone.
No result file exists yet (see docs/EVALS.md), so this page cannot say what that rate actually
is here. Run python scripts/eval_run.py --example evaluator_optimizer --model <spec> --dry to
project the cost of a real run first: the projection matters more for this technique than most,
since its real cost depends on how often questions need a revision, not just how many questions
there are.
The cap-hit rate over time (the share of runs that exhaust max_revisions still failing) and the average revisions per run. A cap-hit rate that rises with no change to the checker's prompt is a sign the questions arriving have shifted, not that the checker got stricter.
Cost per question is variable here, unlike a fixed-step workflow: it depends on how often the checker fails the first draft. Budget for the worst case (every question uses the full cap), not the average, or a bad week of drafts becomes a bad week of the bill too.
The writer and the checker share a blind spot neither prompt was built to catch, so confidently wrong answers pass their own check at the normal rate, with no signal in the trace that anything is different from a correctly-checked answer.
Every draft, every check verdict verbatim, and every revision, in order, for each question, not just the final answer. A checker that starts passing bad drafts is invisible in aggregate metrics until you can see its actual verdicts change.
Find a tool that shows a checking or verifying step before its final answer. Does its documentation say what the check actually tests, or only that a check happens?
Run python -m examples.evaluator_optimizer --model stub:scripted from the repo root. The draft cites dw300-manual#6, the checker answers MISSING since retrieval never returned it, the revision cites one it did, and the check passes. Then edit CHECK_SYSTEM in examples/evaluator_optimizer/run.py to ask something vague ("Is this a good answer?") instead of that citation test, and say why it could not give the same verdict twice.
Write the checking criterion you would use for work you review. Would two people applying it to the same draft reach the same verdict? If not, it is not a criterion yet.
Read DR-14 in the design-review-checklist recipe's design-review-rules.md next to this page's checker. Both catch one quotable, mechanical thing. Now find a rule there that needs reading rather than arithmetic, and say why a single checker pass is less trustworthy on it.
1 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.
Prompt programs and optimizers
Maker’s documentation Checked 09/19/2026A writer produces a draft and a checker flags missed requirements. Feed the feedback into a bounded revision loop.
Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page