Level 03 · Workflows

Write and check

One prompt writes, another checks, and the loop repeats until the check passes.

Sourced

Concept at a glance

Write, check, and revise.

Feedback loopConceptual illustration
Write, check, and revise.Task + criteria leads to Draft. Draft leads to Evaluate. Evaluate leads to Feedback. Feedback leads to Draft as feedback. A failed check feeds a revision. A passing check or a fixed cap ends the loop.Task + criteriaDefine what good meansDraftWrite a candidate answerEvaluateCheck against the criteriaFeedbackRevise if it did not passWrite, check, and revise.Task + criteria leads to Draft. Draft leads to Evaluate. Evaluate leads to Feedback. Feedback leads to Draft as feedback. A failed check feeds a revision. A passing check or a fixed cap ends the loop.Task + criteriaDefine what good meansDraftWrite a candidate answerEvaluateCheck against the criteriaFeedbackRevise if it did not pass

Ending or continuingStop when the check passes or the revision limit is reached.

Read the connections in words
  • Task + criteria → Draft: Write a candidate answer.
  • Draft → Evaluate: Check against the criteria.
  • Evaluate → Feedback: Revise if it did not pass.
  • Feedback → Draft: feedback informs another turn.
Key idea

A failed check feeds a revision. A passing check or a fixed cap ends the loop.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Write and check: see it in practice.

Iterating between generating a candidate and evaluating it against criteria until it passes or a limit stops the loop.

What you’ll walk through

Follow a draft through feedback and revision. Inspect whether the revision improves a stated criterion without losing facts or satisfying a weak reviewer through superficial changes.

The task in this version

Improve this onboarding article without unsupported policy claims.

What you’ll learn to check

Draft diffs, criterion-level feedback, a maximum-attempt stop, and an independent source check.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Improve this onboarding article without unsupported policy claims.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Draft: refunds always take one day. Policy: up to five working days. Review limit: two passes.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

The evaluator needs a meaningful rubric and enough evidence to judge it. Model-generated feedback can itself be mistaken.

1 / 6

Apply this to your project

Describe your task to your own model and use Write and check as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Write and check runs two prompts against each other: one writes, a separate one checks the result against explicit criteria, and if it fails, the first revises and the check runs again. Anthropic describes it as “one LLM call generates a response while another provides evaluation and feedback in a loop”, and calls the workflow “particularly effective when we have clear evaluation criteria, and when iterative refinement provides measurable value”: its examples are literary translation and “Complex search tasks that require multiple rounds of searching and analysis”[1].

“Clear evaluation criteria” is the load-bearing phrase: a checker asked whether an answer is “good” just drafts again with extra steps, since a vague verdict drifts between calls, while one asked a specific checkable question, like whether every citation appears in its sources, answers the same way every time.

The checker’s verdict does change what happens next, and the code branches on it. It is still level 3 because the code owns the loop, not the model: the while condition and the cap are written in advance. Hand the model the loop itself and the stop becomes a model decision: level 5.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Write and check

One call drafts, another checks against one explicit criterion, and the loop repeats until it passes or hits a cap.

Level 3 · Workflows
QuestionQuestionRetrieve sourcesRetrieve sourcesMODELDraftDraftMODELCheckCheckMODELReviseReviseAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 06Your code chose

The question arrives

"How often should the DW-300's filter be cleaned?"
0 tokens · 0 ms

Practical guidance

You can run this loop by hand in any chat app, and it is the move to reach for when a draft is nearly right and “make it better” has stopped changing anything. The one rule is that you write the criteria down before you read the draft. A criterion invented while looking at a draft is an opinion about that draft.

Three messages, in the same conversation.

  1. Ask for the draft. “Write a 200-word notice to tenants about the elevator being out from 4/6/2027 to 4/10/2027. Plain English, no apology longer than one sentence, say where the freight elevator is.”
  2. Hand over the checklist and ask for a verdict, not a rewrite. “Check the draft above against these four rules. Answer each one yes or no and quote the line that proves it: (1) it gives both dates, (2) it says where the freight elevator is, (3) it is under 220 words, (4) no sentence runs past 25 words. Do not rewrite it.”
  3. Fix only what failed. “Fix rules 2 and 4. Change nothing else.”

Stop after two rounds. If the same rule fails a third time, the rule is the problem rather than the draft: either nobody could tell yes from no by reading it, or what you actually want is something you have not written down yet.

The check on the exercise is whether the verdict ever changes anything. If every rule comes back yes on the first pass, add a rule you expect the draft to break and see whether it catches it. A checker that never says no is not a check, it is a delay.

Ask the same thing of any product advertising this. Does its checking step test something specific, or does it ask whether the draft is good? The second is common and mostly cosmetic: a pause and a second bill, not a check, because “is this good” is not something a second pass of the same kind of model answers more reliably than the first pass did. No product in this site’s registry is documented well enough to name here as a ready-made version of this loop, so treat a visible “reviewing” step as an unverified claim until the product’s documentation says what it tests.

When one clear instruction gets it right first time, skip all of this. That is prompt engineering, and it costs one message instead of three.

Implementation details

The example drafts an answer, asks a second, separate prompt whether every citation the draft claims actually appears among the sources it was given, and if not, revises using that verdict verbatim. PASS_TOKEN is the entire contract between the two prompts: the checker either returns it exactly, or returns the specific list of what is missing, and the code never has to interpret anything softer than a string match to know which case it got.

The loop’s shape is a while with two conditions the code owns completely: keep going while the last check failed and the cap has not been reached. max_revisions (default 2) is a plain function argument, not something the model can see or influence. When the cap is hit before a pass, the code records that explicitly and still returns the last draft: shipping an answer that is known to still fail its own check is a real, visible outcome here, not a bug hidden by the loop quietly running forever.

The trace above shows why the checker needs a narrow criterion. It catches the draft citing dw300-manual#6, a real section of the real corpus, that simply was not one of the four sources this particular retrieval handed to the draft step: a citation the checker can verify mechanically, with no judgment call. What it would not catch: every one of those four retrieved sources being wrong for the question, or the drafter and the checker sharing a blind spot, because they are typically driven by the same kind of model and can fail on the same kind of question the same way. That is the case for review and debate instead, where the second opinion is built to differ from the first on purpose.

examples/evaluator_optimizer/run.py · lines 22–107
LEVEL = 3
RETRIEVE_K = 4
MAX_REVISIONS = 2
PASS_TOKEN = "ALL CITATIONS SUPPORTED"
DRAFT_SYSTEM = (
    "You answer questions about Halvorsen appliances using only the numbered sources below. End "
    "your answer with a line starting 'Sources:' listing the citations, like 'dw300-manual#3', "
    "that you used."
)
CHECK_SYSTEM = (
    "You check a draft answer against the source passages it was given. List every citation the "
    "draft claims that does NOT actually appear among the sources below, one per line, as "
    f"'MISSING: <citation>'. If every citation the draft claims is one of the sources, reply with "
    f"exactly '{PASS_TOKEN}' and nothing else."
)
REVISE_SYSTEM = (
    "Revise your previous answer to fix the citation problems named below. Use only the sources "
    "given. Keep the same 'Sources:' line format."
)


def _sources_block(sources: list[Section]) -> str:
    return "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)


def _draft(question: str, sources: list[Section], model: Model, tracer: Tracer) -> str:
    prompt = f"Sources:\n\n{_sources_block(sources)}\n\nQuestion: {question}"
    completion = model.complete([Message(role="system", content=DRAFT_SYSTEM), Message(role="user", content=prompt)], max_tokens=400)
    tracer.record(
        kind="model", decided_by="code", title="Draft an answer", detail=completion.text[:200],
        tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )
    return completion.text


def _check(draft_text: str, sources: list[Section], model: Model, tracer: Tracer) -> str | None:
    """None means the draft passed. Otherwise, the checker's own feedback text."""
    prompt = f"Sources:\n\n{_sources_block(sources)}\n\nDraft answer:\n{draft_text}"
    completion = model.complete([Message(role="system", content=CHECK_SYSTEM), Message(role="user", content=prompt)], max_tokens=200)
    verdict = completion.text.strip()
    tracer.record(
        kind="model", decided_by="code", title="Check citations against the sources", detail=verdict[:200],
        tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )
    return None if PASS_TOKEN in verdict.upper() else verdict


def _revise(question: str, draft_text: str, feedback: str, sources: list[Section], model: Model, tracer: Tracer) -> str:
    prompt = f"Sources:\n\n{_sources_block(sources)}\n\nQuestion: {question}\n\nPrevious answer:\n{draft_text}\n\nProblems to fix:\n{feedback}"
    completion = model.complete([Message(role="system", content=REVISE_SYSTEM), Message(role="user", content=prompt)], max_tokens=400)
    tracer.record(
        kind="model", decided_by="code", title="Revise using the checker's feedback", detail=completion.text[:200],
        tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )
    return completion.text


def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    max_revisions: int = MAX_REVISIONS,
) -> Answer:
    del embedder  # retrieval here is keyword search, not a vector index
    sections = load_sections(corpus_dir)
    sources = [s for s, score in bm25_search(sections, question, k=RETRIEVE_K) if score > 0]
    tracer.record(kind="code", decided_by="code", title="Retrieve sources", detail=", ".join(s.cite for s in sources) or "none")

    draft_text = _draft(question, sources, model, tracer)
    feedback = _check(draft_text, sources, model, tracer)
    revisions = 0
    while feedback is not None and revisions < max_revisions:
        draft_text = _revise(question, draft_text, feedback, sources, model, tracer)
        revisions += 1
        feedback = _check(draft_text, sources, model, tracer)
    if feedback is not None:
        tracer.record(
            kind="code", decided_by="code", title="Stop: revision cap reached",
            detail=f"shipping a draft that still fails its own check after {revisions} revision(s)",
        )

    citations = cited_sources(draft_text)
    return Answer(text=draft_text, citations=citations, retrieved_sources=[s.cite for s in sources])

Run it yourself:

examples/evaluator_optimizer/README.md · lines 16–16
python -m examples.evaluator_optimizer --model stub:scripted

The same shape (a metric-driven loop instead of a vibe-driven one) is what DSPy’s optimizers do to a prompt itself, at build time rather than at answer time: “All optimizers read a numeric score per example,” and each one “tunes one or more of: instructions, demos, or weights”[2] to raise that score. DSPy loops over many training examples to improve the prompt before it ever answers a real question; this page’s loop runs once, at answer time, to improve one answer. Both need the same thing to work at all: a criterion specific enough that two runs of the check agree.

The design review checklist recipe is the worked version of this loop for engineering test and precise measurement alike: one pass drafts a finding against a design-review rule’s own text, a second checks that finding against the rule it cites and drops any that name none. Neither pass reports a measurement or a margin; where a rule is a number against a threshold, code computes it, and a person still decides whether an unmet rule ships or gets fixed.

When you do not need this

Try a single call, or a fixed, code-only check like the one prompt chaining’s example uses (comparing citations by set intersection, no second model call), first if the thing you would check for is something plain code can already test: that is cheaper, always consistent, and does not need a second prompt at all.

Move up to write and check once the failure you are trying to catch needs judgment against a written rule that plain code cannot express directly, but that a second prompt, told the rule in so many words, can apply consistently. This is the site’s draft and check shape.

Failure modes

A checker with no fixed criterion

How to notice it
The checker’s verdict changes between two runs on the same draft, because it was asked something open-ended ("is this good") rather than something specific enough to answer the same way twice.
How to test for it
Run the check step on the exact same draft and sources twice. A checker worth looping on returns the same verdict both times; one that does not is adding cost without adding reliability.

Writer and checker share a blind spot

How to notice it
The checker passes a draft that is confidently wrong in a way neither prompt would ever catch, because both were built from the same kind of model making the same kind of mistake.
How to test for it
Feed the checker a draft with a deliberate error of the kind its own criterion cannot see (a citation that is real and present, but supports the wrong fact) and confirm it passes, which is the specific gap review and debate exists to close.

The cap ships a known-bad answer

How to notice it
The loop reaches max_revisions still failing its own check, and the last draft goes out anyway, silently unless the "Stop: revision cap reached" step is actually surfaced somewhere a person or a downstream system can see it.
How to test for it
Force a draft that can never pass (script the checker to always find fault) and confirm the run still returns an answer rather than hanging or raising, and that the stop is recorded, not just implied by running out of steps.

The checker burns the whole cap on a trivial complaint

How to notice it
A near-miss the checker treats as failing (a citation formatted slightly differently from what it expects) consumes the same revision budget as a genuine problem, leaving fewer chances left for anything that actually matters.
How to test for it
Compare how many revisions a trivially-imperfect draft uses against how many a genuinely wrong one uses; if they are the same, the checker's criterion may be too literal to be worth a full revision cycle.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

2Model calls, best case (passes first check)
6Model calls, worst case (2 revisions)
~350Tokens in, one check call
~1.6sWall time, one revision cycle
Compared with RAG (level 2), one callCost here is not fixed per question the way RAG’s is: a question whose draft passes immediately costs about twice what RAG costs, and one that exhausts the cap costs up to three times that, for the same question.

How to Evaluate It

Follow the reviewer feedback loop to test a brief with a fresh receiving agent, review the proposal against original requirements, and revise within a fixed budget. It includes blind comparisons, parallel providers, and prompts for each role.

60 questionslookupmulti-hopnumericunanswerableconflicting sources

Scored on the same 60-question set as every other technique, with two numbers specific to this loop: the average number of revisions a question used, and the share of questions that hit max_revisions still failing their own check. A high cap-hit rate on real traffic is a sign the checker’s criterion is too strict for what the drafter can realistically satisfy, or that the drafter has a systematic problem the checker keeps finding but the model cannot fix from feedback alone.

No result file exists yet (see docs/EVALS.md), so this page cannot say what that rate actually is here. Run python scripts/eval_run.py --example evaluator_optimizer --model <spec> --dry to project the cost of a real run first: the projection matters more for this technique than most, since its real cost depends on how often questions need a revision, not just how many questions there are.

Run it

What to monitor

The cap-hit rate over time (the share of runs that exhaust max_revisions still failing) and the average revisions per run. A cap-hit rate that rises with no change to the checker's prompt is a sign the questions arriving have shifted, not that the checker got stricter.

Cost at volume

Cost per question is variable here, unlike a fixed-step workflow: it depends on how often the checker fails the first draft. Budget for the worst case (every question uses the full cap), not the average, or a bad week of drafts becomes a bad week of the bill too.

How it fails in production

The writer and the checker share a blind spot neither prompt was built to catch, so confidently wrong answers pass their own check at the normal rate, with no signal in the trace that anything is different from a correctly-checked answer.

What to log

Every draft, every check verdict verbatim, and every revision, in order, for each question, not just the final answer. A checker that starts passing bad drafts is invisible in aggregate metrics until you can see its actual verdicts change.

Try it

  1. Use it

    Find a tool that shows a checking or verifying step before its final answer. Does its documentation say what the check actually tests, or only that a check happens?

  2. Build it

    Run python -m examples.evaluator_optimizer --model stub:scripted from the repo root. The draft cites dw300-manual#6, the checker answers MISSING since retrieval never returned it, the revision cites one it did, and the check passes. Then edit CHECK_SYSTEM in examples/evaluator_optimizer/run.py to ask something vague ("Is this a good answer?") instead of that citation test, and say why it could not give the same verdict twice.

  3. Either lane

    Write the checking criterion you would use for work you review. Would two people applying it to the same draft reach the same verdict? If not, it is not a criterion yet.

  4. Build it

    Read DR-14 in the design-review-checklist recipe's design-review-rules.md next to this page's checker. Both catch one quotable, mechanical thing. Now find a rule there that needs reading rather than arithmetic, and say why a single checker pass is less trustworthy on it.

How it connects

Before, after and instead of this

Read first

Move up when

Often used with

Pages that need this one

Optional: products, tools, and models

1 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Revise a draft against a rubric

A writer produces a draft and a checker flags missed requirements. Feed the feedback into a bounded revision loop.

Out there

Named products, tools and models

Tools1
  • DSPyStanford NLP · prompt programs and optimizers

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Building effective agents · Anthropic, 12/19/2024 (accessed 09/19/2026)
  2. Optimizers: choosing one · DSPy (Stanford NLP) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page