# Write and check

_Level 03 · Workflows · sourced_

One prompt writes, another checks, and the loop repeats until the check passes.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a draft through feedback and revision. Inspect whether the revision improves a stated criterion without losing facts or satisfying a weak reviewer through superficial changes.

**Assumptions:** The evaluator needs a meaningful rubric and enough evidence to judge it. Model-generated feedback can itself be mistaken.

**Design choices:** Use automatic checks for measurable constraints and judgment for qualities that need it. Set a stopping condition so revisions do not continue without useful improvement.

**Request:** Improve this onboarding article without unsupported policy claims.

**Starting evidence:** Draft: refunds always take one day. Policy: up to five working days. Review limit: two passes.

**Action and control:** Check claims against policy, revise, and independently verify the changed sentence.

**Stage records (authored, not executed):**

### Input record

Draft: refunds always take one day. Policy: up to five working days. Review limit: two passes.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use automatic checks for measurable constraints and judgment for qualities that need it. Set a stopping condition so revisions do not continue without useful improvement.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Check claims against policy, revise, and independently verify the changed sentence.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Revision: refunds may take up to five working days. Source supports this claim; other claims still need review.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Draft diffs, criterion-level feedback, a maximum-attempt stop, and an independent source check.

If the result falls short:
Keep the best acceptable version when a revision regresses. Resolve conflicting feedback against the task's priorities instead of repeatedly oscillating.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply the loop to writing, code, plans, or analysis. Choose a small rubric, preserve required facts, and decide when a human review or simple first draft is sufficient.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Revision: refunds may take up to five working days. Source supports this claim; other claims still need review.

**Change something — Reviewer checks only readability:** A clearer one-day claim remains false. Fix the review criteria rather than blindly repeat.

**Decision:** Can a fluent revision pass factual review automatically?

**Answer:** No; verify it against the source.

**Why:** A reviewer can miss errors or reward superficial fixes; repeated polishing may never converge.

**Review criteria:** Draft diffs, criterion-level feedback, a maximum-attempt stop, and an independent source check.

**Recovery:** Keep the best acceptable version when a revision regresses. Resolve conflicting feedback against the task's priorities instead of repeatedly oscillating.

**Adapt it:** Apply the loop to writing, code, plans, or analysis. Choose a small rubric, preserve required facts, and decide when a human review or simple first draft is sufficient.

Write and check runs two prompts against each other: one writes, a separate one checks the result
against explicit criteria, and if it fails, the first revises and the check runs again. Anthropic
describes it as "one LLM call generates a response while another provides evaluation and feedback
in a loop", and calls the workflow "particularly effective when we have clear evaluation criteria,
and when iterative refinement provides measurable value": its examples are literary translation
and "Complex search tasks that require multiple rounds of searching and analysis"[1].

"Clear evaluation criteria" is the load-bearing phrase: a checker asked whether an answer is
"good" just drafts again with extra steps, since a vague verdict drifts between calls, while one
asked a specific checkable question, like whether every citation appears in its sources, answers
the same way every time.

The checker's verdict does change what happens next, and the code branches on it. It is still
level 3 because the code owns the loop, not the model: the `while` condition and the cap are
written in advance. Hand the model the loop itself and the stop becomes a model decision: level 5.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 3 · Write and check. The same steps are described in the sections below._

## Practical guidance

You can run this loop by hand in any chat app, and it is the move to reach for when a draft is
nearly right and "make it better" has stopped changing anything. The one rule is that you write
the criteria down before you read the draft. A criterion invented while looking at a draft is an
opinion about that draft.

Three messages, in the same conversation.

1. Ask for the draft. "Write a 200-word notice to tenants about the elevator being out from
   4/6/2027 to 4/10/2027. Plain English, no apology longer than one sentence, say where the
   freight elevator is."
2. Hand over the checklist and ask for a verdict, not a rewrite. "Check the draft above against
   these four rules. Answer each one yes or no and quote the line that proves it: (1) it gives
   both dates, (2) it says where the freight elevator is, (3) it is under 220 words, (4) no
   sentence runs past 25 words. Do not rewrite it."
3. Fix only what failed. "Fix rules 2 and 4. Change nothing else."

Stop after two rounds. If the same rule fails a third time, the rule is the problem rather than
the draft: either nobody could tell yes from no by reading it, or what you actually want is
something you have not written down yet.

The check on the exercise is whether the verdict ever changes anything. If every rule comes back
yes on the first pass, add a rule you expect the draft to break and see whether it catches it. A
checker that never says no is not a check, it is a delay.

Ask the same thing of any product advertising this. Does its checking step test something
specific, or does it ask whether the draft is good? The second is common and mostly cosmetic: a
pause and a second bill, not a check, because "is this good" is not something a second pass of the
same kind of model answers more reliably than the first pass did. No product in this site's
registry is documented well enough to name here as a ready-made version of this loop, so treat a
visible "reviewing" step as an unverified claim until the product's documentation says what it
tests.

When one clear instruction gets it right first time, skip all of this. That is
[prompt engineering](/gradient_ascent/techniques/prompt-engineering/), and it costs one message
instead of three.

## Implementation details

The example drafts an answer, asks a second, separate prompt whether every citation the draft
claims actually appears among the sources it was given, and if not, revises using that verdict
verbatim. `PASS_TOKEN` is the entire contract between the two prompts: the checker either returns
it exactly, or returns the specific list of what is missing, and the code never has to interpret
anything softer than a string match to know which case it got.

The loop's shape is a `while` with two conditions the code owns completely: keep going while the
last check failed *and* the cap has not been reached. `max_revisions` (default 2) is a plain
function argument, not something the model can see or influence. When the cap is hit before a
pass, the code records that explicitly and still returns the last draft: shipping an answer that
is known to still fail its own check is a real, visible outcome here, not a bug hidden by the
loop quietly running forever.

The trace above shows why the checker needs a *narrow* criterion. It catches the draft citing
`dw300-manual#6`, a real section of the real corpus, that simply was not one of the four sources
this particular retrieval handed to the draft step: a citation the checker can verify
mechanically, with no judgment call. What it would not catch: every one of those four retrieved
sources being wrong for the question, or the drafter and the checker sharing a blind spot,
because they are typically driven by the same kind of model and can fail on the same kind of
question the same way. That is the case for [review and
debate](/gradient_ascent/techniques/debate-review/) instead, where the second opinion is built to differ from the first on purpose.

`examples/evaluator_optimizer/run.py` (lines 22-107)

```python
LEVEL = 3
RETRIEVE_K = 4
MAX_REVISIONS = 2
PASS_TOKEN = "ALL CITATIONS SUPPORTED"
DRAFT_SYSTEM = (
    "You answer questions about Halvorsen appliances using only the numbered sources below. End "
    "your answer with a line starting 'Sources:' listing the citations, like 'dw300-manual#3', "
    "that you used."
)
CHECK_SYSTEM = (
    "You check a draft answer against the source passages it was given. List every citation the "
    "draft claims that does NOT actually appear among the sources below, one per line, as "
    f"'MISSING: <citation>'. If every citation the draft claims is one of the sources, reply with "
    f"exactly '{PASS_TOKEN}' and nothing else."
)
REVISE_SYSTEM = (
    "Revise your previous answer to fix the citation problems named below. Use only the sources "
    "given. Keep the same 'Sources:' line format."
)

def _sources_block(sources: list[Section]) -> str:
    return "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)

def _draft(question: str, sources: list[Section], model: Model, tracer: Tracer) -> str:
    prompt = f"Sources:\n\n{_sources_block(sources)}\n\nQuestion: {question}"
    completion = model.complete([Message(role="system", content=DRAFT_SYSTEM), Message(role="user", content=prompt)], max_tokens=400)
    tracer.record(
        kind="model", decided_by="code", title="Draft an answer", detail=completion.text[:200],
        tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )
    return completion.text

def _check(draft_text: str, sources: list[Section], model: Model, tracer: Tracer) -> str | None:
    """None means the draft passed. Otherwise, the checker's own feedback text."""
    prompt = f"Sources:\n\n{_sources_block(sources)}\n\nDraft answer:\n{draft_text}"
    completion = model.complete([Message(role="system", content=CHECK_SYSTEM), Message(role="user", content=prompt)], max_tokens=200)
    verdict = completion.text.strip()
    tracer.record(
        kind="model", decided_by="code", title="Check citations against the sources", detail=verdict[:200],
        tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )
    return None if PASS_TOKEN in verdict.upper() else verdict

def _revise(question: str, draft_text: str, feedback: str, sources: list[Section], model: Model, tracer: Tracer) -> str:
    prompt = f"Sources:\n\n{_sources_block(sources)}\n\nQuestion: {question}\n\nPrevious answer:\n{draft_text}\n\nProblems to fix:\n{feedback}"
    completion = model.complete([Message(role="system", content=REVISE_SYSTEM), Message(role="user", content=prompt)], max_tokens=400)
    tracer.record(
        kind="model", decided_by="code", title="Revise using the checker's feedback", detail=completion.text[:200],
        tokens_in=completion.tokens_in, tokens_out=completion.tokens_out, ms=completion.ms,
    )
    return completion.text

def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    max_revisions: int = MAX_REVISIONS,
) -> Answer:
    del embedder  # retrieval here is keyword search, not a vector index
    sections = load_sections(corpus_dir)
    sources = [s for s, score in bm25_search(sections, question, k=RETRIEVE_K) if score > 0]
    tracer.record(kind="code", decided_by="code", title="Retrieve sources", detail=", ".join(s.cite for s in sources) or "none")

    draft_text = _draft(question, sources, model, tracer)
    feedback = _check(draft_text, sources, model, tracer)
    revisions = 0
    while feedback is not None and revisions < max_revisions:
        draft_text = _revise(question, draft_text, feedback, sources, model, tracer)
        revisions += 1
        feedback = _check(draft_text, sources, model, tracer)
    if feedback is not None:
        tracer.record(
            kind="code", decided_by="code", title="Stop: revision cap reached",
            detail=f"shipping a draft that still fails its own check after {revisions} revision(s)",
        )

    citations = cited_sources(draft_text)
    return Answer(text=draft_text, citations=citations, retrieved_sources=[s.cite for s in sources])
```

Run it yourself:

`examples/evaluator_optimizer/README.md` (lines 16-16)

```text
python -m examples.evaluator_optimizer --model stub:scripted
```

The same shape (a metric-driven loop instead of a vibe-driven one) is what DSPy's optimizers do
to a prompt itself, at build time rather than at answer time: "All optimizers read a numeric
score per example," and each one "tunes one or more of: instructions, demos, or weights"[2] to raise that score. DSPy loops over many training examples to improve the *prompt*
before it ever answers a real question; this page's loop runs once, at answer time, to improve
one *answer*. Both need the same thing to work at all: a criterion specific enough that two runs
of the check agree.

The [design review checklist](/gradient_ascent/recipes/design-review-checklist/) recipe is the
worked version of this loop for engineering test and precise measurement alike: one pass drafts a
finding against a design-review rule's own text, a second checks that finding against the rule it
cites and drops any that name none. Neither pass reports a measurement or a margin; where a rule
is a number against a threshold, code computes it, and a person still decides whether an unmet
rule ships or gets fixed.

## When you do not need this

Try a single call, or a fixed, code-only check like the one [prompt chaining](/gradient_ascent/techniques/prompt-chaining/)'s example uses (comparing citations by
set intersection, no second model call), first if the thing you would check for is something
plain code can already test: that is cheaper, always consistent, and does not need a second
prompt at all.

Move up to write and check once the failure you are trying to catch needs judgment against a
written rule that plain code cannot express directly, but that a second prompt, told the rule in
so many words, can apply consistently. This is the site's
[draft and check](/gradient_ascent/shapes/#draft-and-check) shape.

## Failure modes

### A checker with no fixed criterion

- **How to notice it:** The checker’s verdict changes between two runs on the same draft, because it was asked something open-ended ("is this good") rather than something specific enough to answer the same way twice.
- **How to test for it:** Run the check step on the exact same draft and sources twice. A checker worth looping on returns the same verdict both times; one that does not is adding cost without adding reliability.

### Writer and checker share a blind spot

- **How to notice it:** The checker passes a draft that is confidently wrong in a way neither prompt would ever catch, because both were built from the same kind of model making the same kind of mistake.
- **How to test for it:** Feed the checker a draft with a deliberate error of the kind its own criterion cannot see (a citation that is real and present, but supports the wrong fact) and confirm it passes, which is the specific gap review and debate exists to close.

### The cap ships a known-bad answer

- **How to notice it:** The loop reaches max_revisions still failing its own check, and the last draft goes out anyway, silently unless the "Stop: revision cap reached" step is actually surfaced somewhere a person or a downstream system can see it.
- **How to test for it:** Force a draft that can never pass (script the checker to always find fault) and confirm the run still returns an answer rather than hanging or raising, and that the stop is recorded, not just implied by running out of steps.

### The checker burns the whole cap on a trivial complaint

- **How to notice it:** A near-miss the checker treats as failing (a citation formatted slightly differently from what it expects) consumes the same revision budget as a genuine problem, leaving fewer chances left for anything that actually matters.
- **How to test for it:** Compare how many revisions a trivially-imperfect draft uses against how many a genuinely wrong one uses; if they are the same, the checker's criterion may be too literal to be worth a full revision cycle.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, best case (passes first check):** 2
- **Model calls, worst case (2 revisions):** 6
- **Tokens in, one check call:** ~350
- **Wall time, one revision cycle:** ~1.6s

**Compared with RAG (level 2), one call.** Cost here is not fixed per question the way RAG’s is: a question whose draft passes immediately costs about twice what RAG costs, and one that exhausts the cap costs up to three times that, for the same question.

## How to Evaluate It

Follow [the reviewer feedback loop](/gradient_ascent/examples/reviewer-feedback-loop/) to test a brief with a fresh receiving agent, review the proposal against original requirements, and revise within a fixed budget. It includes blind comparisons, parallel providers, and prompts for each role.

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

Scored on the same 60-question set as every other technique, with two numbers specific to this
loop: the average number of revisions a question used, and the share of questions that hit
`max_revisions` still failing their own check. A high cap-hit rate on real traffic is a sign the
checker's criterion is too strict for what the drafter can realistically satisfy, or that the
drafter has a systematic problem the checker keeps finding but the model cannot fix from feedback
alone.

No result file exists yet (see `docs/EVALS.md`), so this page cannot say what that rate actually
is here. Run `python scripts/eval_run.py --example evaluator_optimizer --model <spec> --dry` to
project the cost of a real run first: the projection matters more for this technique than most,
since its real cost depends on how often questions need a revision, not just how many questions
there are.

## Run it

**What to monitor.** The cap-hit rate over time (the share of runs that exhaust max_revisions still failing) and the average revisions per run. A cap-hit rate that rises with no change to the checker's prompt is a sign the questions arriving have shifted, not that the checker got stricter.

**Cost at volume.** Cost per question is variable here, unlike a fixed-step workflow: it depends on how often the checker fails the first draft. Budget for the worst case (every question uses the full cap), not the average, or a bad week of drafts becomes a bad week of the bill too.

**How it fails in production.** The writer and the checker share a blind spot neither prompt was built to catch, so confidently wrong answers pass their own check at the normal rate, with no signal in the trace that anything is different from a correctly-checked answer.

**What to log.** Every draft, every check verdict verbatim, and every revision, in order, for each question, not just the final answer. A checker that starts passing bad drafts is invisible in aggregate metrics until you can see its actual verdicts change.

## Try it

1. **Use it.** Find a tool that shows a checking or verifying step before its final answer. Does its documentation say what the check actually tests, or only that a check happens?
2. **Build it.** Run python -m examples.evaluator_optimizer --model stub:scripted from the repo root. The draft cites dw300-manual#6, the checker answers MISSING since retrieval never returned it, the revision cites one it did, and the check passes. Then edit CHECK_SYSTEM in examples/evaluator_optimizer/run.py to ask something vague ("Is this a good answer?") instead of that citation test, and say why it could not give the same verdict twice.
3. **Either lane.** Write the checking criterion you would use for work you review. Would two people applying it to the same draft reach the same verdict? If not, it is not a criterion yet.
4. **Build it.** Read DR-14 in the design-review-checklist recipe's design-review-rules.md next to this page's checker. Both catch one quotable, mechanical thing. Now find a rule there that needs reading rather than arithmetic, and say why a single checker pass is less trustworthy on it.


## Sources

1. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic, 2024-12-19 (accessed 2026-09-19)
2. [Optimizers: choosing one](https://dspy.ai/current/diving-deeper/choosing-an-optimizer/) — DSPy (Stanford NLP) (accessed 2026-09-19)


Last reviewed 2026-09-19.
