Drafting with a reviewer
One prompt drafts a piece of writing and another checks it against a rubric, repeating until the draft passes.
SourcedNeeds level 3
Every week a product team turns a handful of bullet points about what shipped into an update email a customer would actually want to read. Someone has to make sure the email covers everything that shipped, and separately, that it reads like the rest of the company’s writing: no unverified superlatives, no promises the bullets didn’t make.
Drafting with a reviewer splits those two checks apart, because they’re different kinds of question. Whether every bullet made it into the draft is something plain code can check for itself. Whether the draft’s tone matches the style guide is not: that needs a second prompt, told the rule, checking the first prompt’s work.
Example run
Optional: inspect the implementation trace
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Drafting with a reviewer, assembled
A fixed chain turns bullet points into an outline and checks it covers every one, then a separate checker loops on the draft until it passes a style rubric.
The run, step by step
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
The week's bullet points arrive
"- Shipped dark mode - Fixed the sync bug reported by several users - Raised the free-tier project limit from 3 to 5"
Walkthrough
The pipeline composes the runnable code already on the prompt chaining and evaluator-optimizer pages, in the order described above, with this job’s own outline-and-rubric content in place of their own citation-checking example. Neither example writes email in the repository; the diagram is their two loops chained and pointed at this job.
The run above starts with three bullets and an outline that covers all of them, so the chain’s own
gate passes without a second model call. Decide in advance what a dropped bullet does: the gate
can ask for the outline once more, or stop and hand the bullets back, but it must not pass a
draft that was never going to mention the thing that shipped. The draft that follows reads fine
on its surface but uses
“amazing” (a word the style rubric flags as an unverified superlative), and the checker’s one-word
verdict is specific enough that the revision fixes exactly that, nothing else, on the first try.
PASS_TOKEN is the entire contract the checker and the code share: the code never has to judge
what “good” means, only whether that exact token came back.
What to measure
Evaluator-optimizer and prompt chaining are each scored on their own by the site’s shared
60-question set (see docs/EVALS.md); this job needs its own set instead, since it drafts an
email rather than answering a question about documents. Build one from real or synthetic bullet
lists with a known-good outline and a small style rubric written down in advance. Three numbers:
outline coverage (does every bullet’s subject appear, the same test the chain’s own gate runs);
the average revisions per draft and the share that hit the cap still failing the rubric, the two
numbers evaluator-optimizer’s own page tracks; and, run twice on the same draft, whether the
checker’s verdict agrees with itself: a checker that doesn’t is adding cost without adding
reliability, which its own page warns against directly. No result file exists for either
technique on this task yet, so this recipe cannot claim a score for any of it.
Variations
- Add a second, unrelated checker for a second rubric (length, a required legal footer) rather than widening one prompt’s criteria; two narrow checks stay more testable than one broad one.
- Move to review and debate once the same kind of error keeps passing the checker: evidence the writer and the checker share a blind spot rather than that the rubric needs one more rule.
- Route by content type first with routing if the team drafts more than one kind of message (a release note, a status update) needing a different outline shape and a different rubric.
- Add human approval as a last step before sending, for the first several weeks, until the rubric has actually been checked against enough real drafts to trust unattended.
Design choices
Why this level, and when to use another approach
Two techniques compose this recipe, in sequence rather than as alternatives. Prompt chaining runs first: rewrite the bullets into a short outline, then check in plain code that every bullet’s subject actually appears in it: a set-membership test, the same shape prompt chaining’s own citation check uses, needing no second model call because the thing being checked is objective. Write and check runs second, once there’s a full draft: a separate prompt checks it against the style rubric, and if it fails, the draft revises and the check runs again, up to a fixed cap.
The split matters because the two checks need different tools. Evaluator-optimizer’s own page makes the case for using plain code wherever a check can be code: it’s cheaper, always consistent, and doesn’t need a second prompt at all. Whether the outline mentions “sync bug” is exactly that kind of check. Whether a sentence reads as an unapproved superlative is not. No fixed string test tells “amazing” from a plain factual claim in general, so that check needs a criterion specific enough for a second prompt to apply consistently, which is what write and check is for.
Both stay level 3 because the code owns the loop in each case: a fixed number of chain steps, and
a while condition and a revision cap written before either prompt runs. Handing either loop to
the model (let it decide whether another revision is worth the cost, or how many outline passes
to try) is what would raise this to level 5, and nothing here needs that: the criteria are fixed
in advance, and a checker told a specific rule answers the same way on the same draft every time.
The one real risk at this level is the one evaluator-optimizer’s own page names directly: a writer
and a checker built from the same kind of model can share a blind spot, passing a draft that’s
confidently wrong in a way neither prompt would catch. Its own upgrade path is exactly that
failure: move to review and debate once one
reviewer shares the author’s blind spots, when a second opinion built to differ on purpose earns
its cost.
Techniques this recipe uses
The highest level it needs is level 3.
Write and check
SourcedOne prompt writes, another checks, and the loop repeats until the check passes.
Produce something that has to meet a standard, and check it before anyone sees it
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Marketing copy against brand and legal rules
- An instrument control script checked against the documented command set and run on a simulator
- A measurement report checked figure by figure against the numbers code computed
- Code against its tests
- A SQL query against the schema
- A report against a required template
- A test procedure against the requirement it verifies
Last reviewed 09/18/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page