Topics at every level

Reviewing work you did not do

Checking work you did not do yourself before it goes anywhere.

Sourced

Concept at a glance

Trace claims back to something you can check.

SequenceConceptual illustration
Trace claims back to something you can check.Proposed work leads to Independent checks. Independent checks leads to Your decision. Review the evidence and the result, not just how convincing the prose sounds.Proposed workAn answer, edit, or actionIndependent checksSources, tests, constraintsYour decisionAccept, correct, or rejectTrace claims back to something you can check.Proposed work leads to Independent checks. Independent checks leads to Your decision. Review the evidence and the result, not just how convincing the prose sounds.Proposed workAn answer, edit, or actionIndependent checksSources, tests, constraintsYour decisionAccept, correct, or reject
Read the connections in words
  • Proposed work → Independent checks: Sources, tests, constraints.
  • Independent checks → Your decision: Accept, correct, or reject.
Key idea

Review the evidence and the result, not just how convincing the prose sounds.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Reviewing work you did not do: see it in practice.

Independently checking another party's work before relying on or releasing it.

What you’ll walk through

Follow a generated artifact through checks targeted at its claims and consequences. Inspect why readable prose or plausible code should not receive the same review as an independently verified result.

The task in this version

Check this budget and announcement before sharing.

What you’ll learn to check

Source-backed fact checks, recomputed totals, marked corrections, and a final review decision.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example

Check this budget and announcement before sharing.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Room $120 + materials $90 + snacks $40; draft total $230. Tools promised but unconfirmed.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Review effort is limited. The reviewer needs the source evidence or expected behavior for the parts they are asked to approve.

1 / 6

Apply this to your project

Describe your task to your own model and use Reviewing work you did not do as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Reviewing is checking work you did not do yourself before it goes anywhere. It is a different skill from doing the work, and it gets harder rather than easier as a system does more on its own, because more of the work happened somewhere you were not watching.

Two things decide whether a piece of work can be reviewed at all. There has to be something to check it against (a source, a record, a number you can recompute) and enough of the work has to be visible to check, not just its conclusion. When neither is true, “review” means reading something plausible and agreeing with it.

How much review a result needs is not a fixed amount. It scales with what a wrong answer would cost and how long the mistake would sit there before anyone noticed. That is the same question delegating asks before handing a task over at all, asked again about the thing that came back; calibrating trust is what the answers to it, kept over time, add up to. All four skills on the operator craft topic start from what the request said, which is briefing.

This page is sourced, not measured: the checks below are drawn from primary sources, and no result file exists for any of them, so no number here is one this site took.

Practical guidance

Reviewing one answer starts with a request you can type back to whatever produced it: “List every number, date, name and citation you used, and tell me exactly where each one came from.” That turns a paragraph you’d judge by feel into a list you can check line by line: for each item, either the source it points to says it or it doesn’t. When it can’t produce one, or the source named doesn’t actually contain the figure, that’s the failure, and the usual cause is that it summarized or estimated instead of quoting.

Then three passes, in order: the first two need no expertise, the third needs all of yours.

  1. Open the source for each item and confirm it says what the answer says. Do not accept a summary of the source; open the document, email or page itself.
  2. Check what’s missing against what you asked for. If you asked for five points and got four, nothing in the answer will flag the gap.
  3. Read for judgment: is this the right answer to the right question? This is the part only you can do.

Reviewing a day of unattended work is different: nothing can be stopped mid-run, so there is only a record. Read the actions taken before any summary, irreversible ones first (sent, deleted, paid, published), and open two or three at random to confirm the record matches what was done.

Fluency is not evidence. A fabricated answer reads exactly as well as a correct one, and a 2025 paper by researchers at OpenAI and Georgia Tech argues this is structural: models “hallucinate because the training and evaluation procedures reward guessing over acknowledging uncertainty,” because “language models are optimized to be good test-takers, and guessing when uncertain improves test performance.”[2]

What erodes first is attention, not trust. A 2019 survey puts the core of automation complacency at “the degree of attention devoted to monitoring automated tasks (specifically, the lack thereof)”[3]. A string of clean checks is exactly when your attention is most likely to slip, which argues for a fixed sample rate over trusting that today’s batch looks fine.

Skip the three passes on something low-stakes you’ll reread yourself anyway. And when there’s nothing to check against, no source, no record, no number you can recompute, that isn’t a review you can do; ask for the source before you sign off, not after.

Implementation details

A review surface is something a builder chooses to expose. A system that returns only a final answer leaves a reviewer nothing but plausibility to judge; one that shows its sources and the steps that produced them lets a reviewer check the parts that are actually checkable. Anthropic’s own guidance for agent builders is blunt about where that effort should come from: even where automated tests already ran, “human review remains crucial for ensuring solutions align with broader system requirements.”[1] A test suite checks what it was written to check, not whether the change was the right one to make.

examples/reviewing builds one small piece of a surface like that. Given a drafted answer and the citations it names, it reports which figures the answer states are carried by no section it cites, and which citations carry none of them. figures_in reads the answer’s own numbers. Getting this loose is the whole difficulty: a checker that flags everything gets ignored exactly the way an approval step does, and one that matches too eagerly reports clean when it should not.

examples/reviewing/run.py · lines 68–78
def figures_in(text: str) -> list[str]:
    """Every figure the text states, in one canonical spelling each, sorted and deduplicated."""
    figures = []
    for word in _PERCENT_RE.sub(r"\1%", text).split():
        word = word.strip(_TRIM)
        parts = [word] if _ISO_DATE_RE.match(word) else _RANGE_RE.split(word)
        for part in parts:
            figure = _figure(part)
            if figure is not None:
                figures.append(figure)
    return sorted(set(figures))

_figure, the helper called on each token, is where the judgment sits. It compares figures as values rather than as text, so $1,200 and 1200 are one figure and 52 is not a match for 1152. A substring search would have accepted this, reporting a citation as support for a price it says nothing about. A token with a letter before its digits (HLV-2205, DW300, v2.1, dw300-manual#3) states no quantity, so nothing is claimed about it. A range states both of its ends; an ISO date is one figure rather than three.

examples/reviewing/run.py · lines 100–140
def run(answer: Answer, sections: dict[str, Section], tracer: Tracer) -> ReviewReport:
    figures = figures_in(answer.text)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Read the figures the answer states",
        detail=", ".join(figures) or "none",
    )

    flags: list[Flag] = []
    found_somewhere: set[str] = set()
    for citation in answer.citations:
        section = sections.get(citation)
        if section is None:
            tracer.record(
                kind="code", decided_by="code", title=f"Open {citation}", detail="not in the corpus"
            )
            flags.append(Flag(citation, "cited section does not exist"))
            continue
        here = sorted(set(figures) & set(figures_in(section.text)))
        found_somewhere.update(here)
        tracer.record(
            kind="code",
            decided_by="code",
            title=f"Open {citation}",
            detail=f"{section.title}: {', '.join(here) or 'no claimed figure'}",
        )
        if figures and not here:
            flags.append(Flag(citation, "section carries none of the answer's figures"))

    for figure in figures:
        if figure not in found_somewhere:
            flags.append(Flag(figure, "figure appears in no cited section"))

    tracer.record(
        kind="code",
        decided_by="code",
        title="Report",
        detail=f"{len(flags)} thing(s) to look at across {len(answer.citations)} citation(s)",
    )
    return ReviewReport(figures_claimed=figures, checked=list(answer.citations), flags=flags)

This is a presence check, not a truth check. It says a number appears in the text the answer points at. It does not say the section supports the claim, that the right sources were chosen, or that the answer is complete, and two limits are pinned as tests rather than hidden: units are dropped, so a figure can match with the wrong unit, and a date written in prose will not match the same date written 2026-09-18. No model is called anywhere in it, so every step is decided_by: "code", and the 60-question set does not score it: it answers no question about the corpus, it checks an answer someone else produced. The number worth tracking is its own flag rate on real output, which is a claim about that system, not about models in general.

At each level

  • Conventional software: there is no output to review, only code to test.
  • Direct prompting: one reply, reviewed against what you already know or can look up in the time you were willing to spend.
  • Added context: the sources are on screen, so the cheap checks become possible, and a citation that is present is not yet a citation that supports the sentence.
  • Workflows: the pipeline is fixed, so you can decide once, in advance, which step is worth a person’s eyes rather than deciding per result.
  • Tool use: a proposed action can be reviewed before it runs, which is a cheaper and more useful check than reading about it afterwards.
  • Agent loops: nobody reads the whole run, so review becomes a sample plus the final result, and the sample rate is now a number someone has to choose.
  • Teams of Agents: several outputs agree with each other, which is not evidence: they can be wrong together, and often from the same starting assumption.
  • Always-on agents: review is entirely after the fact, so the question becomes what would have to change to catch this before the next run, not this one.

Practices

  • Check the specific claims first (figures, dates, names, citations) then what is missing, then the judgment. The order matters: the first two are fast and the third is where your expertise actually earns its keep.
  • Open the source, not a summary of the source written by the system whose work you are checking.
  • Set the sample rate before the week starts, not while reading the output.
  • Read the record of what was done before any summary of what was done, and the irreversible actions before the rest.
  • Count the flags and the skips. A run with none of either is a reason to look harder.

Run it

What to monitor

The share of problems found by a reviewer versus found later by someone downstream: a customer, an auditor, the person who received the output. That ratio moving the wrong way says review is missing things, well before anyone complains about quality.

Cost at volume

Reviewing every result costs a fixed amount per result and does not survive growth. What does survive is a machine check on the parts that are mechanical (a figure that must appear in a cited source, a required field that must be filled) plus a sample of the rest, which keeps a person's time on the judgment that no check can make.

How it fails in production

The review step is still in the process and nobody is doing it. Approvals come back in seconds, flags stop appearing, and the record shows a hundred consecutive clean results. That is what both a reliable system and an unread queue look like.

What to log

For every reviewed item: what was changed, rejected or flagged, and why. That log is what tells the two cases above apart, and it is the raw material a trust record is built from.

Try it

  1. Use it

    Take the last piece of model-generated work you accepted without much scrutiny. Pull out three specific claims (a figure, a date, a citation) and check each against its source. Note how long it took; that number is what a review costs you.

  2. Build it

    From the repo root run python -m examples.reviewing --scenario clean, then --scenario mismatch, then --scenario missing, and read the three reports. Then add a scenario of your own in examples/reviewing/__main__.py whose answer spells a corpus figure differently ("the drain pump costs 52 dollars" against a list price written $52.00) and confirm it still reports clean. Then break it on purpose: change 52 to 5.2 and read the two flags that come back.

  3. Either lane

    Write down the sample rate you actually use on work you are supposed to be reviewing: one in three, one in ten, whatever it honestly is. Then write down the rate you would defend to someone whose money or safety depends on it. If the two differ, one of them is wrong.

How it connects

Before, after and instead of this

Optional: products, tools, and models

1 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Verify a measurement writeup

Compare every reported value and conclusion with the test log, calculations, and stated uncertainty.

Out there

Named products, tools and models

Products1
  • JulesGoogle · coding agent

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Building Effective AI Agents · Anthropic (accessed 09/19/2026)
  2. Why Language Models Hallucinate · arXiv (Kalai, Nachum, Vempala and Zhang; OpenAI and Georgia Tech), 09/04/2025 (accessed 09/19/2026)
  3. Automation-Induced Complacency Potential: Development and Validation of a New Scale · Frontiers in Psychology (Merritt et al.), 02/19/2019 (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page