# Reviewing work you did not do

_Topics at every level · sourced_

Checking work you did not do yourself before it goes anywhere.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a generated artifact through checks targeted at its claims and consequences. Inspect why readable prose or plausible code should not receive the same review as an independently verified result.

**Assumptions:** Review effort is limited. The reviewer needs the source evidence or expected behavior for the parts they are asked to approve.

**Design choices:** Check important facts, calculations, omissions, and commitments first. Use deterministic checks where available and judgment where purpose or ambiguity matters.

**Request:** Check this budget and announcement before sharing.

**Starting evidence:** Room $120 + materials $90 + snacks $40; draft total $230. Tools promised but unconfirmed.

**Action and control:** Recompute totals and verify claims before polishing style.

**Stage records (authored, not executed):**

### Input record

Room $120 + materials $90 + snacks $40; draft total $230. Tools promised but unconfirmed.

What changed: Establish the facts supplied for this version of the task.

### Design note

Check important facts, calculations, omissions, and commitments first. Use deterministic checks where available and judgment where purpose or ambiguity matters.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Recompute totals and verify claims before polishing style.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Total is $250. Tools claim unresolved; retain draft status.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Source-backed fact checks, recomputed totals, marked corrections, and a final review decision.

If the result falls short:
When one claim fails, inspect related assumptions and correct the source of the error. Do not discard sound work automatically or approve the remainder without thought.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for reports, budgets, code, or plans. Scale review to consequence and uncertainty; a rough private draft needs less scrutiny than a shared operational decision.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Total is $250. Tools claim unresolved; retain draft status.

**Change something — Check only grammar and tone:** The $20 error and unsupported tools promise survive.

**Decision:** What should precede cosmetic editing?

**Answer:** Verify totals and operational claims.

**Why:** Fluent writing can hide bad totals and invented facts; prioritize consequential errors rather than cosmetic edits.

**Review criteria:** Source-backed fact checks, recomputed totals, marked corrections, and a final review decision.

**Recovery:** When one claim fails, inspect related assumptions and correct the source of the error. Do not discard sound work automatically or approve the remainder without thought.

**Adapt it:** Use this for reports, budgets, code, or plans. Scale review to consequence and uncertainty; a rough private draft needs less scrutiny than a shared operational decision.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a generated artifact through checks targeted at its claims and consequences. Inspect why readable prose or plausible code should not receive the same review as an independently verified result.

**Assumptions:** Review effort is limited. The reviewer needs the source evidence or expected behavior for the parts they are asked to approve.

**Design choices:** Check important facts, calculations, omissions, and commitments first. Use deterministic checks where available and judgment where purpose or ambiguity matters.

**Request:** Review a generated measurement report before accepting it.

**Starting evidence:** Table labels voltage in V, but one source file is in mV. Arithmetic appears internally consistent.

**Action and control:** Check units and source transformations independently, then inspect conclusions and requirements coverage.

**Stage records (authored, not executed):**

### Input record

Table labels voltage in V, but one source file is in mV. Arithmetic appears internally consistent.

What changed: Establish the facts supplied for this version of the task.

### Design note

Check important facts, calculations, omissions, and commitments first. Use deterministic checks where available and judgment where purpose or ambiguity matters.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Check units and source transformations independently, then inspect conclusions and requirements coverage.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Flag the mixed-unit result; normalize and recompute before accepting the pass/fail conclusion.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Recompute sample rows and trace each unit conversion to its source.

If the result falls short:
When one claim fails, inspect related assumptions and correct the source of the error. Do not discard sound work automatically or approve the remainder without thought.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for reports, budgets, code, or plans. Scale review to consequence and uncertainty; a rough private draft needs less scrutiny than a shared operational decision.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Flag the mixed-unit result; normalize and recompute before accepting the pass/fail conclusion.

**Change something — Review only whether the script ran without errors:** The unit error survives. Runtime success does not establish a valid measurement calculation.

**Decision:** Can successful execution replace engineering review?

**Answer:** No; check units and requirement interpretation.

**Why:** Independent review should target consequential semantic errors, not only syntax.

**Review criteria:** Recompute sample rows and trace each unit conversion to its source.

**Recovery:** When one claim fails, inspect related assumptions and correct the source of the error. Do not discard sound work automatically or approve the remainder without thought.

**Adapt it:** Use this for reports, budgets, code, or plans. Scale review to consequence and uncertainty; a rough private draft needs less scrutiny than a shared operational decision.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a generated artifact through checks targeted at its claims and consequences. Inspect why readable prose or plausible code should not receive the same review as an independently verified result.

**Assumptions:** Review effort is limited. The reviewer needs the source evidence or expected behavior for the parts they are asked to approve.

**Design choices:** Check important facts, calculations, omissions, and commitments first. Use deterministic checks where available and judgment where purpose or ambiguity matters.

**Request:** Review the weekly report before approving distribution.

**Starting evidence:** Draft says Beacon complete. Source says development complete, acceptance testing pending.

**Action and control:** Compare report language with the precise source status and identify omitted qualifications.

**Stage records (authored, not executed):**

### Input record

Draft says Beacon complete. Source says development complete, acceptance testing pending.

What changed: Establish the facts supplied for this version of the task.

### Design note

Check important facts, calculations, omissions, and commitments first. Use deterministic checks where available and judgment where purpose or ambiguity matters.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Compare report language with the precise source status and identify omitted qualifications.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Revise to development complete; acceptance pending. Do not approve a broader completion claim.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Inspect source-backed status, omitted qualifiers, recipients, and approval version.

If the result falls short:
When one claim fails, inspect related assumptions and correct the source of the error. Do not discard sound work automatically or approve the remainder without thought.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for reports, budgets, code, or plans. Scale review to consequence and uncertainty; a rough private draft needs less scrutiny than a shared operational decision.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Revise to development complete; acceptance pending. Do not approve a broader completion claim.

**Change something — Review only the executive summary's tone:** The misleading completion statement remains even if the summary reads well.

**Decision:** Is development complete equivalent to project accepted?

**Answer:** No; preserve the distinct completion criteria.

**Why:** Review must preserve distinctions that affect decisions, especially in compressed summaries.

**Review criteria:** Inspect source-backed status, omitted qualifiers, recipients, and approval version.

**Recovery:** When one claim fails, inspect related assumptions and correct the source of the error. Do not discard sound work automatically or approve the remainder without thought.

**Adapt it:** Use this for reports, budgets, code, or plans. Scale review to consequence and uncertainty; a rough private draft needs less scrutiny than a shared operational decision.

Reviewing is checking work you did not do yourself before it goes anywhere. It is a different
skill from doing the work, and it gets harder rather than easier as a system does more on its
own, because more of the work happened somewhere you were not watching.

Two things decide whether a piece of work can be reviewed at all. There has to be something to
check it against (a source, a record, a number you can recompute) and enough of the work has
to be visible to check, not just its conclusion. When neither is true, "review" means reading
something plausible and agreeing with it.

How much review a result needs is not a fixed amount. It scales with what a wrong answer would
cost and how long the mistake would sit there before anyone noticed. That is the same question
[delegating](/gradient_ascent/techniques/delegating/) asks before handing a task over at all,
asked again about the thing that came back; [calibrating
trust](/gradient_ascent/techniques/trust/) is what the answers to it, kept over time, add up to. All four skills on the
[operator craft](/gradient_ascent/techniques/operator-craft/) topic start from what the request
said, which is [briefing](/gradient_ascent/techniques/briefing/).

This page is sourced, not measured: the checks below are drawn from primary sources, and no
result file exists for any of them, so no number here is one this site took.

## Practical guidance

Reviewing one answer starts with a request you can type back to whatever produced it: "List
every number, date, name and citation you used, and tell me exactly where each one came from."
That turns a paragraph you'd judge by feel into a list you can check line by line: for each item,
either the source it points to says it or it doesn't. When it can't produce one, or the source
named doesn't actually contain the figure, that's the failure, and the usual cause is that it
summarized or estimated instead of quoting.

Then three passes, in order: the first two need no expertise, the third needs all of yours.

1. Open the source for each item and confirm it says what the answer says. Do not accept a
   summary of the source; open the document, email or page itself.
2. Check what's missing against what you asked for. If you asked for five points and got four,
   nothing in the answer will flag the gap.
3. Read for judgment: is this the right answer to the right question? This is the part only you
   can do.

**Reviewing a day of unattended work** is different: nothing can be stopped mid-run, so there is
only a record. Read the actions taken before any summary, irreversible ones first (sent, deleted,
paid, published), and open two or three at random to confirm the record matches what was done.

Fluency is not evidence. A fabricated answer reads exactly as well as a correct one, and a 2025
paper by researchers at OpenAI and Georgia Tech argues this is structural: models "hallucinate
because the training and evaluation procedures reward guessing over acknowledging uncertainty,"
because "language models are optimized to be good test-takers, and guessing when uncertain
improves test performance."[2]

What erodes first is attention, not trust. A 2019 survey puts the core of automation complacency
at "the degree of attention devoted to monitoring automated tasks (specifically, the lack
thereof)"[3]. A string of clean checks is exactly when your attention is most likely to
slip, which argues for a fixed sample rate over trusting that today's batch looks fine.

Skip the three passes on something low-stakes you'll reread yourself anyway. And when there's
nothing to check against, no source, no record, no number you can recompute, that isn't a review
you can do; ask for the source before you sign off, not after.

## Implementation details

A review surface is something a builder chooses to expose. A system that returns only a final
answer leaves a reviewer nothing but plausibility to judge; one that shows its sources and the
steps that produced them lets a reviewer check the parts that are actually checkable. Anthropic's
own guidance for agent builders is blunt about where that effort should come from: even where
automated tests already ran, "human review remains crucial for ensuring solutions align with
broader system requirements."[1] A test suite checks what it was written to check, not
whether the change was the right one to make.

`examples/reviewing` builds one small piece of a surface like that. Given a drafted answer and
the citations it names, it reports which figures the answer states are carried by no section it
cites, and which citations carry none of them. `figures_in` reads the answer's own numbers.
Getting this loose is the whole difficulty: a checker that flags everything gets ignored exactly
the way an approval step does, and one that matches too eagerly reports clean when it should not.

`examples/reviewing/run.py` (lines 68-78)

```python
def figures_in(text: str) -> list[str]:
    """Every figure the text states, in one canonical spelling each, sorted and deduplicated."""
    figures = []
    for word in _PERCENT_RE.sub(r"\1%", text).split():
        word = word.strip(_TRIM)
        parts = [word] if _ISO_DATE_RE.match(word) else _RANGE_RE.split(word)
        for part in parts:
            figure = _figure(part)
            if figure is not None:
                figures.append(figure)
    return sorted(set(figures))
```

`_figure`, the helper called on each token, is where the judgment sits. It compares figures as
values rather than as text, so `$1,200` and `1200` are one figure and `52` is not a match for
`1152`. A substring search would have accepted this, reporting a citation as support for a
price it says nothing about. A token with a letter before its digits (`HLV-2205`, `DW300`,
`v2.1`, `dw300-manual#3`) states no quantity, so nothing is claimed about it. A range states both
of its ends; an ISO date is one figure rather than three.

`examples/reviewing/run.py` (lines 100-140)

```python
def run(answer: Answer, sections: dict[str, Section], tracer: Tracer) -> ReviewReport:
    figures = figures_in(answer.text)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Read the figures the answer states",
        detail=", ".join(figures) or "none",
    )

    flags: list[Flag] = []
    found_somewhere: set[str] = set()
    for citation in answer.citations:
        section = sections.get(citation)
        if section is None:
            tracer.record(
                kind="code", decided_by="code", title=f"Open {citation}", detail="not in the corpus"
            )
            flags.append(Flag(citation, "cited section does not exist"))
            continue
        here = sorted(set(figures) & set(figures_in(section.text)))
        found_somewhere.update(here)
        tracer.record(
            kind="code",
            decided_by="code",
            title=f"Open {citation}",
            detail=f"{section.title}: {', '.join(here) or 'no claimed figure'}",
        )
        if figures and not here:
            flags.append(Flag(citation, "section carries none of the answer's figures"))

    for figure in figures:
        if figure not in found_somewhere:
            flags.append(Flag(figure, "figure appears in no cited section"))

    tracer.record(
        kind="code",
        decided_by="code",
        title="Report",
        detail=f"{len(flags)} thing(s) to look at across {len(answer.citations)} citation(s)",
    )
    return ReviewReport(figures_claimed=figures, checked=list(answer.citations), flags=flags)
```

This is a presence check, not a truth check. It says a number appears in the text the answer
points at. It does not say the section supports the claim, that the right sources were chosen,
or that the answer is complete, and two limits are pinned as tests rather than hidden: units are
dropped, so a figure can match with the wrong unit, and a date written in prose will not match
the same date written `2026-09-18`. No model is called anywhere in it, so every step is
`decided_by: "code"`, and the 60-question set does not score it: it answers no question about
the corpus, it checks an answer someone else produced. The number worth tracking is its own
flag rate on real output, which is a claim about that system, not about models in general.

## At each level

- [Conventional software](/gradient_ascent/levels/0/): there is no output to review, only code to test.
- [Direct prompting](/gradient_ascent/levels/1/): one reply, reviewed against what you already know or
  can look up in the time you were willing to spend.
- [Added context](/gradient_ascent/levels/2/): the sources are on screen, so the cheap checks become
  possible, and a citation that is present is not yet a citation that supports the sentence.
- [Workflows](/gradient_ascent/levels/3/): the pipeline is fixed, so you can decide once, in
  advance, which step is worth a person's eyes rather than deciding per result.
- [Tool use](/gradient_ascent/levels/4/): a proposed action can be reviewed before it runs, which
  is a cheaper and more useful check than reading about it afterwards.
- [Agent loops](/gradient_ascent/levels/5/): nobody reads the whole run, so review becomes a sample
  plus the final result, and the sample rate is now a number someone has to choose.
- [Teams of Agents](/gradient_ascent/levels/6/): several outputs agree with each other, which
  is not evidence: they can be wrong together, and often from the same starting assumption.
- [Always-on agents](/gradient_ascent/levels/7/): review is entirely after the fact, so the
  question becomes what would have to change to catch this before the next run, not this one.

## Practices

- Check the specific claims first (figures, dates, names, citations) then what is missing, then
  the judgment. The order matters: the first two are fast and the third is where your expertise
  actually earns its keep.
- Open the source, not a summary of the source written by the system whose work you are checking.
- Set the sample rate before the week starts, not while reading the output.
- Read the record of what was done before any summary of what was done, and the irreversible
  actions before the rest.
- Count the flags and the skips. A run with none of either is a reason to look harder.

## Run it

**What to monitor.** The share of problems found by a reviewer versus found later by someone downstream:
  a customer, an auditor, the person who received the output. That ratio moving the wrong way
  says review is missing things, well before anyone complains about quality.

**Cost at volume.** Reviewing every result costs a fixed amount per result and does not survive
  growth. What does survive is a machine check on the parts that are mechanical (a figure that
  must appear in a cited source, a required field that must be filled) plus a sample of the rest,
  which keeps a person's time on the judgment that no check can make.

**How it fails in production.** The review step is still in the process and nobody is doing it. Approvals
  come back in seconds, flags stop appearing, and the record shows a hundred consecutive clean
  results. That is what both a reliable system and an unread queue look like.

**What to log.** For every reviewed item: what was changed, rejected or flagged, and why. That log
  is what tells the two cases above apart, and it is the raw material a trust record is built
  from.

## Try it

1. **Use it.** Take the last piece of model-generated work you accepted without much scrutiny. Pull out three specific claims (a figure, a date, a citation) and check each against its source. Note how long it took; that number is what a review costs you.
2. **Build it.** From the repo root run python -m examples.reviewing --scenario clean, then --scenario mismatch, then --scenario missing, and read the three reports. Then add a scenario of your own in examples/reviewing/__main__.py whose answer spells a corpus figure differently ("the drain pump costs 52 dollars" against a list price written $52.00) and confirm it still reports clean. Then break it on purpose: change 52 to 5.2 and read the two flags that come back.
3. **Either lane.** Write down the sample rate you actually use on work you are supposed to be reviewing: one in three, one in ten, whatever it honestly is. Then write down the rate you would defend to someone whose money or safety depends on it. If the two differ, one of them is wrong.


## Sources

1. [Building Effective AI Agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic (accessed 2026-09-19)
2. [Why Language Models Hallucinate](https://arxiv.org/abs/2509.04664) — arXiv (Kalai, Nachum, Vempala and Zhang; OpenAI and Georgia Tech), 2025-09-04 (accessed 2026-09-19)
3. [Automation-Induced Complacency Potential: Development and Validation of a New Scale](https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2019.00225/full) — Frontiers in Psychology (Merritt et al.), 2019-02-19 (accessed 2026-09-19)


Last reviewed 2026-09-19.
