A review surface is something a builder chooses to expose. A system that returns only a final
answer leaves a reviewer nothing but plausibility to judge; one that shows its sources and the
steps that produced them lets a reviewer check the parts that are actually checkable. Anthropic’s
own guidance for agent builders is blunt about where that effort should come from: even where
automated tests already ran, “human review remains crucial for ensuring solutions align with
broader system requirements.”[1] A test suite checks what it was written to check, not
whether the change was the right one to make.
examples/reviewing builds one small piece of a surface like that. Given a drafted answer and
the citations it names, it reports which figures the answer states are carried by no section it
cites, and which citations carry none of them. figures_in reads the answer’s own numbers.
Getting this loose is the whole difficulty: a checker that flags everything gets ignored exactly
the way an approval step does, and one that matches too eagerly reports clean when it should not.
examples/reviewing/run.py · lines 68–78
def figures_in(text: str) -> list[str]:
"""Every figure the text states, in one canonical spelling each, sorted and deduplicated."""
figures = []
for word in _PERCENT_RE.sub(r"\1%", text).split():
word = word.strip(_TRIM)
parts = [word] if _ISO_DATE_RE.match(word) else _RANGE_RE.split(word)
for part in parts:
figure = _figure(part)
if figure is not None:
figures.append(figure)
return sorted(set(figures))
_figure, the helper called on each token, is where the judgment sits. It compares figures as
values rather than as text, so $1,200 and 1200 are one figure and 52 is not a match for
1152. A substring search would have accepted this, reporting a citation as support for a
price it says nothing about. A token with a letter before its digits (HLV-2205, DW300,
v2.1, dw300-manual#3) states no quantity, so nothing is claimed about it. A range states both
of its ends; an ISO date is one figure rather than three.
examples/reviewing/run.py · lines 100–140
def run(answer: Answer, sections: dict[str, Section], tracer: Tracer) -> ReviewReport:
figures = figures_in(answer.text)
tracer.record(
kind="code",
decided_by="code",
title="Read the figures the answer states",
detail=", ".join(figures) or "none",
)
flags: list[Flag] = []
found_somewhere: set[str] = set()
for citation in answer.citations:
section = sections.get(citation)
if section is None:
tracer.record(
kind="code", decided_by="code", title=f"Open {citation}", detail="not in the corpus"
)
flags.append(Flag(citation, "cited section does not exist"))
continue
here = sorted(set(figures) & set(figures_in(section.text)))
found_somewhere.update(here)
tracer.record(
kind="code",
decided_by="code",
title=f"Open {citation}",
detail=f"{section.title}: {', '.join(here) or 'no claimed figure'}",
)
if figures and not here:
flags.append(Flag(citation, "section carries none of the answer's figures"))
for figure in figures:
if figure not in found_somewhere:
flags.append(Flag(figure, "figure appears in no cited section"))
tracer.record(
kind="code",
decided_by="code",
title="Report",
detail=f"{len(flags)} thing(s) to look at across {len(answer.citations)} citation(s)",
)
return ReviewReport(figures_claimed=figures, checked=list(answer.citations), flags=flags)
This is a presence check, not a truth check. It says a number appears in the text the answer
points at. It does not say the section supports the claim, that the right sources were chosen,
or that the answer is complete, and two limits are pinned as tests rather than hidden: units are
dropped, so a figure can match with the wrong unit, and a date written in prose will not match
the same date written 2026-09-18. No model is called anywhere in it, so every step is
decided_by: "code", and the 60-question set does not score it: it answers no question about
the corpus, it checks an answer someone else produced. The number worth tracking is its own
flag rate on real output, which is a claim about that system, not about models in general.