Thread

Checking the work

How you tell whether it worked, from a person reading one answer to a scored set and a recorded trace. The check changes shape at every level; the question does not.

The organizing sentenceEvery level has a way to be wrong that the level below could not be, and a check that costs less than the mistake.
Compare the mechanisms

Three questions. Three different kinds of evidence.

These checks complement each other; they are not a progression from weak to strong.

Is the record valid?

  1. 1Returned data
  2. 2Schema + deterministic rules
  3. 3Accept or reject the record

Check types, required fields, arithmetic, permissions, and source IDs.

BoundaryA valid record can contain false claims.

Did the task succeed?

  1. 1Task + outcome
  2. 2Rubric or external check
  3. 3Quality judgment

Grade usefulness, factual support, completeness, and unwanted effects.

BoundaryA model grader needs calibration; agreement is not proof.

Where did it go wrong?

  1. 1Observed run events
  2. 2Trace + artifact inspection
  3. 3Failure diagnosis

Follow the evidence, tool calls, state changes, and errors behind an outcome.

BoundaryTelemetry alone does not establish answer quality.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Checking the work: see it in practice.

Matching verification methods to the kinds of errors a system can make.

What you’ll walk through

Follow a convincing result through different kinds of verification. Separate checking the format, the supporting evidence, and the effect of any action.

The task in this version

Check a well-formatted but wrong support answer.

What you’ll learn to check

Layered checks with the exact defect each detects, a remaining blind spot, and an appropriate human escalation.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Check a well-formatted but wrong support answer.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Valid fields and citation. Citation gives duration; answer claims water damage is covered.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Each check establishes something specific. Valid formatting does not prove a factual claim, and an approved plan does not prove successful execution.

1 / 6

Apply this to your project

Describe your task to your own model and use Checking the work as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

A convincing response is not the same as a successful task. Check the returned record, the meaning of the answer, and—when tools act—the state left in the environment.

Keep the checks distinct

Check What it can establish What it cannot establish alone
Schema and deterministic rules Required fields, types, arithmetic, allowed values Truth or completeness of arbitrary prose
Source review Whether evidence supports a claim Whether retrieval found every relevant source
A model reviewer A rubric-based judgment or proposed critique Correctness merely because it is another call
An evaluation Performance under defined task conditions Performance outside those conditions
Tracing Recorded calls, evidence, actions, and errors A quality judgment without a criterion

Check the actual outcome

Inspect the artifact or external state that should result from the task. A success message can be wrong. Refusal or clarification can be the correct outcome when evidence or authority is missing.

Evaluations can grade one run, estimate performance across cases, or compare versions. Graders can be code, people, models, or a combination.

Make comparisons meaningful

Separate prompt-development cases from held-out evaluation. Repeat tasks when variation matters. Inspect failures as well as averages, and record model, configuration, prompts, tools, data, and harness version.

If several things change together, an improved score does not identify which change helped.

Reviewers can share blind spots

A reviewer using the same incorrect source can reinforce a mistake. Separate agents may have separate context without statistically independent errors. Complement model review with checkable constraints and independently verified evidence.

Calibrate model graders against human judgments. Disagreements are cases to inspect, not just scores to combine.

Diagnose the failure you have

If the needed passage never entered context, inspect retrieval and access. If it was present but contradicted, inspect generation. If the proposed action was correct but the tool failed, inspect execution and recovery.

Record useful events and artifacts with appropriate redaction and retention. Observability helps locate a failure; evaluation defines why it counts as one.

Where this comes from

Primary sources

  1. Demystifying evals for AI agents · Anthropic (accessed 09/20/2026)

Last reviewed 09/20/2026. Pages unreviewed for 90 days are flagged for another pass.