# Checking the work

_Thread · sourced_

How you tell whether it worked, from a person reading one answer to a scored set and a recorded trace. The check changes shape at every level; the question does not.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a convincing result through different kinds of verification. Separate checking the format, the supporting evidence, and the effect of any action.

**Assumptions:** Each check establishes something specific. Valid formatting does not prove a factual claim, and an approved plan does not prove successful execution.

**Design choices:** Choose independent checks for consequential claims and actions. Combine automated validation with source review instead of asking the same model only whether it was right.

**Request:** Check a well-formatted but wrong support answer.

**Starting evidence:** Valid fields and citation. Citation gives duration; answer claims water damage is covered.

**Action and control:** Validate structure, claim support, and policy correctness separately.

**Stage records (authored, not executed):**

### Input record

Valid fields and citation. Citation gives duration; answer claims water damage is covered.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose independent checks for consequential claims and actions. Combine automated validation with source review instead of asking the same model only whether it was right.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Validate structure, claim support, and policy correctness separately.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Schema passes; evidence fails. Duration does not establish water-damage coverage.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Layered checks with the exact defect each detects, a remaining blind spot, and an appropriate human escalation.

If the result falls short:
When checks disagree, identify which property failed and repair that layer. Preserve what was actually observed and label what remains unknown.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to any generated artifact. Ask what evidence would convince you the result works for its intended purpose, not merely that it looks finished.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Schema passes; evidence fails. Duration does not establish water-damage coverage.

**Change something — Two reviewers praise clarity:** Clarity agreement still supplies no exclusion evidence. Preserve the failed check.

**Decision:** Do structure and consensus establish truth?

**Answer:** No; check actual evidence.

**Why:** A well-formed answer can still be false; multiple agreeing agents can still repeat the same error.

**Review criteria:** Layered checks with the exact defect each detects, a remaining blind spot, and an appropriate human escalation.

**Recovery:** When checks disagree, identify which property failed and repair that layer. Preserve what was actually observed and label what remains unknown.

**Adapt it:** Apply this to any generated artifact. Ask what evidence would convince you the result works for its intended purpose, not merely that it looks finished.


> Every level has a way to be wrong that the level below could not be, and a check that costs less than the mistake.

A convincing response is not the same as a successful task. Check the returned record, the meaning of the answer, and—when tools act—the state left in the environment.

## Keep the checks distinct

| Check | What it can establish | What it cannot establish alone |
|---|---|---|
| Schema and deterministic rules | Required fields, types, arithmetic, allowed values | Truth or completeness of arbitrary prose |
| Source review | Whether evidence supports a claim | Whether retrieval found every relevant source |
| A model reviewer | A rubric-based judgment or proposed critique | Correctness merely because it is another call |
| An evaluation | Performance under defined task conditions | Performance outside those conditions |
| Tracing | Recorded calls, evidence, actions, and errors | A quality judgment without a criterion |

## Check the actual outcome

Inspect the artifact or external state that should result from the task. A success message can be wrong. Refusal or clarification can be the correct outcome when evidence or authority is missing.

Evaluations can grade one run, estimate performance across cases, or compare versions. Graders can be code, people, models, or a combination.

## Make comparisons meaningful

Separate prompt-development cases from held-out evaluation. Repeat tasks when variation matters. Inspect failures as well as averages, and record model, configuration, prompts, tools, data, and harness version.

If several things change together, an improved score does not identify which change helped.

## Reviewers can share blind spots

A reviewer using the same incorrect source can reinforce a mistake. Separate agents may have separate context without statistically independent errors. Complement model review with checkable constraints and independently verified evidence.

Calibrate model graders against human judgments. Disagreements are cases to inspect, not just scores to combine.

## Diagnose the failure you have

If the needed passage never entered context, inspect retrieval and access. If it was present but contradicted, inspect generation. If the proposed action was correct but the tool failed, inspect execution and recovery.

Record useful events and artifacts with appropriate redaction and retention. Observability helps locate a failure; evaluation defines why it counts as one.


## Sources

1. [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) — Anthropic (accessed 2026-09-20)


## Pages that carry it

- [Structured output](/gradient_ascent/techniques/structured-output/) (sourced): Getting answers in a fixed format such as JSON.
- [Write and check](/gradient_ascent/techniques/evaluator-optimizer/) (sourced): One prompt writes, another checks, and the loop repeats until the check passes.
- [Review and debate](/gradient_ascent/techniques/debate-review/) (sourced): Agents that check, or argue with, each other's work.
- [Evals](/gradient_ascent/techniques/evals/) (sourced): Measuring whether a change made the results better.
- [Observability](/gradient_ascent/techniques/observability/) (sourced): Recording what each run did, so a bad result can be traced to the step that caused it.

Last reviewed 2026-09-20.
