Primary sources
- Demystifying evals for AI agents · Anthropic (accessed 09/20/2026)
How you tell whether it worked, from a person reading one answer to a scored set and a recorded trace. The check changes shape at every level; the question does not.
These checks complement each other; they are not a progression from weak to strong.
Check types, required fields, arithmetic, permissions, and source IDs.
BoundaryA valid record can contain false claims.
Grade usefulness, factual support, completeness, and unwanted effects.
BoundaryA model grader needs calibration; agreement is not proof.
Follow the evidence, tool calls, state changes, and errors behind an outcome.
BoundaryTelemetry alone does not establish answer quality.
A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.
Matching verification methods to the kinds of errors a system can make.
Follow a convincing result through different kinds of verification. Separate checking the format, the supporting evidence, and the effect of any action.
Check a well-formatted but wrong support answer.
Layered checks with the exact defect each detects, a remaining blind spot, and an appropriate human escalation.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Check a well-formatted but wrong support answer.
Authored case. Select any record below; nothing is sent to a model.What changed: Establish the facts supplied for this version of the task.
Each check establishes something specific. Valid formatting does not prove a factual claim, and an approved plan does not prove successful execution.
Describe your task to your own model and use Checking the work as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
A convincing response is not the same as a successful task. Check the returned record, the meaning of the answer, and—when tools act—the state left in the environment.
| Check | What it can establish | What it cannot establish alone |
|---|---|---|
| Schema and deterministic rules | Required fields, types, arithmetic, allowed values | Truth or completeness of arbitrary prose |
| Source review | Whether evidence supports a claim | Whether retrieval found every relevant source |
| A model reviewer | A rubric-based judgment or proposed critique | Correctness merely because it is another call |
| An evaluation | Performance under defined task conditions | Performance outside those conditions |
| Tracing | Recorded calls, evidence, actions, and errors | A quality judgment without a criterion |
Inspect the artifact or external state that should result from the task. A success message can be wrong. Refusal or clarification can be the correct outcome when evidence or authority is missing.
Evaluations can grade one run, estimate performance across cases, or compare versions. Graders can be code, people, models, or a combination.
Separate prompt-development cases from held-out evaluation. Repeat tasks when variation matters. Inspect failures as well as averages, and record model, configuration, prompts, tools, data, and harness version.
If several things change together, an improved score does not identify which change helped.
A reviewer using the same incorrect source can reinforce a mistake. Separate agents may have separate context without statistically independent errors. Complement model review with checkable constraints and independently verified evidence.
Calibrate model graders against human judgments. Disagreements are cases to inspect, not just scores to combine.
If the needed passage never entered context, inspect retrieval and access. If it was present but contradicted, inspect generation. If the proposed action was correct but the tool failed, inspect execution and recovery.
Record useful events and artifacts with appropriate redaction and retention. Observability helps locate a failure; evaluation defines why it counts as one.
Last reviewed 09/20/2026. Pages unreviewed for 90 days are flagged for another pass.