# Evals

_Topics at every level · sourced_

Measuring whether a change made the results better.


## Try this in a recipe
- [Turn an invoice into a checked record](/gradient_ascent/recipes/invoice-matching.md): Extract a useful JSON record, preserve missing fields, and catch a total that does not reconcile.

## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a candidate change through a set of representative tasks and inspect how success is judged. Compare improvements, regressions, and cases the headline score hides.

**Assumptions:** The evaluation set and rubric define what the score means. A small or contaminated set can exaggerate improvement.

**Design choices:** Include common cases, consequential failures, and a held-out comparison. Separate format checks, factual quality, and action correctness when they matter independently.

**Request:** Compare two answer versions on a held-out review set.

**Starting evidence:** Four fictional cases: two routine, missing evidence, conflicting source. B improves wording but invents one answer.

**Action and control:** Inspect case-level judgments and failure slices, not just an average.

**Stage records (authored, not executed):**

### Toy review set · fixed

R1: routine question.
R2: routine question.
M1: missing evidence.
C1: conflicting source.
Rubric: 0 = incorrect/unsupported; 1 = correct but unclear; 2 = correct and clear.
Separate release rule: inspect unsupported claims regardless of average.

What changed: These authored judgments are invented teaching data, not a benchmark or measured model run.

### Comparison plan

Compare A and B on the same four cases.
Keep routine, missing-evidence, and conflict slices visible.
Track unsupported claims separately from clarity.
Do not edit the evaluation set to favor the preferred candidate.

What changed: The plan makes the acceptance criteria and coverage explicit before choosing a winner.

### Case-level judgments

R1: A=1, B=2; B is clearer.
R2: A=1, B=2; B is clearer.
M1: A=1, B=0; B invents an answer.
C1: A=1, B=1; conflict remains qualified.
Toy total: A=4/8; B=5/8.
Unsupported-claim count: A=0; B=1.

What changed: B's higher total coexists with a new consequential failure.

### Release decision · hold B

Routine clarity improved.
Missing-evidence behavior regressed.
Conflict handling unchanged in this toy set.
Decision under the stated rule: investigate and repair B before release.
No conclusion about real-world quality follows from four invented cases.

What changed: The release decision depends on failure consequences, not only the largest number.

### Coverage change · do not hide it

If M1 is removed: A=3/6; B=5/6.
Missing-evidence coverage becomes zero.
The apparent advantage grows because the difficult case vanished.
Record any justified exclusion and evaluate that risk separately.

What changed: The score changed without either candidate becoming better.

### Your evaluation contract

Replace toy cases with representative authorized tasks and reference judgments.
Choose criteria and failure severity before comparing.
Keep a held-out set separate from tuning.
Inspect grader disagreement and uncertainty.
Make the release decision accountable to the task's actual consequences.

What changed: Use the table format, not its invented scores or a universal pass threshold.

**Sample result:** Hold B for review: routine improvements do not compensate for the unsupported claim. Toy judgments are not measured model quality.

**Change something — Remove the missing-evidence case:** The apparent score rises because coverage shrank. Record exclusions and retain critical cases.

**Decision:** Is a higher aggregate score always a release signal?

**Answer:** No; inspect coverage and consequential failures.

**Why:** Aggregate scores can hide critical failures; do not tune against the final test set or treat a model grader as ground truth.

**Review criteria:** Per-case outcomes, sliced metrics, disagreement review, uncertainty, and a documented ship/hold decision.

**Recovery:** If results are mixed, inspect failure slices and judge the tradeoff against the task. Do not repeatedly tune on the final test set and still call it held out.

**Adapt it:** Use evaluation for prompts, models, retrieval, or complete workflows. Choose examples from the intended use and keep a simple baseline for comparison.

Evals are how a claim that a change made things better gets checked instead of assumed. A
golden set is a fixed list of questions with a known right answer, or a rubric for judging one,
run under defined conditions. It can support a comparison before and after a change, but evaluations can also grade a single run or estimate performance across a dataset. Grading can use deterministic code, people, models, or a combination. A rubric grader is itself a model call and can be wrong, so some share of its verdicts
needs checking by a person. None of this belongs to one level: a single prompt, a fixed workflow
and an agent that runs for hours all make claims only an eval can check.

This is the site's own second principle, on the [Method page](/gradient_ascent/method/): a page
here may say a level helped only when a result file backs it. This is a topic, not a level.

This page is sourced, not measured: how an eval works is checked against primary sources, but no
eval on this site has a scored result file yet (see `docs/EVALS.md`), so what follows describes
the measuring rather than reporting any of it.

## Practical guidance

When a coworker tells you a new prompt "tested better," reply with two sentences: "How was it
graded, exact match or a model reading it?" and "Did you run the old version on the identical
question set, or a different one?" Those two questions catch most of what makes a "tested
better" claim unreliable, and they cost you nothing to ask.

If the answer to the second is "a different set" or "I don't remember," there is no comparison
yet, whatever the number says. LangSmith's own documentation describes running evaluations as the
way to "catch regressions, and track quality over time"[3], which only works if
nothing about the test itself moved between the two runs. On the first question, either grading
method is fine: Anthropic's own guidance calls exact-match grading "perfect for tasks with
clear-cut, categorical answers like sentiment analysis (positive, negative, neutral)"[1],
and points to model-based grading on a rating scale for qualities like tone and coherence instead.
What matters is that the method was the same both times.

The one thing worth building yourself: twenty real questions from your own work, with the answer
you'd expect written next to each, kept in a document. When the tool changes (a new model, an
edited prompt, a new version) run the same twenty through it and read the answers against what
you wrote down. That catches most of what a coworker's "seems better" would miss, and it takes
one afternoon to build, once.

If a rubric grader, a second model judging the answer, is involved, OpenAI's own documentation on
graders warns that "Models being trained sometimes learn to exploit weaknesses in model
graders," and that the tell is a model that "will score highly on model grader evals but score
poorly on expert human evaluations"[2]. You don't need to build that detection
yourself; ask whether anyone has checked a handful of the automated grades by hand, and if the
answer is no, treat the score as unverified.

A one-time launch score is also not the whole story: Braintrust's documentation distinguishes
evaluation before a change ships from evaluation that keeps running on production traffic once
the right answer isn't known in advance[4]. Building the harness that does either is
not your job; that belongs to whoever owns the system, covered in Build it below. Yours is the
two sentences and your own twenty questions.

## Implementation details

This site measures every technique against one running task: 60 synthetic questions about a
synthetic document set, in five kinds (lookup, multi-hop, numeric, unanswerable, conflicting
sources), 12 of each, in `evals/questions.json`. Every question carries its own grading contract:
`accept` and `require` patterns for exact matching, or a `rubric` list for a grader model to
check against. `scripts/eval_run.py` runs one example, or all the examples that do this site's
own task, against that set.

Before a real run spends anything, `--dry` projects its cost. It never calls a model; it runs
every example through a stand-in that counts real input tokens and reports the `max_tokens` cap
as the worst-case output, so the number it prints is a ceiling, not a guess:

`scripts/eval_run.py` (lines 452-486)

```python
class DryRunModel:
    """Stands in for a real model during `--dry`. Makes no network call, and projects an upper
    bound rather than a likely run.

    Input tokens are counted for real, with `count_tokens`, over the prompt the example actually
    builds. Output tokens are reported as the `max_tokens` the example asked for, on every call:
    that is the most the provider can bill for output, since the example caps every call. Whenever
    tools are offered it calls the first one, every time, so a tool-using example runs to its own
    step cap or token cap. That is a ceiling except where control flow branches on the model's
    own words: `routing` parses one word to pick a handler, so it projects its cheapest branch.
    """

    def __init__(self, requested_id: str) -> None:
        self.model_id = requested_id
        self.calls = 0

    def complete(
        self,
        messages: list[Message],
        *,
        tools: list[dict] | None = None,
        schema: dict | None = None,
        max_tokens: int = 1024,
    ) -> Completion:
        self.calls += 1
        tokens_in = sum(count_tokens(content_text(m.content)) for m in messages)
        tool_calls: list[ToolCall] = []
        if tools:
            first = tools[0]
            arg_name = next(iter(first["parameters"]["properties"]), "query")
            tool_calls = [ToolCall(name=first["name"], arguments={arg_name: "projected"})]
        text = "" if tool_calls else "[dry run projection, no model called]"
        return Completion(
            text=text, tool_calls=tool_calls, tokens_in=tokens_in, tokens_out=max_tokens, ms=0.0, model_id=self.model_id
        )
```

That ceiling holds for an example whose control flow does not read the model's text. `routing`
picks its branch from the model's answer, so the stand-in's placeholder text sends it down the
cheapest branch and the projection comes in low; `docs/EVALS.md` says to project a branching
example from the branch you expect to be busiest instead.

A real run caches every response by the model id and a hash of the exact prompt, so re-running
after a small prompt edit only pays for the questions whose prompt actually changed, and a run
can be stopped early with `--budget-tokens` and resumed later without re-paying for what already
ran. A run against the stub model (the one used for testing) is refused a result file unless
the caller passes `--allow-stub`, and is marked `"stub": true` even then, because a stub answers
nothing real:

`scripts/eval_run.py` (lines 972-983)

```python
def stub_refusal(example: str, *, is_stub: bool, allow_stub: bool) -> str | None:
    """The line to print when a stub run is refused a result file, or None when it may write one.

    A stub answers nothing real, so a result file from one measures nothing. It is refused unless
    the caller asks for it outright, and even then the summary carries `"stub": true` so the site
    can refuse to chart it.
    """
    # Named, rather than three lines inside `main`, so a page can pin it by name: a pinned line
    # range over this file has now slid three times, once per wave that grew the runner.
    if is_stub and not allow_stub:
        return f"[{example}] stub model: result not written (pass --allow-stub to write one anyway, marked stub=true)."
    return None
```

A result file (`evals/results/<example>/<model-id>.json`) records `score_overall` over graded
questions only, a per-kind breakdown, `citation_hit_rate`, tokens in and out, wall time, and
`model_decided_steps` (the count of trace steps the model itself chose, defined in
`examples/common/trace.py`) alongside the run date and commit, so a chart built from it can be
traced back to exactly what produced it. An `ungraded` question (no grader configured, or a
grader reply that was not readable as PASS or FAIL) is counted and excluded from the score rather
than scored as wrong, and 10% of every rubric grader's verdicts are written to a `.review.json`
file for a person to check by hand.

No example in this repository has a non-stub result file yet: nothing here has been measured
against a live model, local or metered. `docs/EVALS.md` lists exactly which of the site's
examples this question set scores and which it does not, and why: an example that does a
different task, such as extracting a record instead of answering a question, gets a different
measurement described on its own page rather than a meaningless number from this one.

## Grading engineering work

The 60-question set grades text answers, which is one shape of grade among several. For a reader
who writes code for electronics test, measurement or design, the shape follows the job, and two of
the three below are not a score at all.

**A drafted script is a pass or a fail.** When a model drafts commands for an instrument from that
instrument's programming manual, there is nothing to rate. Each command is in the documented
command set or it is not, and the script either runs on the simulated instrument with an empty
error queue or it does not. Grading is running it, so the grade is repeatable and no rubric model
takes part in it. [Drafting an instrument
control script](/gradient_ascent/recipes/instrument-script-from-the-manual/) works that case through.

**A label is a confusion matrix, not an accuracy number.** Sorting failing units and operator
notes into causes is classification into a fixed set, and the two directions of a mistake cost
different amounts: calling a real defect a fixture problem can let a bad board ship, while calling
a fixture problem a defect costs an engineer an afternoon. A single accuracy figure averages the
two together and hides the one that matters.
[Sorting failing units into causes](/gradient_ascent/recipes/test-failure-triage/) names which
direction to drive toward zero and which cheaper one to accept more of.

**A reported measurement is not graded by a model at all.** A margin, an uncertainty, a Cpk or a
verdict is arithmetic, and the check on it is that code computed it and a person can reproduce it.
Where a model writes the prose around numbers code produced, the grade is mechanical: every figure
in the draft appears in the computed results, character for character, or the draft does not ship.
[Turning a measurement session into a report](/gradient_ascent/recipes/measurement-writeup/) runs
exactly that check in code.

How many graded examples exist to work with depends on the setting rather than on the technique. A
production line yields thousands of labeled units a month, so a rate means something. Design
verification has five prototype boards and one sweep, so there is no rate to compute and the
honest eval is a person reading every case. A precise measurement may happen once, and what gets
checked there is the uncertainty budget behind the number, not a score over a set. Prescribing a
set size the reader cannot reach is the fastest way to lose them.

## When you do not need this

Skip building a golden set and a grader, and just read five or ten real outputs by hand, when a
change is small, reversible, and you are the only person who has to trust the verdict: a prompt
tweak checked before lunch does not need a maintained question set to back it. That is still
evaluation, just done directly instead of automated;
[prompt engineering](/gradient_ascent/techniques/prompt-engineering/)'s own advice to check a
new version against the same cases the old one had to pass is exactly this, at spreadsheet scale.

Build the harness once any of that stops holding: the same comparison gets made more than once,
more than one person has to trust the number, or a wrong verdict is expensive enough that "it
read fine to me" is not a good enough answer on its own.

And there is nothing to evaluate before there is a claim to check. An eval measures whether a
specific change made a specific, already-running task better or worse; get the task running
first.

## Failure modes

### Grader hacking

- **How to notice it:** A model or a prompt scores well against a rubric grader, but a person reading the same answers by hand rates them worse: the split a maker's own guidance names as the sign of a model that has learned to exploit the grader rather than do the task.
- **How to test for it:** Run the hand-check sample this site's own runner writes for every rubric verdict, and compare its pass rate against the grader's own pass rate on the same questions.

### An ungraded question counted as a zero

- **How to notice it:** A report's accuracy number is lower than it should be because a question the grader could not parse a verdict from was folded into the score as a failure instead of excluded and counted separately.
- **How to test for it:** Check a result file's ungraded count against its overall score; a report that never mentions ungraded questions may be silently treating every one of them as wrong.

### Before and after were never the same test

- **How to notice it:** A "tested better" claim turns out to compare two different question sets, two different grading rules, or two runs of a rubric grader whose own verdicts are not perfectly repeatable.
- **How to test for it:** Re-run the old version against the exact question file and grading contract the new version used, rather than trusting a score that was recorded before the test itself changed.

### A rubric with nothing specific to check

- **How to notice it:** The grader's verdict on the same answer changes between two runs, because the rubric asks something open-ended (is this good) instead of one specific, checkable claim.
- **How to test for it:** Run the grader on the same answer twice and see whether the verdict is stable. If it moves, no amount of hand-checking makes the number underneath it trustworthy.

### A golden set that stopped matching the real task

- **How to notice it:** The score holds steady release after release, but the questions arriving in production have moved on from what the golden set covers, so the number is stable and unrepresentative at the same time.
- **How to test for it:** Sample real traffic and check what share of it resembles a question actually in the set; a low share means the score is still answering yesterday's question.

## At each level

- [Conventional software](/gradient_ascent/levels/0/): there is no model call to grade, so this is ordinary
  software testing: fixed inputs, known outputs, run exhaustively rather than sampled, the way
  [level 0](/gradient_ascent/techniques/order-zero/)'s own search can be tested exhaustively
  rather than sampled.
- [Direct prompting](/gradient_ascent/levels/1/): grading one prompt's answer is the simplest case;
  exact match works when a task has one right phrasing, the way
  [structured output](/gradient_ascent/techniques/structured-output/) can be scored on whether
  the reply is valid at all, separately from whether its values are right, and a rubric grader is
  needed as soon as a task has no one right phrasing.
- [Added context](/gradient_ascent/levels/2/): what got retrieved changes the answer, so grading has
  to check citations as well as the final text: whether
  [RAG](/gradient_ascent/techniques/rag/)'s answer used the passage that actually holds the
  fact, not merely a plausible one.
- [Workflows](/gradient_ascent/levels/3/): a fixed chain of steps can be graded step by step,
  not only on the final output: [prompt chaining](/gradient_ascent/techniques/prompt-chaining/)'s
  own citation check is exactly this, so a wrong route or a dropped step shows up even when a
  later step happens to recover.
- [Tool use](/gradient_ascent/levels/4/): grading has to check whether the right tool was called
  with the right arguments, since [function
  calling](/gradient_ascent/techniques/function-calling/)'s own failure modes show a wrong tool call can still produce an answer that
  reads fine.
- [Agent loops](/gradient_ascent/levels/5/): the model decides when to stop, so an eval has to weigh
  how many steps and tool calls [a single agent](/gradient_ascent/techniques/single-agent/) took,
  and whether it stopped too early or kept going too long, alongside whether the final answer is
  right.
- [Teams of Agents](/gradient_ascent/levels/6/): one agent's output being graded correct is not
  enough for a [review or debate](/gradient_ascent/techniques/debate-review/) setup: the
  agreement between the agents needs measuring too, since two agents can share a blind spot and
  agree while both are wrong.
- [Always-on agents](/gradient_ascent/levels/7/): there is no single run to grade before it
  ships, the reason [long-running tasks](/gradient_ascent/techniques/long-horizon/)' own eval
  section gives for why the site's question set does not apply to it. Evaluation shifts from a
  one-time offline pass on a fixed set toward continuous scoring of what the agent actually did
  once it was already running[4].

## Practices

- Ask how a number was graded before trusting it: exact match, or a model reading a rubric. The
  two carry different failure modes.
- Match the grade to the job before picking a tool. A pass or fail from running the thing, a
  confusion matrix with one named costly direction, and a rubric read by a model are three
  different measurements, and only the third needs a grader model at all.
- You do not have to build the machinery. [Evaluation
  frameworks](/gradient_ascent/techniques/eval-frameworks/) sets six tools that run a set, grade it and compare runs side by side, in
  their own documentation's words.
- If a grader model is used, check some of its verdicts by hand. This site's own runner writes
  10% of them to a file for exactly that.
- Never treat an ungraded question as a wrong answer, and never let a report quietly do the same;
  a grader that returns nothing readable is a gap, not a zero.
- Re-run the same question set before and after a change, not a new set each time, or the
  comparison is not measuring the change.
- Report cost and latency next to the score, not separately. A prompt that scores two points
  higher at ten times the tokens is a different trade than the score alone shows; the
  [operations](/gradient_ascent/techniques/ops/) topic is where those numbers are managed.

## Run it

**What to monitor.** The hand-check agreement rate between a rubric grader and a person, per run, not just
  once at launch. A grader whose PASS rate climbs across unrelated commits is a sign it has
  started rewarding its own habits instead of the checklist.

**Cost at volume.** A dry run's projected tokens, times the question count, times how many examples
  or levels are being scored, sets the ceiling before a metered run starts. The response cache
  means a prompt edit only pays again for the questions whose prompt actually changed.

**How it fails in production.** A grader model is updated by its maker and its verdicts shift with no change
  to the thing being graded, which looks in a result file exactly like the graded system getting
  better or worse.

**What to log.** Every grader verdict next to the question and answer it graded, not only the
  aggregate score, plus the model id of both the system under test and the grader, the run date
  and the commit, so a question about one number can be answered by rereading a file instead of
  running anything again.

## Try it

1. **Use it.** Find a benchmark score a product or model page states, and check whether the page says how the answer was graded and whether a person checked any of the verdicts by hand. If it doesn't say, that's the gap this page is about.
2. **Build it.** Run python scripts/eval_run.py --example rag --model stub --dry from the repo root and read the projected tokens it prints, then change TOP_K in examples/rag/run.py from 4 to 2 and run it again. Did the projection change, and does that match what you'd expect from fewer retrieved chunks?
3. **Either lane.** Pick one of the five question kinds in evals/questions.json (lookup, multi-hop, numeric, unanswerable, conflicting sources) and write, on paper, one new question of that kind against evals/corpus/: the question, the accepted answer pattern, and which section it must cite.


## Sources

1. [Define success criteria and build evaluations](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests) — Anthropic (Claude Platform Docs) (accessed 2026-09-19)
2. [Graders](https://developers.openai.com/api/docs/guides/graders) — OpenAI (API documentation) (accessed 2026-09-19)
3. [Evaluation](https://docs.langchain.com/langsmith/evaluation) — LangChain (LangSmith documentation) (accessed 2026-09-19)
4. [Evaluate systematically](https://www.braintrust.dev/docs/evaluate) — Braintrust (accessed 2026-09-19)


Last reviewed 2026-09-19.
