Topics at every level

Evals

Measuring whether a change made the results better.

Sourced

Concept at a glance

Compare a change on the same test cases.

SequenceConceptual illustration
Compare a change on the same test cases.Fixed test set leads to Run + grade. Run + grade leads to Compare versions. A score is meaningful only when the cases and grading match the job you care about.Fixed test setInputs and expected behaviorRun + gradeApply the same criteriaCompare versionsInspect wins and failuresCompare a change on the same test cases.Fixed test set leads to Run + grade. Run + grade leads to Compare versions. A score is meaningful only when the cases and grading match the job you care about.Fixed test setInputs and expected behaviorRun + gradeApply the same criteriaCompare versionsInspect wins and failures
Read the connections in words
  • Fixed test set → Run + grade: Apply the same criteria.
  • Run + grade → Compare versions: Inspect wins and failures.
Key idea

A score is meaningful only when the cases and grading match the job you care about.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Evals: see it in practice.

Systematic measurement of performance against defined tasks, criteria, and reference judgments.

What you’ll walk through

Follow a candidate change through a set of representative tasks and inspect how success is judged. Compare improvements, regressions, and cases the headline score hides.

The task in this version

Compare two answer versions on a held-out review set.

What you’ll learn to check

Per-case outcomes, sliced metrics, disagreement review, uncertainty, and a documented ship/hold decision.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Compare two answer versions on a held-out review set.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Toy review set · fixed
AUTHORED TEACHING RECORD · NOT A LIVE RUN
R1: routine question. R2: routine question. M1: missing evidence. C1: conflicting source. Rubric: 0 = incorrect/unsupported; 1 = correct but unclear; 2 = correct and clear. Separate release rule: inspect unsupported claims regardless of average.

What changed: These authored judgments are invented teaching data, not a benchmark or measured model run.

WHY THIS MATTERS

What this case assumes

The evaluation set and rubric define what the score means. A small or contaminated set can exaggerate improvement.

1 / 6
In this topic

1 pages under evals

Each one goes further into a part of this page than this page does.

Evaluation frameworks

Sourced

The tools that run test sets and graders for you, and what to check before trusting their numbers.

Apply this to your project

Describe your task to your own model and use Evals as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Evals are how a claim that a change made things better gets checked instead of assumed. A golden set is a fixed list of questions with a known right answer, or a rubric for judging one, run under defined conditions. It can support a comparison before and after a change, but evaluations can also grade a single run or estimate performance across a dataset. Grading can use deterministic code, people, models, or a combination. A rubric grader is itself a model call and can be wrong, so some share of its verdicts needs checking by a person. None of this belongs to one level: a single prompt, a fixed workflow and an agent that runs for hours all make claims only an eval can check.

This is the site’s own second principle, on the Method page: a page here may say a level helped only when a result file backs it. This is a topic, not a level.

This page is sourced, not measured: how an eval works is checked against primary sources, but no eval on this site has a scored result file yet (see docs/EVALS.md), so what follows describes the measuring rather than reporting any of it.

Practical guidance

When a coworker tells you a new prompt “tested better,” reply with two sentences: “How was it graded, exact match or a model reading it?” and “Did you run the old version on the identical question set, or a different one?” Those two questions catch most of what makes a “tested better” claim unreliable, and they cost you nothing to ask.

If the answer to the second is “a different set” or “I don’t remember,” there is no comparison yet, whatever the number says. LangSmith’s own documentation describes running evaluations as the way to “catch regressions, and track quality over time”[3], which only works if nothing about the test itself moved between the two runs. On the first question, either grading method is fine: Anthropic’s own guidance calls exact-match grading “perfect for tasks with clear-cut, categorical answers like sentiment analysis (positive, negative, neutral)”[1], and points to model-based grading on a rating scale for qualities like tone and coherence instead. What matters is that the method was the same both times.

The one thing worth building yourself: twenty real questions from your own work, with the answer you’d expect written next to each, kept in a document. When the tool changes (a new model, an edited prompt, a new version) run the same twenty through it and read the answers against what you wrote down. That catches most of what a coworker’s “seems better” would miss, and it takes one afternoon to build, once.

If a rubric grader, a second model judging the answer, is involved, OpenAI’s own documentation on graders warns that “Models being trained sometimes learn to exploit weaknesses in model graders,” and that the tell is a model that “will score highly on model grader evals but score poorly on expert human evaluations”[2]. You don’t need to build that detection yourself; ask whether anyone has checked a handful of the automated grades by hand, and if the answer is no, treat the score as unverified.

A one-time launch score is also not the whole story: Braintrust’s documentation distinguishes evaluation before a change ships from evaluation that keeps running on production traffic once the right answer isn’t known in advance[4]. Building the harness that does either is not your job; that belongs to whoever owns the system, covered in Build it below. Yours is the two sentences and your own twenty questions.

Implementation details

This site measures every technique against one running task: 60 synthetic questions about a synthetic document set, in five kinds (lookup, multi-hop, numeric, unanswerable, conflicting sources), 12 of each, in evals/questions.json. Every question carries its own grading contract: accept and require patterns for exact matching, or a rubric list for a grader model to check against. scripts/eval_run.py runs one example, or all the examples that do this site’s own task, against that set.

Before a real run spends anything, --dry projects its cost. It never calls a model; it runs every example through a stand-in that counts real input tokens and reports the max_tokens cap as the worst-case output, so the number it prints is a ceiling, not a guess:

scripts/eval_run.py · lines 452–486
class DryRunModel:
    """Stands in for a real model during `--dry`. Makes no network call, and projects an upper
    bound rather than a likely run.

    Input tokens are counted for real, with `count_tokens`, over the prompt the example actually
    builds. Output tokens are reported as the `max_tokens` the example asked for, on every call:
    that is the most the provider can bill for output, since the example caps every call. Whenever
    tools are offered it calls the first one, every time, so a tool-using example runs to its own
    step cap or token cap. That is a ceiling except where control flow branches on the model's
    own words: `routing` parses one word to pick a handler, so it projects its cheapest branch.
    """

    def __init__(self, requested_id: str) -> None:
        self.model_id = requested_id
        self.calls = 0

    def complete(
        self,
        messages: list[Message],
        *,
        tools: list[dict] | None = None,
        schema: dict | None = None,
        max_tokens: int = 1024,
    ) -> Completion:
        self.calls += 1
        tokens_in = sum(count_tokens(content_text(m.content)) for m in messages)
        tool_calls: list[ToolCall] = []
        if tools:
            first = tools[0]
            arg_name = next(iter(first["parameters"]["properties"]), "query")
            tool_calls = [ToolCall(name=first["name"], arguments={arg_name: "projected"})]
        text = "" if tool_calls else "[dry run projection, no model called]"
        return Completion(
            text=text, tool_calls=tool_calls, tokens_in=tokens_in, tokens_out=max_tokens, ms=0.0, model_id=self.model_id
        )

That ceiling holds for an example whose control flow does not read the model’s text. routing picks its branch from the model’s answer, so the stand-in’s placeholder text sends it down the cheapest branch and the projection comes in low; docs/EVALS.md says to project a branching example from the branch you expect to be busiest instead.

A real run caches every response by the model id and a hash of the exact prompt, so re-running after a small prompt edit only pays for the questions whose prompt actually changed, and a run can be stopped early with --budget-tokens and resumed later without re-paying for what already ran. A run against the stub model (the one used for testing) is refused a result file unless the caller passes --allow-stub, and is marked "stub": true even then, because a stub answers nothing real:

scripts/eval_run.py · lines 972–983
def stub_refusal(example: str, *, is_stub: bool, allow_stub: bool) -> str | None:
    """The line to print when a stub run is refused a result file, or None when it may write one.

    A stub answers nothing real, so a result file from one measures nothing. It is refused unless
    the caller asks for it outright, and even then the summary carries `"stub": true` so the site
    can refuse to chart it.
    """
    # Named, rather than three lines inside `main`, so a page can pin it by name: a pinned line
    # range over this file has now slid three times, once per wave that grew the runner.
    if is_stub and not allow_stub:
        return f"[{example}] stub model: result not written (pass --allow-stub to write one anyway, marked stub=true)."
    return None

A result file (evals/results/<example>/<model-id>.json) records score_overall over graded questions only, a per-kind breakdown, citation_hit_rate, tokens in and out, wall time, and model_decided_steps (the count of trace steps the model itself chose, defined in examples/common/trace.py) alongside the run date and commit, so a chart built from it can be traced back to exactly what produced it. An ungraded question (no grader configured, or a grader reply that was not readable as PASS or FAIL) is counted and excluded from the score rather than scored as wrong, and 10% of every rubric grader’s verdicts are written to a .review.json file for a person to check by hand.

No example in this repository has a non-stub result file yet: nothing here has been measured against a live model, local or metered. docs/EVALS.md lists exactly which of the site’s examples this question set scores and which it does not, and why: an example that does a different task, such as extracting a record instead of answering a question, gets a different measurement described on its own page rather than a meaningless number from this one.

Grading engineering work

The 60-question set grades text answers, which is one shape of grade among several. For a reader who writes code for electronics test, measurement or design, the shape follows the job, and two of the three below are not a score at all.

A drafted script is a pass or a fail. When a model drafts commands for an instrument from that instrument’s programming manual, there is nothing to rate. Each command is in the documented command set or it is not, and the script either runs on the simulated instrument with an empty error queue or it does not. Grading is running it, so the grade is repeatable and no rubric model takes part in it. Drafting an instrument control script works that case through.

A label is a confusion matrix, not an accuracy number. Sorting failing units and operator notes into causes is classification into a fixed set, and the two directions of a mistake cost different amounts: calling a real defect a fixture problem can let a bad board ship, while calling a fixture problem a defect costs an engineer an afternoon. A single accuracy figure averages the two together and hides the one that matters. Sorting failing units into causes names which direction to drive toward zero and which cheaper one to accept more of.

A reported measurement is not graded by a model at all. A margin, an uncertainty, a Cpk or a verdict is arithmetic, and the check on it is that code computed it and a person can reproduce it. Where a model writes the prose around numbers code produced, the grade is mechanical: every figure in the draft appears in the computed results, character for character, or the draft does not ship. Turning a measurement session into a report runs exactly that check in code.

How many graded examples exist to work with depends on the setting rather than on the technique. A production line yields thousands of labeled units a month, so a rate means something. Design verification has five prototype boards and one sweep, so there is no rate to compute and the honest eval is a person reading every case. A precise measurement may happen once, and what gets checked there is the uncertainty budget behind the number, not a score over a set. Prescribing a set size the reader cannot reach is the fastest way to lose them.

When you do not need this

Skip building a golden set and a grader, and just read five or ten real outputs by hand, when a change is small, reversible, and you are the only person who has to trust the verdict: a prompt tweak checked before lunch does not need a maintained question set to back it. That is still evaluation, just done directly instead of automated; prompt engineering’s own advice to check a new version against the same cases the old one had to pass is exactly this, at spreadsheet scale.

Build the harness once any of that stops holding: the same comparison gets made more than once, more than one person has to trust the number, or a wrong verdict is expensive enough that “it read fine to me” is not a good enough answer on its own.

And there is nothing to evaluate before there is a claim to check. An eval measures whether a specific change made a specific, already-running task better or worse; get the task running first.

Failure modes

Grader hacking

How to notice it
A model or a prompt scores well against a rubric grader, but a person reading the same answers by hand rates them worse: the split a maker's own guidance names as the sign of a model that has learned to exploit the grader rather than do the task.
How to test for it
Run the hand-check sample this site's own runner writes for every rubric verdict, and compare its pass rate against the grader's own pass rate on the same questions.

An ungraded question counted as a zero

How to notice it
A report's accuracy number is lower than it should be because a question the grader could not parse a verdict from was folded into the score as a failure instead of excluded and counted separately.
How to test for it
Check a result file's ungraded count against its overall score; a report that never mentions ungraded questions may be silently treating every one of them as wrong.

Before and after were never the same test

How to notice it
A "tested better" claim turns out to compare two different question sets, two different grading rules, or two runs of a rubric grader whose own verdicts are not perfectly repeatable.
How to test for it
Re-run the old version against the exact question file and grading contract the new version used, rather than trusting a score that was recorded before the test itself changed.

A rubric with nothing specific to check

How to notice it
The grader's verdict on the same answer changes between two runs, because the rubric asks something open-ended (is this good) instead of one specific, checkable claim.
How to test for it
Run the grader on the same answer twice and see whether the verdict is stable. If it moves, no amount of hand-checking makes the number underneath it trustworthy.

A golden set that stopped matching the real task

How to notice it
The score holds steady release after release, but the questions arriving in production have moved on from what the golden set covers, so the number is stable and unrepresentative at the same time.
How to test for it
Sample real traffic and check what share of it resembles a question actually in the set; a low share means the score is still answering yesterday's question.

At each level

  • Conventional software: there is no model call to grade, so this is ordinary software testing: fixed inputs, known outputs, run exhaustively rather than sampled, the way level 0’s own search can be tested exhaustively rather than sampled.
  • Direct prompting: grading one prompt’s answer is the simplest case; exact match works when a task has one right phrasing, the way structured output can be scored on whether the reply is valid at all, separately from whether its values are right, and a rubric grader is needed as soon as a task has no one right phrasing.
  • Added context: what got retrieved changes the answer, so grading has to check citations as well as the final text: whether RAG’s answer used the passage that actually holds the fact, not merely a plausible one.
  • Workflows: a fixed chain of steps can be graded step by step, not only on the final output: prompt chaining’s own citation check is exactly this, so a wrong route or a dropped step shows up even when a later step happens to recover.
  • Tool use: grading has to check whether the right tool was called with the right arguments, since function calling’s own failure modes show a wrong tool call can still produce an answer that reads fine.
  • Agent loops: the model decides when to stop, so an eval has to weigh how many steps and tool calls a single agent took, and whether it stopped too early or kept going too long, alongside whether the final answer is right.
  • Teams of Agents: one agent’s output being graded correct is not enough for a review or debate setup: the agreement between the agents needs measuring too, since two agents can share a blind spot and agree while both are wrong.
  • Always-on agents: there is no single run to grade before it ships, the reason long-running tasks’ own eval section gives for why the site’s question set does not apply to it. Evaluation shifts from a one-time offline pass on a fixed set toward continuous scoring of what the agent actually did once it was already running[4].

Practices

  • Ask how a number was graded before trusting it: exact match, or a model reading a rubric. The two carry different failure modes.
  • Match the grade to the job before picking a tool. A pass or fail from running the thing, a confusion matrix with one named costly direction, and a rubric read by a model are three different measurements, and only the third needs a grader model at all.
  • You do not have to build the machinery. Evaluation frameworks sets six tools that run a set, grade it and compare runs side by side, in their own documentation’s words.
  • If a grader model is used, check some of its verdicts by hand. This site’s own runner writes 10% of them to a file for exactly that.
  • Never treat an ungraded question as a wrong answer, and never let a report quietly do the same; a grader that returns nothing readable is a gap, not a zero.
  • Re-run the same question set before and after a change, not a new set each time, or the comparison is not measuring the change.
  • Report cost and latency next to the score, not separately. A prompt that scores two points higher at ten times the tokens is a different trade than the score alone shows; the operations topic is where those numbers are managed.

Run it

What to monitor

The hand-check agreement rate between a rubric grader and a person, per run, not just once at launch. A grader whose PASS rate climbs across unrelated commits is a sign it has started rewarding its own habits instead of the checklist.

Cost at volume

A dry run's projected tokens, times the question count, times how many examples or levels are being scored, sets the ceiling before a metered run starts. The response cache means a prompt edit only pays again for the questions whose prompt actually changed.

How it fails in production

A grader model is updated by its maker and its verdicts shift with no change to the thing being graded, which looks in a result file exactly like the graded system getting better or worse.

What to log

Every grader verdict next to the question and answer it graded, not only the aggregate score, plus the model id of both the system under test and the grader, the run date and the commit, so a question about one number can be answered by rereading a file instead of running anything again.

Try it

  1. Use it

    Find a benchmark score a product or model page states, and check whether the page says how the answer was graded and whether a person checked any of the verdicts by hand. If it doesn't say, that's the gap this page is about.

  2. Build it

    Run python scripts/eval_run.py --example rag --model stub --dry from the repo root and read the projected tokens it prints, then change TOP_K in examples/rag/run.py from 4 to 2 and run it again. Did the projection change, and does that match what you'd expect from fewer retrieved chunks?

  3. Either lane

    Pick one of the five question kinds in evals/questions.json (lookup, multi-hop, numeric, unanswerable, conflicting sources) and write, on paper, one new question of that kind against evals/corpus/: the question, the accepted answer pattern, and which section it must cite.

How it connects

Before, after and instead of this

Optional: products, tools, and models

8 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 2 more examples
In practice

Compare two support prompts

Run both on the same labeled questions, grade with the same criteria, and inspect where the results differ.

Out there

Named products, tools and models

Tools9
  • BraintrustBraintrust · eval platform
  • DeepEvalConfident AI · eval framework
  • InspectUK AI Security Institute · eval framework
  • LangSmithLangChain · eval and tracing platform
  • lm-evaluation-harnessEleutherAI · benchmark runner
  • OpenAI EvalsOpenAI · eval frameworkRetires 2026-11-30
  • PhoenixArize AI · AI observability and evaluation
  • promptfoopromptfoo · eval runner
  • Ragasopen source · evals for retrieval

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Define success criteria and build evaluations · Anthropic (Claude Platform Docs) (accessed 09/19/2026)
  2. Graders · OpenAI (API documentation) (accessed 09/19/2026)
  3. Evaluation · LangChain (LangSmith documentation) (accessed 09/19/2026)
  4. Evaluate systematically · Braintrust (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page