Topics at every level

Evaluation frameworks

The tools that run test sets and graders for you, and what to check before trusting their numbers.

Sourced

Concept at a glance

Automate the evaluation, keep the judgment visible.

SequenceConceptual illustration
Automate the evaluation, keep the judgment visible.Cases + graders leads to Evaluation runner. Evaluation runner leads to Report + review. The framework runs the machinery; you remain responsible for the test and grader quality.Cases + gradersDefine the experimentEvaluation runnerExecute and record resultsReport + reviewInspect scores and examplesAutomate the evaluation, keep the judgment visible.Cases + graders leads to Evaluation runner. Evaluation runner leads to Report + review. The framework runs the machinery; you remain responsible for the test and grader quality.Cases + gradersDefine the experimentEvaluation runnerExecute and record resultsReport + reviewInspect scores and examples
Read the connections in words
  • Cases + graders → Evaluation runner: Execute and record results.
  • Evaluation runner → Report + review: Inspect scores and examples.
Key idea

The framework runs the machinery; you remain responsible for the test and grader quality.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Evaluation frameworks: see it in practice.

Infrastructure for organizing datasets, executing runs, applying graders, and reporting evaluation results.

What you’ll walk through

Follow an evaluation configuration into reproducible case results and a report. Inspect what the framework automates and what still depends on the quality of the cases and graders.

The task in this version

Run a reproducible evaluation over recorded responses.

What you’ll learn to check

Dataset version, reproducible run manifest, grader spot checks, failed-run accounting, and inspectable results.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Run a reproducible evaluation over recorded responses.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Manifest: dataset D3, prompt P2, grader G1. Ten expected cases; one request failed.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

A runnable harness cannot make an unsuitable rubric meaningful. Versions, inputs, and grading configuration need to be recorded.

1 / 6

Apply this to your project

Describe your task to your own model and use Evaluation frameworks as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Evaluation frameworks is a page under evals: tools that hold a dataset of examples, run a program or prompt against every one of them, grade each result, and let two runs be compared, instead of a team building that machinery from scratch. Some also record a trace of what happened inside each run, not only the final answer.

Five of them are set out side by side below, described only in their own documentation’s words, read September 19, 2026. There is no ranking and no “best for” here: what fits a team depends on things this page cannot know: what is already being traced, who has to read a result, and whether the data being graded may leave the machine at all. A sixth tool, OpenAI Evals, is mid-shutdown and is reported separately rather than compared. Prompt optimization is the technique that consumes whatever one of these produces.

This page is sourced, not measured: what each framework does comes from its own documentation, and no result from any of them exists for this site, so no number below is one it scored.

Practical guidance

You’ll meet these tools as someone else’s number, not something you set up yourself: “our evals run in Braintrust,” “CI runs promptfoo,” “check the LangSmith experiment.” Before repeating that number in a meeting, ask whoever set it up four questions, in these words.

“Was this graded by an exact match, or by a model reading a rubric?” If a model graded it, ask a second question right after: “Has anyone checked a sample of its verdicts against a person’s own read?” A dashboard’s one aggregate number rarely says which on its own; it’s usually a setting or a second report someone has to go find, and “I don’t know” is itself worth having as an answer.

“Is this the same dataset and grading setup as the last number you showed me, or did either change?” A framework that versions both separately makes this easy to answer; one that doesn’t makes it easy to compare two different things without anyone noticing.

“How many items came back ungraded, and were they left out of the score or counted as failures?” A malformed grader reply or a run that errored out shouldn’t silently become a wrong answer, and a well-built dashboard can tell you which happened.

“Where does the data being graded actually go?” If the tool is hosted, your prompts and answers are now sitting with a second company under that company’s own terms, whatever your model maker’s page promises. A locally run tool and a cloud dashboard answer this very differently, which is safety, privacy and governance’s point about every intermediary on the route, applied here to a dashboard rather than the model itself.

If two or more of these come back as “I’m not sure” or “I’d have to check,” treat the number as a rumor rather than a result until someone actually answers them; a good dashboard makes all four easy to answer on the spot, and a bad one makes even the person who built it stop and go looking.

None of this requires reading a line of the tool’s code, and building the harness that produces the number isn’t yours to do either: that decision belongs to whoever owns the system, weighed against everything Build it below lays out. Yours is asking these four questions before you trust what the dashboard tells you.

Implementation details

The same four questions, asked of five tools, answered in each one’s own current documentation. Nothing below is this site’s assessment of any of them.

promptfoo. Calls itself “an open-source CLI and library for evaluating and red-teaming LLM apps.” On where the data goes: “This software runs completely locally. The evals run on your machine and talk directly with the LLM.” No license on the page read. Grading is metrics a user defines, displayed in “matrix views that let you quickly evaluate outputs across many prompts”[1]. promptfoo announced on March 9, 2026 that it was joining OpenAI, saying “Promptfoo will remain open source and we will continue to serve users and customers”[10].

Braintrust. Calls itself “the active observability platform for instrumenting, understanding, and improving agents”[2]. On where the data goes: hosted, with a self-hosting guide that “offers a self-hosted deployment option that separates data storage from platform management” and marks self-hosting as available only on the Enterprise plan[4]. No license on the pages read. Grading is its autoevals library, which “bundles together a variety of automatic evaluation methods including” LLM-as-a-judge, heuristic (it gives Levenshtein distance) and statistical (BLEU)[3].

LangSmith. On where the data goes, its docs tell a reader setting up an instance to “choose between cloud, hybrid, or self-hosted”, and say “All options include observability, evaluation, prompt engineering, and deployment”[5]. No license on the pages read. Grading is done by evaluators, which “are workspace-level resources that score application performance”, run over a dataset (“a collection of examples”) where each experiment “captures outputs, evaluator scores, and execution traces for every example in the dataset”[6].

DeepEval. Calls itself “a simple-to-use, open-source LLM evaluation framework, for evaluating large-language model systems.” On where the data goes: its metrics “run locally on your machine”, with an opt-in cloud: using the deepeval platform “will allow you to generate sharable testing reports on the cloud.” License: “DeepEval is licensed under Apache 2.0.” Grading uses “LLM-as-a-judge and other NLP models”[7].

Ragas. Describes itself as “Objective metrics, intelligent test generation, and data-driven insights for LLM apps.” On where the data goes: installed from PyPI and run locally. License: an Apache-2.0 badge on its README. Grading is metrics “both LLM-based and traditional”, and the quickstart ships one template today, rag_eval, “Evaluate RAG systems”, with agent, benchmark, prompt and workflow templates listed under “Coming Soon”[8].

OpenAI Evals is the sixth, described apart because it is being withdrawn: the hosted platform, not the separate open-source repository of that name. OpenAI’s deprecations page gives the dates: announced June 3, 2026, “Existing evals become read-only” on Oct 31, 2026, and “The Evals dashboard and API are scheduled to shut down” on Nov 30, 2026, with a linked migration path to promptfoo[9]. Two more eval tools sit in this site’s registry with no description here, Inspect and lm-evaluation-harness: five is what fits on a page, not a shortlist.

This site’s own scripts/eval_run.py, described in full on the evals page, is not a general framework: one script, one task, no dataset format and no dashboard. Two of its defaults are worth lifting out. It refuses to write a scored result file for a stub run unless told to, and it draws a fixed 10% sample of every rubric grader’s verdicts for hand-checking:

scripts/eval_run.py · lines 762–770
def sample_for_review(items: list[dict], frac: float = REVIEW_FRACTION, *, seed: int = 0) -> list[dict]:
    """A `frac` sample of the grader's verdicts to hand-check. Deterministic given `seed`: the
    same run re-sampled with the same seed picks the same questions, and a different seed picks
    a different set, so a reviewer can take a second sample without re-running anything."""
    if not items:
        return []
    count = max(1, round(len(items) * frac))
    indexes = sorted(random.Random(seed).sample(range(len(items)), count))
    return [items[i] for i in indexes]

Ungraded is a tracked outcome rather than a silent failure: a grader reply that does not parse as PASS or FAIL is None, counted and excluded from the score, never scored as wrong:

scripts/eval_run.py · lines 679–689
def parse_verdict(text: str) -> bool | None:
    """Read a grader's reply. `True` for PASS, `False` for FAIL, `None` for anything else.

    The grader is told to answer with one word, so the first line, stripped of surrounding
    punctuation, must be exactly that word. Malformed output is ungraded, never correct: a
    grader that has drifted or been talked into prose must show up as an ungraded count on the
    result file, not as a silent run of zeros.
    """
    first_line = text.strip().splitlines()[0] if text.strip() else ""
    word = first_line.strip().strip(".,:;!*_`\"'()[]").upper()
    return VERDICTS.get(word)

One thing to settle before installing any of these: whether the grade you need is a score over text at all. A drafted instrument script is graded by running it, a triage label by a confusion matrix with one costly direction, and a reported measurement is never graded by a model at all. Evals sets out the three shapes.

When you do not need this

Before adopting any of these, try the version that takes an afternoon: twenty questions in a spreadsheet, the answers pasted in beside them, read by a person. Most teams asking which framework to pick have not yet written down what a good answer looks like, and a framework will not do that for them: it will run whatever they have, faster. A tool is worth installing once the same comparison has to be made repeatedly, or more than one person has to trust the number, which is the threshold evals sets for building a golden set at all, applied now to buying instead of writing.

A rubric model is not a shortcut past reading a sample by hand either. Every one of the five above still needs someone reading a slice of the verdicts, which is what this site’s own runner does by default rather than on request.

Scoring the path, not only the answer

Everything above scores an answer against a golden set. From level 4 up there is a second thing to score: the path the run took to get there. A LangSmith tutorial builds a customer support agent and runs three kinds of evaluation on it. One of the three is the trajectory, defined there as “Evaluate whether the agent took the expected path (e.g., of tool calls) to arrive at the final answer”[11]. Its concepts page makes the same split when it asks what “good” looks like for an agent, listing “correct tool selection and proper argument formatting or trajectory that the agent took”[6].

This sits in the evals topic rather than on a rung because it changes nothing about who decides the next step. It only widens what a grader reads: from the last message to every step before it. The reason to bother starts where the steps are the model’s own. The wrong tool, the right tool with the wrong arguments, six calls where one would do, and a loop that never stops are all failures a final-answer grader can score as a pass.

You do not need it below level 4. A single call has no path, and in a fixed chain the path is the one your code wrote, so scoring it measures your code rather than the model.

The failure mode is scoring the path as an exact sequence. A run that reaches the right answer a different way then reads as a failure, and the score punishes the agent for not being your reference implementation. LangSmith’s own tutorial gives the shape that avoids it: a trajectory evaluation “Compares the actual sequence of steps the agent took against an expected sequence” and “Calculates a score based on how many of the expected steps were completed correctly”[11], which is partial credit rather than a match. Report it beside the answer score, never instead of it: a perfect path to a wrong answer is still a wrong answer.

Failure modes

An aggregate score with no visible grading method

How to notice it
A dashboard reports one number, and nobody looking at it can tell whether it came from an exact pattern match or a model reading a rubric, or how many items were graded by each.
How to test for it
Find the per-item grading method in the tool, not just the summary score, before repeating the number in a meeting.

Two experiments compared across a changed dataset or grader

How to notice it
A 'before' and 'after' number in the same dashboard turn out to come from different dataset versions, different grading configurations, or a grader model the vendor updated between the two runs.
How to test for it
Check the dataset version and grader configuration recorded on each experiment, not just its score, before treating a difference as real.

Ungraded items folded into the failure count

How to notice it
A framework's default report treats an errored run, a timeout, or an unparseable grader reply the same as a real failure, so the score looks worse than the system that was actually tested, or better, if such items are silently dropped instead.
How to test for it
Find how many items were ungraded or errored on a given run, and read the tool's own docs for how those are counted before trusting the pass rate.

Eval data sent to a hosted service that should have stayed local

How to notice it
A team adopts a hosted dashboard for convenience and later realizes the prompts and outputs being graded include data that was never supposed to leave the machine it ran on.
How to test for it
Read the deployment options out of each tool's own documentation before adopting it, the way the five entries above do, rather than after data has already been sent.

A shut-down platform leaves old numbers with no way to reproduce them

How to notice it
A score reported months ago came from a platform that has since been withdrawn, and nobody can re-run the same evaluation to check whether it still holds.
How to test for it
Before citing an old score as still meaningful, check the maker's own deprecations page for the tool. OpenAI's gives two dates for Evals (read-only, then shut down) which is the kind of notice worth finding before the second one passes.

How to Evaluate It

There is no example on this page, because adopting a framework is a decision rather than a technique the 60-question set can run. What can be checked is the framework itself, and the check is the same whichever one is installed: a claim that a system improved needs the same dataset and the same grading configuration run before and after, a hand-checked sample of any rubric grader’s verdicts, and an honest count of what went ungraded instead of being scored as a failure. If a tool cannot tell you which of those three it did, that is the finding.

For this site’s own runner, python scripts/eval_run.py --example rag --model stub --dry projects what a real run would cost before anything is spent; it is described in full on the evals page. Whether each tool above offers the same projection is a question for its own documentation.

Run it

What to monitor

The hand-check agreement rate between a rubric grader and a person, tracked over time in whichever tool is running the evals, not just read once at adoption and never checked again.

Cost at volume

A hosted service comes with a bill; a self-hosted, open-source one comes with an operator instead. Neither is the main number: the model calls being graded, and the grader model's own calls on top of them, are what scale with the size of the set.

How it fails in production

A grading configuration or dataset changes inside the tool with no version recorded, and a later comparison against an old score is quietly comparing two different tests without anyone noticing.

What to log

The dataset version, the grading configuration, and which items were ungraded, for every run, exactly what an experiment or a result file needs to carry so a number can be checked later without re-running anything.

Try it

  1. Use it

    Find a dashboard or report at work, or in a project you use, that shows an eval score. Can you tell from it how items were graded, and how many were not graded at all? Ask whoever set it up; the answer is usually available and rarely on screen.

  2. Build it

    Run python scripts/eval_run.py --example rag --model stub --dry from the repo root and read the projected token count. Then open the documentation of one tool above and look for its own dry run or cost projection. Note whether you found one, and how long it took to find.

  3. Either lane

    Pick one of the five tools above and read its documentation for how it treats an item that errored or whose grader reply could not be parsed. Excluded from the score, counted as a failure, or not stated anywhere? All three answers happen.

How it connects

Before, after and instead of this

Read first

Optional: products, tools, and models

6 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Repeat a regression evaluation

A runner loads saved cases, calls each candidate system, applies graders, and records comparable reports.

Out there

Named products, tools and models

Tools6
  • BraintrustBraintrust · eval platform
  • DeepEvalConfident AI · eval framework
  • LangSmithLangChain · eval and tracing platform
  • PhoenixArize AI · AI observability and evaluation
  • promptfoopromptfoo · eval runner
  • Ragasopen source · evals for retrieval

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Promptfoo documentation · promptfoo (accessed 09/19/2026)
  2. Get started with Braintrust · Braintrust (accessed 09/19/2026)
  3. Autoevals · Braintrust (accessed 09/19/2026)
  4. Self-hosting Braintrust · Braintrust (accessed 09/19/2026)
  5. LangSmith Observability · LangChain (LangSmith documentation) (accessed 09/19/2026)
  6. Evaluation concepts · LangChain (LangSmith documentation) (accessed 09/19/2026)
  7. DeepEval · Confident AI (accessed 09/19/2026)
  8. Ragas · Ragas maintainers (repository README) (accessed 09/19/2026)
  9. Deprecations · OpenAI (API documentation) (accessed 09/19/2026)
  10. Promptfoo is joining OpenAI · promptfoo, 03/09/2026 (accessed 09/19/2026)
  11. Evaluate a complex agent · LangChain (LangSmith documentation) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page