# Evaluation frameworks

_Topics at every level · sourced_

The tools that run test sets and graders for you, and what to check before trusting their numbers.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an evaluation configuration into reproducible case results and a report. Inspect what the framework automates and what still depends on the quality of the cases and graders.

**Assumptions:** A runnable harness cannot make an unsuitable rubric meaningful. Versions, inputs, and grading configuration need to be recorded.

**Design choices:** Use automation for repeatable execution and reporting, with human review where grading is uncertain. Keep errors distinct from scored failures.

**Request:** Run a reproducible evaluation over recorded responses.

**Starting evidence:** Manifest: dataset D3, prompt P2, grader G1. Ten expected cases; one request failed.

**Action and control:** Account for every case and preserve versions so missing data cannot silently inflate results.

**Stage records (authored, not executed):**

### Input record

Manifest: dataset D3, prompt P2, grader G1. Ten expected cases; one request failed.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use automation for repeatable execution and reporting, with human review where grading is uncertain. Keep errors distinct from scored failures.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Account for every case and preserve versions so missing data cannot silently inflate results.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Nine answered, one failed. Record the failure rather than shrink the denominator.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Dataset version, reproducible run manifest, grader spot checks, failed-run accounting, and inspectable results.

If the result falls short:
When a run is interrupted or a grader changes, preserve the records and compare only compatible results. Investigate missing cases instead of presenting an incomplete average.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Adapt this to your preferred testing stack. The important contract is reproducible inputs, attributable outputs, explicit grading, and inspectable failures.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Nine answered, one failed. Record the failure rather than shrink the denominator.

**Change something — Grader accepts any answer with a citation:** Wrong answers with irrelevant citations pass. Spot-check and fix the criterion.

**Decision:** Does an automated dashboard ensure valid evaluation?

**Answer:** No; audit graders and missing cases.

**Why:** Missing cases, grading bugs, and inconsistent configurations can inflate scores.

**Review criteria:** Dataset version, reproducible run manifest, grader spot checks, failed-run accounting, and inspectable results.

**Recovery:** When a run is interrupted or a grader changes, preserve the records and compare only compatible results. Investigate missing cases instead of presenting an incomplete average.

**Adapt it:** Adapt this to your preferred testing stack. The important contract is reproducible inputs, attributable outputs, explicit grading, and inspectable failures.

Evaluation frameworks is a page under [evals](/gradient_ascent/techniques/evals/): tools that hold
a dataset of examples, run a program or prompt against every one of them, grade each result, and
let two runs be compared, instead of a team building that machinery from scratch. Some also record
a trace of what happened inside each run, not only the final answer.

Five of them are set out side by side below, described only in their own documentation's words, read
September 19, 2026. There is no ranking and no "best for" here: what fits a team depends on things this
page cannot know: what is already being traced, who has to read a result, and whether the data
being graded may leave the machine at all. A sixth tool, OpenAI Evals, is mid-shutdown and is
reported separately rather than compared. [Prompt
optimization](/gradient_ascent/techniques/prompt-optimization/) is the technique that consumes whatever one of these produces.

This page is sourced, not measured: what each framework does comes from its own documentation,
and no result from any of them exists for this site, so no number below is one it scored.

## Practical guidance

You'll meet these tools as someone else's number, not something you set up yourself: "our evals
run in Braintrust," "CI runs promptfoo," "check the LangSmith experiment." Before repeating that
number in a meeting, ask whoever set it up four questions, in these words.

"Was this graded by an exact match, or by a model reading a rubric?" If a model graded it, ask a
second question right after: "Has anyone checked a sample of its verdicts against a person's own
read?" A dashboard's one aggregate number rarely says which on its own; it's usually a setting or
a second report someone has to go find, and "I don't know" is itself worth having as an answer.

"Is this the same dataset and grading setup as the last number you showed me, or did either
change?" A framework that versions both separately makes this easy to answer; one that doesn't
makes it easy to compare two different things without anyone noticing.

"How many items came back ungraded, and were they left out of the score or counted as
failures?" A malformed grader reply or a run that errored out shouldn't silently become a wrong
answer, and a well-built dashboard can tell you which happened.

"Where does the data being graded actually go?" If the tool is hosted, your prompts and answers
are now sitting with a second company under that company's own terms, whatever your model
maker's page promises. A locally run tool and a cloud dashboard answer this very differently,
which is [safety, privacy and governance](/gradient_ascent/techniques/safety/)'s point about
every intermediary on the route, applied here to a dashboard rather than the model itself.

If two or more of these come back as "I'm not sure" or "I'd have to check," treat the number as
a rumor rather than a result until someone actually answers them; a good dashboard makes all four
easy to answer on the spot, and a bad one makes even the person who built it stop and go looking.

None of this requires reading a line of the tool's code, and building the harness that produces
the number isn't yours to do either: that decision belongs to whoever owns the system, weighed
against everything Build it below lays out. Yours is asking these four questions before you
trust what the dashboard tells you.

## Implementation details

The same four questions, asked of five tools, answered in each one's own current documentation.
Nothing below is this site's assessment of any of them.

**promptfoo.** Calls itself "an open-source CLI and library for evaluating and red-teaming LLM
apps." On where the data goes: "This software runs completely locally. The evals run on your
machine and talk directly with the LLM." No license on the page read. Grading is metrics a user
defines, displayed in "matrix views that let you quickly evaluate outputs across many
prompts"[1]. promptfoo announced on March 9, 2026 that it was joining OpenAI, saying
"Promptfoo will remain open source and we will continue to serve users and customers"[10].

**Braintrust.** Calls itself "the active observability platform for instrumenting, understanding,
and improving agents"[2]. On where the data goes: hosted, with a self-hosting guide that
"offers a self-hosted deployment option that separates data storage from platform management" and
marks self-hosting as available only on the Enterprise plan[4]. No license on the pages
read. Grading is its autoevals library, which "bundles together a variety of automatic evaluation
methods including" LLM-as-a-judge, heuristic (it gives Levenshtein distance) and statistical
(BLEU)[3].

**LangSmith.** On where the data goes, its docs tell a reader setting up an instance to "choose
between cloud, hybrid, or self-hosted", and say "All options include observability, evaluation,
prompt engineering, and deployment"[5]. No license on the pages read. Grading is done by
evaluators, which "are workspace-level resources that score application performance", run over a
dataset ("a collection of examples") where each experiment "captures outputs, evaluator scores,
and execution traces for every example in the dataset"[6].

**DeepEval.** Calls itself "a simple-to-use, open-source LLM evaluation framework, for evaluating
large-language model systems." On where the data goes: its metrics "run locally on your machine",
with an opt-in cloud: using the `deepeval` platform "will allow you to generate sharable testing
reports on the cloud." License: "DeepEval is licensed under Apache 2.0." Grading uses
"LLM-as-a-judge and other NLP models"[7].

**Ragas.** Describes itself as "Objective metrics, intelligent test generation, and data-driven
insights for LLM apps." On where the data goes: installed from PyPI and run locally. License: an
Apache-2.0 badge on its README. Grading is metrics "both LLM-based and traditional", and the
quickstart ships one template today, `rag_eval`, "Evaluate RAG systems", with agent, benchmark,
prompt and workflow templates listed under "Coming Soon"[8].

**OpenAI Evals** is the sixth, described apart because it is being withdrawn: the hosted platform,
not the separate open-source repository of that name. OpenAI's deprecations page gives the dates:
announced June 3, 2026, "Existing evals become read-only" on Oct 31, 2026, and "The Evals dashboard
and API are scheduled to shut down" on Nov 30, 2026, with a linked migration path to
promptfoo[9]. Two more eval tools sit in this site's registry with no description here,
Inspect and lm-evaluation-harness: five is what fits on a page, not a shortlist.

This site's own `scripts/eval_run.py`, described in full on
[the evals page](/gradient_ascent/techniques/evals/), is not a general framework: one script,
one task, no dataset format and no dashboard. Two of its defaults are worth lifting out. It
refuses to write a scored result file for a stub run unless told to, and it draws a fixed 10%
sample of every rubric grader's verdicts for hand-checking:

`scripts/eval_run.py` (lines 762-770)

```python
def sample_for_review(items: list[dict], frac: float = REVIEW_FRACTION, *, seed: int = 0) -> list[dict]:
    """A `frac` sample of the grader's verdicts to hand-check. Deterministic given `seed`: the
    same run re-sampled with the same seed picks the same questions, and a different seed picks
    a different set, so a reviewer can take a second sample without re-running anything."""
    if not items:
        return []
    count = max(1, round(len(items) * frac))
    indexes = sorted(random.Random(seed).sample(range(len(items)), count))
    return [items[i] for i in indexes]
```

Ungraded is a tracked outcome rather than a silent failure: a grader reply that does not parse
as PASS or FAIL is `None`, counted and excluded from the score, never scored as wrong:

`scripts/eval_run.py` (lines 679-689)

```python
def parse_verdict(text: str) -> bool | None:
    """Read a grader's reply. `True` for PASS, `False` for FAIL, `None` for anything else.

    The grader is told to answer with one word, so the first line, stripped of surrounding
    punctuation, must be exactly that word. Malformed output is ungraded, never correct: a
    grader that has drifted or been talked into prose must show up as an ungraded count on the
    result file, not as a silent run of zeros.
    """
    first_line = text.strip().splitlines()[0] if text.strip() else ""
    word = first_line.strip().strip(".,:;!*_`\"'()[]").upper()
    return VERDICTS.get(word)
```

One thing to settle before installing any of these: whether the grade you need is a score over
text at all. A drafted instrument script is graded by running it, a triage label by a confusion
matrix with one costly direction, and a reported measurement is never graded by a model at all.
[Evals](/gradient_ascent/techniques/evals/) sets out the three shapes.

## When you do not need this

Before adopting any of these, try the version that takes an afternoon: twenty questions in a
spreadsheet, the answers pasted in beside them, read by a person. Most teams asking which framework
to pick have not yet written down what a good answer looks like, and a framework will not do that
for them: it will run whatever they have, faster. A tool is worth installing once the same
comparison has to be made repeatedly, or more than one person has to trust the number, which is the
threshold [evals](/gradient_ascent/techniques/evals/) sets for building a golden set at all,
applied now to buying instead of writing.

A rubric model is not a shortcut past reading a sample by hand either. Every one of the five above
still needs someone reading a slice of the verdicts, which is what this site's own runner does by
default rather than on request.

## Scoring the path, not only the answer

Everything above scores an answer against a golden set. From level 4 up there is a second thing to
score: the path the run took to get there. A LangSmith tutorial builds a customer support agent
and runs three kinds of evaluation on it. One of the three is the trajectory, defined there as
"Evaluate whether the agent took the expected path (e.g., of tool calls) to arrive at the final
answer"[11]. Its concepts page makes the same split when it asks what "good" looks like
for an agent, listing "correct tool selection and proper argument formatting or trajectory that
the agent took"[6].

This sits in the evals topic rather than on a rung because it changes nothing about who decides
the next step. It only widens what a grader reads: from the last message to every step before it.
The reason to bother starts where the steps are the model's own. The wrong tool, the right tool
with the wrong arguments, six calls where one would do, and a loop that never stops are all
failures a final-answer grader can score as a pass.

You do not need it below level 4. A single call has no path, and in a
[fixed chain](/gradient_ascent/techniques/prompt-chaining/) the path is the one your code wrote,
so scoring it measures your code rather than the model.

The failure mode is scoring the path as an exact sequence. A run that reaches the right answer a
different way then reads as a failure, and the score punishes the agent for not being your
reference implementation. LangSmith's own tutorial gives the shape that avoids it: a trajectory
evaluation "Compares the actual sequence of steps the agent took against an expected sequence" and
"Calculates a score based on how many of the expected steps were completed
correctly"[11], which is partial credit rather than a match. Report it beside the answer score, never
instead of it: a perfect path to a wrong answer is still a wrong answer.

## Failure modes

### An aggregate score with no visible grading method

- **How to notice it:** A dashboard reports one number, and nobody looking at it can tell whether it came from an exact pattern match or a model reading a rubric, or how many items were graded by each.
- **How to test for it:** Find the per-item grading method in the tool, not just the summary score, before repeating the number in a meeting.

### Two experiments compared across a changed dataset or grader

- **How to notice it:** A 'before' and 'after' number in the same dashboard turn out to come from different dataset versions, different grading configurations, or a grader model the vendor updated between the two runs.
- **How to test for it:** Check the dataset version and grader configuration recorded on each experiment, not just its score, before treating a difference as real.

### Ungraded items folded into the failure count

- **How to notice it:** A framework's default report treats an errored run, a timeout, or an unparseable grader reply the same as a real failure, so the score looks worse than the system that was actually tested, or better, if such items are silently dropped instead.
- **How to test for it:** Find how many items were ungraded or errored on a given run, and read the tool's own docs for how those are counted before trusting the pass rate.

### Eval data sent to a hosted service that should have stayed local

- **How to notice it:** A team adopts a hosted dashboard for convenience and later realizes the prompts and outputs being graded include data that was never supposed to leave the machine it ran on.
- **How to test for it:** Read the deployment options out of each tool's own documentation before adopting it, the way the five entries above do, rather than after data has already been sent.

### A shut-down platform leaves old numbers with no way to reproduce them

- **How to notice it:** A score reported months ago came from a platform that has since been withdrawn, and nobody can re-run the same evaluation to check whether it still holds.
- **How to test for it:** Before citing an old score as still meaningful, check the maker's own deprecations page for the tool. OpenAI's gives two dates for Evals (read-only, then shut down) which is the kind of notice worth finding before the second one passes.

## How to Evaluate It

There is no example on this page, because adopting a framework is a decision rather than a
technique the 60-question set can run. What can be checked is the framework itself, and the check
is the same whichever one is installed: a claim that a system improved needs the same dataset and
the same grading configuration run before and after, a hand-checked sample of any rubric grader's
verdicts, and an honest count of what went ungraded instead of being scored as a failure. If a
tool cannot tell you which of those three it did, that is the finding.

For this site's own runner, `python scripts/eval_run.py --example rag --model stub --dry` projects
what a real run would cost before anything is spent; it is described in full on
[the evals page](/gradient_ascent/techniques/evals/). Whether each tool above offers the same
projection is a question for its own documentation.

## Run it

**What to monitor.** The hand-check agreement rate between a rubric grader and a person, tracked over time
  in whichever tool is running the evals, not just read once at adoption and never checked again.

**Cost at volume.** A hosted service comes with a bill; a self-hosted, open-source one comes with an
  operator instead. Neither is the main number: the model calls being graded, and the grader
  model's own calls on top of them, are what scale with the size of the set.

**How it fails in production.** A grading configuration or dataset changes inside the tool with no version
  recorded, and a later comparison against an old score is quietly comparing two different tests
  without anyone noticing.

**What to log.** The dataset version, the grading configuration, and which items were ungraded, for
  every run, exactly what an experiment or a result file needs to carry so a number can be
  checked later without re-running anything.

## Try it

1. **Use it.** Find a dashboard or report at work, or in a project you use, that shows an eval score. Can you tell from it how items were graded, and how many were not graded at all? Ask whoever set it up; the answer is usually available and rarely on screen.
2. **Build it.** Run python scripts/eval_run.py --example rag --model stub --dry from the repo root and read the projected token count. Then open the documentation of one tool above and look for its own dry run or cost projection. Note whether you found one, and how long it took to find.
3. **Either lane.** Pick one of the five tools above and read its documentation for how it treats an item that errored or whose grader reply could not be parsed. Excluded from the score, counted as a failure, or not stated anywhere? All three answers happen.


## Sources

1. [Promptfoo documentation](https://www.promptfoo.dev/docs/intro/) — promptfoo (accessed 2026-09-19)
2. [Get started with Braintrust](https://www.braintrust.dev/docs) — Braintrust (accessed 2026-09-19)
3. [Autoevals](https://www.braintrust.dev/docs/sdks/typescript/related/autoevals/overview) — Braintrust (accessed 2026-09-19)
4. [Self-hosting Braintrust](https://www.braintrust.dev/docs/admin/self-hosting) — Braintrust (accessed 2026-09-19)
5. [LangSmith Observability](https://docs.langchain.com/langsmith/observability) — LangChain (LangSmith documentation) (accessed 2026-09-19)
6. [Evaluation concepts](https://docs.langchain.com/langsmith/evaluation-concepts) — LangChain (LangSmith documentation) (accessed 2026-09-19)
7. [DeepEval](https://github.com/confident-ai/deepeval) — Confident AI (accessed 2026-09-19)
8. [Ragas](https://github.com/vibrantlabsai/ragas) — Ragas maintainers (repository README) (accessed 2026-09-19)
9. [Deprecations](https://developers.openai.com/api/docs/deprecations) — OpenAI (API documentation) (accessed 2026-09-19)
10. [Promptfoo is joining OpenAI](https://www.promptfoo.dev/blog/promptfoo-joining-openai/) — promptfoo, 2026-03-09 (accessed 2026-09-19)
11. [Evaluate a complex agent](https://docs.langchain.com/langsmith/evaluate-complex-agent) — LangChain (LangSmith documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
