The same four questions, asked of five tools, answered in each one’s own current documentation.
Nothing below is this site’s assessment of any of them.
promptfoo. Calls itself “an open-source CLI and library for evaluating and red-teaming LLM
apps.” On where the data goes: “This software runs completely locally. The evals run on your
machine and talk directly with the LLM.” No license on the page read. Grading is metrics a user
defines, displayed in “matrix views that let you quickly evaluate outputs across many
prompts”[1]. promptfoo announced on March 9, 2026 that it was joining OpenAI, saying
“Promptfoo will remain open source and we will continue to serve users and customers”[10].
Braintrust. Calls itself “the active observability platform for instrumenting, understanding,
and improving agents”[2]. On where the data goes: hosted, with a self-hosting guide that
“offers a self-hosted deployment option that separates data storage from platform management” and
marks self-hosting as available only on the Enterprise plan[4]. No license on the pages
read. Grading is its autoevals library, which “bundles together a variety of automatic evaluation
methods including” LLM-as-a-judge, heuristic (it gives Levenshtein distance) and statistical
(BLEU)[3].
LangSmith. On where the data goes, its docs tell a reader setting up an instance to “choose
between cloud, hybrid, or self-hosted”, and say “All options include observability, evaluation,
prompt engineering, and deployment”[5]. No license on the pages read. Grading is done by
evaluators, which “are workspace-level resources that score application performance”, run over a
dataset (“a collection of examples”) where each experiment “captures outputs, evaluator scores,
and execution traces for every example in the dataset”[6].
DeepEval. Calls itself “a simple-to-use, open-source LLM evaluation framework, for evaluating
large-language model systems.” On where the data goes: its metrics “run locally on your machine”,
with an opt-in cloud: using the deepeval platform “will allow you to generate sharable testing
reports on the cloud.” License: “DeepEval is licensed under Apache 2.0.” Grading uses
“LLM-as-a-judge and other NLP models”[7].
Ragas. Describes itself as “Objective metrics, intelligent test generation, and data-driven
insights for LLM apps.” On where the data goes: installed from PyPI and run locally. License: an
Apache-2.0 badge on its README. Grading is metrics “both LLM-based and traditional”, and the
quickstart ships one template today, rag_eval, “Evaluate RAG systems”, with agent, benchmark,
prompt and workflow templates listed under “Coming Soon”[8].
OpenAI Evals is the sixth, described apart because it is being withdrawn: the hosted platform,
not the separate open-source repository of that name. OpenAI’s deprecations page gives the dates:
announced June 3, 2026, “Existing evals become read-only” on Oct 31, 2026, and “The Evals dashboard
and API are scheduled to shut down” on Nov 30, 2026, with a linked migration path to
promptfoo[9]. Two more eval tools sit in this site’s registry with no description here,
Inspect and lm-evaluation-harness: five is what fits on a page, not a shortlist.
This site’s own scripts/eval_run.py, described in full on
the evals page, is not a general framework: one script,
one task, no dataset format and no dashboard. Two of its defaults are worth lifting out. It
refuses to write a scored result file for a stub run unless told to, and it draws a fixed 10%
sample of every rubric grader’s verdicts for hand-checking:
scripts/eval_run.py · lines 762–770
def sample_for_review(items: list[dict], frac: float = REVIEW_FRACTION, *, seed: int = 0) -> list[dict]:
"""A `frac` sample of the grader's verdicts to hand-check. Deterministic given `seed`: the
same run re-sampled with the same seed picks the same questions, and a different seed picks
a different set, so a reviewer can take a second sample without re-running anything."""
if not items:
return []
count = max(1, round(len(items) * frac))
indexes = sorted(random.Random(seed).sample(range(len(items)), count))
return [items[i] for i in indexes]
Ungraded is a tracked outcome rather than a silent failure: a grader reply that does not parse
as PASS or FAIL is None, counted and excluded from the score, never scored as wrong:
scripts/eval_run.py · lines 679–689
def parse_verdict(text: str) -> bool | None:
"""Read a grader's reply. `True` for PASS, `False` for FAIL, `None` for anything else.
The grader is told to answer with one word, so the first line, stripped of surrounding
punctuation, must be exactly that word. Malformed output is ungraded, never correct: a
grader that has drifted or been talked into prose must show up as an ungraded count on the
result file, not as a silent run of zeros.
"""
first_line = text.strip().splitlines()[0] if text.strip() else ""
word = first_line.strip().strip(".,:;!*_`\"'()[]").upper()
return VERDICTS.get(word)
One thing to settle before installing any of these: whether the grade you need is a score over
text at all. A drafted instrument script is graded by running it, a triage label by a confusion
matrix with one costly direction, and a reported measurement is never graded by a model at all.
Evals sets out the three shapes.