This site measures every technique against one running task: 60 synthetic questions about a
synthetic document set, in five kinds (lookup, multi-hop, numeric, unanswerable, conflicting
sources), 12 of each, in evals/questions.json. Every question carries its own grading contract:
accept and require patterns for exact matching, or a rubric list for a grader model to
check against. scripts/eval_run.py runs one example, or all the examples that do this site’s
own task, against that set.
Before a real run spends anything, --dry projects its cost. It never calls a model; it runs
every example through a stand-in that counts real input tokens and reports the max_tokens cap
as the worst-case output, so the number it prints is a ceiling, not a guess:
scripts/eval_run.py · lines 452–486
class DryRunModel:
"""Stands in for a real model during `--dry`. Makes no network call, and projects an upper
bound rather than a likely run.
Input tokens are counted for real, with `count_tokens`, over the prompt the example actually
builds. Output tokens are reported as the `max_tokens` the example asked for, on every call:
that is the most the provider can bill for output, since the example caps every call. Whenever
tools are offered it calls the first one, every time, so a tool-using example runs to its own
step cap or token cap. That is a ceiling except where control flow branches on the model's
own words: `routing` parses one word to pick a handler, so it projects its cheapest branch.
"""
def __init__(self, requested_id: str) -> None:
self.model_id = requested_id
self.calls = 0
def complete(
self,
messages: list[Message],
*,
tools: list[dict] | None = None,
schema: dict | None = None,
max_tokens: int = 1024,
) -> Completion:
self.calls += 1
tokens_in = sum(count_tokens(content_text(m.content)) for m in messages)
tool_calls: list[ToolCall] = []
if tools:
first = tools[0]
arg_name = next(iter(first["parameters"]["properties"]), "query")
tool_calls = [ToolCall(name=first["name"], arguments={arg_name: "projected"})]
text = "" if tool_calls else "[dry run projection, no model called]"
return Completion(
text=text, tool_calls=tool_calls, tokens_in=tokens_in, tokens_out=max_tokens, ms=0.0, model_id=self.model_id
)
That ceiling holds for an example whose control flow does not read the model’s text. routing
picks its branch from the model’s answer, so the stand-in’s placeholder text sends it down the
cheapest branch and the projection comes in low; docs/EVALS.md says to project a branching
example from the branch you expect to be busiest instead.
A real run caches every response by the model id and a hash of the exact prompt, so re-running
after a small prompt edit only pays for the questions whose prompt actually changed, and a run
can be stopped early with --budget-tokens and resumed later without re-paying for what already
ran. A run against the stub model (the one used for testing) is refused a result file unless
the caller passes --allow-stub, and is marked "stub": true even then, because a stub answers
nothing real:
scripts/eval_run.py · lines 972–983
def stub_refusal(example: str, *, is_stub: bool, allow_stub: bool) -> str | None:
"""The line to print when a stub run is refused a result file, or None when it may write one.
A stub answers nothing real, so a result file from one measures nothing. It is refused unless
the caller asks for it outright, and even then the summary carries `"stub": true` so the site
can refuse to chart it.
"""
# Named, rather than three lines inside `main`, so a page can pin it by name: a pinned line
# range over this file has now slid three times, once per wave that grew the runner.
if is_stub and not allow_stub:
return f"[{example}] stub model: result not written (pass --allow-stub to write one anyway, marked stub=true)."
return None
A result file (evals/results/<example>/<model-id>.json) records score_overall over graded
questions only, a per-kind breakdown, citation_hit_rate, tokens in and out, wall time, and
model_decided_steps (the count of trace steps the model itself chose, defined in
examples/common/trace.py) alongside the run date and commit, so a chart built from it can be
traced back to exactly what produced it. An ungraded question (no grader configured, or a
grader reply that was not readable as PASS or FAIL) is counted and excluded from the score rather
than scored as wrong, and 10% of every rubric grader’s verdicts are written to a .review.json
file for a person to check by hand.
No example in this repository has a non-stub result file yet: nothing here has been measured
against a live model, local or metered. docs/EVALS.md lists exactly which of the site’s
examples this question set scores and which it does not, and why: an example that does a
different task, such as extracting a record instead of answering a question, gets a different
measurement described on its own page rather than a meaningless number from this one.