The example is a small reporting tool: it reads trace files in the exact shape
examples/common/trace.py writes, and a price table the caller supplies, and reports cost and
latency per question and per level. No price is written into the code: a maker’s per-token price
is specific to one model and changes without notice, so hard-coding one here would eventually be
wrong and would read as this site’s own claim about a real price rather than the caller’s.
examples/ops/run.py · lines 74–87
def cost_for_trace(trace: dict, prices: PriceTable) -> QuestionCost:
tokens_in = sum(step["tokens_in"] for step in trace["steps"])
tokens_out = sum(step["tokens_out"] for step in trace["steps"])
ms = sum(step["ms"] for step in trace["steps"])
usd = _cost(trace["model_id"], tokens_in, tokens_out, prices)
return QuestionCost(
example=trace["example"],
level=trace["level"],
model_id=trace["model_id"],
tokens_in=tokens_in,
tokens_out=tokens_out,
ms=ms,
usd=usd,
)
A trace’s model id might not be in the price table at all: a model retired since the table was
built, a typo, a local model with no per-token price because nothing is billed for it. The
estimator reports that as usd: None, not as free, and names every unpriced model id it found so
the gap is visible instead of silently zeroed out:
examples/ops/run.py · lines 94–107
def summarize_by_level(costs: list[QuestionCost]) -> list[LevelSummary]:
summaries = []
for level in sorted({c.level for c in costs}):
subset = [c for c in costs if c.level == level]
priced = [c.usd for c in subset if c.usd is not None]
summaries.append(
LevelSummary(
level=level,
n=len(subset),
mean_usd=_mean(priced) if priced else None,
mean_ms=_mean([c.ms for c in subset]),
)
)
return summaries
python -m examples.ops --demo needs no files: it writes two synthetic trace files with a real
Tracer (one shaped like a level 1 chat call, one like the level 2 RAG run this site’s own rag
page illustrates) and prices them against a small table that is explicitly made up for the demo,
kept apart from examples.ops.run itself so nothing about the reusable code depends on it. A real
report would point --traces at recorded trace.json files and --prices at a table built from
a maker’s current, dated price page instead.
tests/test_example_ops.py uses only made-up prices (fake-small, fake-big), never a real
vendor’s figures, and checks the arithmetic directly: 1,000 input and 1,000 output tokens against
a $1/$2-per-1,000-token table comes to exactly $3.00, a trace with an unpriced model id reports
None rather than $0, and summarize_by_level groups and averages correctly across several
traces at the same level. That is the whole measurement for this example: it answers no question,
so the site’s 60-question set does not score it, and docs/EVALS.md says so with the reason.
This site’s trace.json is its own small format, built for stepping through one recorded run on
a page, not for production monitoring across every call a system makes. A real deployment
generally reaches for a shared standard instead: OpenTelemetry describes itself as “vendor- and
tool-agnostic”, an “observability framework and toolkit” for producing traces, metrics and logs
that many different backends can read, and says that “The backend (storage) and the frontend
(visualization) of telemetry data are intentionally left to other tools”[4]. These are
the same tokens-in, tokens-out and wall-time fields this example reads out of a trace file, but
emitted in a shape a tracing backend already knows how to store, query and alert on.