Recipe

Turn a measurement session into a report somebody can review

Turn computed measurements and notebook notes into a report. Code owns the figures, the model writes the prose, and a person checks the finished draft.

SourcedNeeds level 1

This is an engineering-test page. Five revision C prototypes of the Orbeck SRB-5030 spent three days on the bench, swept over line, load and temperature, and the sweep is done: the margins are computed, the uncertainty budget behind them is computed, the guardbanded verdicts are computed. What is left is the afternoon’s other job, turning a lab notebook (evals/bench/corpus/characterization-notebook.md) and a table of numbers into a short characterization report somebody else can read without re-deriving any of it. Code has already decided every number in that report. What it has not decided is how to say it in sentences that connect one finding to the next and do not quietly drop a caveat the notebook is honest about.

The run, stepped

compute_results reads evals/bench/data/characterization-2026-09.csv and returns the corner margins for all five boards, the uncertainty budget for the thinnest one, and board SRB5030-2609-0005’s line regulation figures and verdicts at all three ambients. The corner budget is the same one that page computes, for the same five readings: board SRB5030-2609-0003 at 9.0 V in, 3.000 A out, 70 degC ambient, on the 10 V range, meter at 23 degC, one year calibration row.

Contribution Half-width or spread Standard uncertainty
Meter accuracy, 1 year, 10 V range, 23 degC 222.2 uV 128.3 uV
Resolution, 10 uV per count 5.0 uV 2.9 uV
Repeatability, 5 readings s = 80.7 uV 36.1 uV
Leads and connections 200 uV 115.5 uV
Combined 176.4 uV
Expanded, k = 2 352.7 uV

That number belongs to those five readings and to nothing else, which matters because two nearby figures are in the documents this report is written from. The meter’s manual works a budget for a different reading and gets 349.5 uV, and the notebook works one for a generic reading of this session and gets 351.6 uV. A report that quotes either of those for this corner is quoting the wrong measurement, and the check below rejects it for that reason and not by accident.

Against that budget the corner margin is 20.4 mV, about fifty-eight times the expanded uncertainty, so the measurement is not in doubt even though the margin is thin. Board SRB5030-2609-0005’s line regulation at 25 degC is a different case: 0.299 percent against a 0.300 percent limit, with 0.0075 percentage points of expanded uncertainty on the difference. guarded_verdict returns "cannot say" there, "pass" at 0 degC and "fail" at 70 degC, and all three verdicts are figures the model is handed, not conclusions it draws. The k of 2 is stated with every one of those numbers because it is part of them: it covers roughly 95 percent of where the true value could be, on the usual assumption that the contributions are independent and roughly normal once they are combined, and it is not a guarantee that the value is inside.

Every figure is formatted once, as a Figure(label, text) pair, and that formatted string is the only form the model is allowed to reproduce:

View code: compute results
examples/bench_measurement_writeup/run.py · lines 174–239
def compute_results(csv_path: Path = CHARACTERIZATION_CSV) -> Results:
    """Read the characterization sweep and compute the corner margins, one uncertainty budget
    and the line-regulation verdicts this report needs. Nothing here is a model's arithmetic, and
    nothing the model is handed lets it redo this arithmetic differently.
    """
    rows = _read_rows(csv_path)
    serials = sorted({r["serial"] for r in rows})
    days = sorted({r["timestamp"][:10] for r in rows})

    figures: list[Figure] = [
        Figure("first day of the sweep", _mdy(days[0])),
        Figure("last day of the sweep", _mdy(days[-1])),
        Figure("output voltage minimum (V)", f"{VOUT_MIN_V:.3f}"),
        Figure("line regulation limit (%)", f"{LINE_REG_MAX_PCT:.3f}"),
        Figure("coverage factor (k)", f"{COVERAGE_FACTOR:.0f}"),
        Figure("corner ambient (degC)", CORNER_AMBIENT_TEXT),
        Figure("corner input voltage (V)", CORNER_VIN_V),
        Figure("corner load current (A)", CORNER_IOUT_A),
        Figure("uncertain block input voltage (V)", UNCERTAIN_BLOCK_VIN_TEXT),
        Figure("lead and connection uncertainty assumption (uV)", f"{LEAD_HALF_WIDTH_V * 1e6:.0f}"),
    ]

    corner_margins_mv: dict[str, float] = {}
    for serial in serials:
        mean = statistics.fmean(_readings(rows, serial, CORNER_TAMB_C, CORNER_VIN_V, CORNER_IOUT_A))
        margin_v = margin_to_limit(mean, VOUT_MIN_V, side="lower")
        corner_margins_mv[serial] = 1000.0 * margin_v
        figures.append(Figure(f"{serial} corner output voltage (V)", f"{mean:.5f}"))
        figures.append(Figure(f"{serial} corner margin (mV)", f"{1000.0 * margin_v:.1f}"))

    thin_serial = min(corner_margins_mv, key=corner_margins_mv.get)
    corner_readings = _readings(rows, thin_serial, CORNER_TAMB_C, CORNER_VIN_V, CORNER_IOUT_A)
    corner_expanded_v = expanded_uncertainty(
        combined_uncertainty(dc_voltage_budget(corner_readings, range_v=DC_RANGE_V))
    )
    corner_expanded_uv = 1e6 * corner_expanded_v
    corner_verdict = guarded_verdict(
        statistics.fmean(corner_readings), corner_expanded_v, lower=VOUT_MIN_V, upper=VOUT_MAX_V
    )
    figures.append(Figure("thin-margin board corner expanded uncertainty (uV)", f"{corner_expanded_uv:.1f}"))

    line_reg_pct = {tamb: line_regulation_pct(rows, MARGINAL_LINE_SERIAL, tamb) for tamb in ("0.0", "25.0", "70.0")}
    reg_readings = _readings(rows, MARGINAL_LINE_SERIAL, "25.0", "32.0", "1.000")
    per_reading = combined_uncertainty(
        dc_voltage_budget(reg_readings, range_v=DC_RANGE_V, lead_half_width_v=None)
    )
    reg_expanded_v = expanded_uncertainty(per_reading * math.sqrt(2.0))
    line_reg_uncertainty_pct = 100.0 * reg_expanded_v / VOUT_NOM_V
    figures.append(
        Figure("marginal-line board regulation uncertainty (percentage points)", f"{line_reg_uncertainty_pct:.4f}")
    )
    line_reg_verdicts: dict[str, str] = {}
    for tamb, pct in line_reg_pct.items():
        figures.append(Figure(f"marginal-line board regulation at {tamb} degC (%)", f"{pct:.3f}"))
        line_reg_verdicts[tamb] = guarded_verdict(pct, line_reg_uncertainty_pct, upper=LINE_REG_MAX_PCT)

    return Results(
        figures=tuple(figures),
        corner_margins_mv=corner_margins_mv,
        thin_serial=thin_serial,
        corner_expanded_uncertainty_uv=corner_expanded_uv,
        corner_verdict=corner_verdict,
        line_reg_pct=line_reg_pct,
        line_reg_uncertainty_pct=line_reg_uncertainty_pct,
        line_reg_verdicts=line_reg_verdicts,
    )

The notebook’s own prose (evals/bench/corpus/characterization-notebook.md, loaded whole) and the figures block go into one prompt. The system message is explicit that every number in the reply has to be copied from the figures, character for character, and that anything else (a count, a comparison between ambients) gets written in words, not digits. One call, and the model’s only job is the sentences:

View code: run
examples/bench_measurement_writeup/run.py · lines 300–360
def run(
    notes: str,
    model: Model,
    tracer: Tracer,
    *,
    csv_path: Path = CHARACTERIZATION_CSV,
) -> Report:
    """One model call. `notes` is free text a requester may add (which finding to lead with, a
    house style note); it is appended to the prompt and never reaches the check, because it never
    contributes a number of its own.
    """
    results = compute_results(csv_path)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Compute the figures from the characterization data",
        detail=(
            f"{len(results.figures)} figures; thin-margin board {results.thin_serial} "
            f"({results.corner_margins_mv[results.thin_serial]:.1f} mV, {results.corner_verdict})"
        ),
    )

    notebook = load_bench_documents()["characterization-notebook"]
    tracer.record(kind="code", decided_by="code", title="Load the session notebook", detail=f"{len(notebook)} characters")

    user_parts = [
        f"Figures (use only these numbers, exactly as written):\n{_figures_block(results.figures)}",
        f"Notebook:\n{notebook}",
    ]
    if notes and notes.strip():
        user_parts.append(f"Additional guidance from the requester: {notes.strip()}")
    messages = [
        Message(role="system", content=SYSTEM_PROMPT),
        Message(role="user", content="\n\n".join(user_parts)),
    ]
    tracer.record(
        kind="code",
        decided_by="code",
        title="Build the prompt with the figures and the notebook",
        detail=f"{len(results.figures)} figures, {len(notebook)}-character notebook",
    )

    completion = model.complete(messages, max_tokens=700)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model to draft the report",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )

    unsupported = unsupported_numbers(completion.text, results.figures)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check the draft's numbers against the figures",
        detail="every number is supported" if not unsupported else f"unsupported: {', '.join(unsupported)}",
    )
    return Report(text=completion.text, figures=results.figures, unsupported=unsupported)

What it costs

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls per session report
2,724 / 340Tokens in / out
25Figures computed and handed over
a fraction of a centCost per run, at typical rates
Compared with characterize-a-designThat page finds the same corner margin and the same cannot-say verdict for zero model calls: grouping, a subtraction, a root sum of squares. This page spends exactly one call on top of the same arithmetic, to turn the finding into prose a reviewer does not have to reassemble from a table themselves.

This is an engineering-test cost, counted per run and per session, not per unit: the notebook this recipe reads took three days on the bench, and the report about it is written once. A tenth of a cent a run is not a number worth optimizing; the number worth comparing it against is the afternoon a person would otherwise spend writing the same paragraphs by hand, and the half hour lost if a report ships with a number nobody checked, which is what the check above exists to catch before it costs that much.

How it fails on a real bench, specifically

A right number attached to the wrong claim

How to notice it
The draft says SRB5030-2609-0003 held 20.4 mV of margin, correctly, and then says SRB5030-2609-0002 held it, which is 58.2 mV in the figures. Both numbers are real figures code produced; the check only confirms each token appears somewhere among them, not which board or which claim it is sitting next to.
How to test for it
A person checks the binding, not just the digits: for every figure the draft quotes, read the sentence around it against the label compute_results gave that figure. This is the one thing the automated check cannot do, and it is why a person still reads every report before it ships.

A caveat from the notebook left out

How to notice it
The notebook's own section 6 says the 200 uV lead and connection contribution is an assumption, not something measured on this harness, and that everything above it moves if that number is wrong. A draft that quotes every figure correctly but drops that sentence reads as more certain than the session actually was, and the check has nothing to compare an omission against: there is no missing number to flag.
How to test for it
Keep a short list of the notebook's own open items (section 7) next to the draft and confirm each one that is still genuinely open appears somewhere in the report. A number-only check cannot do this; a person with the notebook open in the other window can, in under a minute.

A short figure that is common enough to be a coincidence

How to notice it
The coverage factor, 2, is one of this recipe's own figures, because docs/THE-BENCH.md requires stating it next to any expanded uncertainty. A bare "2" is also one of the most common tokens in ordinary prose, so the check would wave through an unrelated wrong number that happened to also be a bare 2. Nothing on this bench currently puts a fabricated single-digit figure next to a real finding, but the check's guarantee is about matching tokens, not about matching what a number means.
How to test for it
Treat a short bare figure (one or two digits, no unit attached) as the one class of number the check is weakest on, and give it the same read-the-sentence attention as a claim-binding problem, not the confidence a five- or six-character figure earns.

A model copies a real number from the notebook that code never produced

How to notice it
The notebook goes into the prompt whole, and it carries its own uncertainty figure, 351.6 uV for a generic reading of this session rather than for the corner this report is about. It is a right number for the wrong reading, and it is only a microvolt away from the right one, so nothing about it looks wrong on the page.
How to test for it
The check rejects it, because 351.6 is not one of the figures. When that happens, do not treat the check as the thing that is wrong: confirm which reading the report is about, and redraft from the figures rather than from numbers the source document worked out for itself.

How to evaluate it

There is no 60-question set to grade this against: scripts/eval_run.py measures document question answering over evals/corpus/, and this recipe reads a different document set and a different data set entirely. What it measures instead is the share of drafts in which every figure in the prose is one that code produced, which unsupported_numbers turns into a plain pass or a fail rather than a judgment call: run a session’s worth of drafts, count how many come back clean, and a number under 100 percent means the prompt is losing the “copy the figure, do not compute one” instruction, not that the arithmetic underneath is wrong.

That check is necessary and not sufficient, for the reasons the list above it gives. Before trusting this recipe on a real session, keep a small set of past reports (five or six is enough at this volume, since there is no golden run to score against and no second unit to catch a mistake the way production volume would) and read each one against its own notebook, checking the claim each number sits next to and the open items the report was supposed to carry forward.

How to adapt it

Every figure above, and the notebook it is drawn from, is this bench’s own. A real characterization session has its own datasheet limits, its own uncertainty budget and its own notebook format; the part that ports is the shape, not the numbers: compute every figure first, format each one exactly once, hand the model the figures and the source text and nothing else it could compute a number from, and check the draft’s numbers against the figures before anyone reads it.

That shape is not specific to a bench. It is turning one piece of text into another with the facts pinned down in advance: a test report written from a results table a suite already produced, release notes drafted from a list of commits, a status update written from what a tracker already says about a sprint. In every case the numbers or the facts exist before the model is asked to write anything, so the same rule applies: a model writes the sentences, code supplies and checks the facts, and nothing downstream trusts the prose for a number it could have gotten from the table instead.

Where this stays level 1 and does not climb to the kind of pipeline accuracy-specs-from-the-manual needs, a schema, a bounded retry and a person’s recorded sign-off: one draft, read by one person, with nothing else consuming it. It would move to write and check the day a report started generating unattended on a schedule with nobody reading each one before it went out, or the day something downstream, a dashboard, another program, a second model, started parsing this report’s prose instead of a person reading it. Neither is true here, and a page that added a second model pass anyway would be spending a call to guard against a risk this recipe does not have.

Design choices

Why this level, and when to use another approach

Code computes every figure this report can quote: the corner margin for all five boards, the expanded uncertainty behind the thinnest one, and the three-ambient line regulation figures and guardbanded verdict for the one board whose margin the measurement cannot quite decide. None of that arithmetic is new to this page; it is the same dc_voltage_budget, combined_uncertainty, expanded_uncertainty, guarded_verdict and margin_to_limit functions every other page on this bench that touches uncertainty imports rather than reimplements. A model never produces a reported measurement, an uncertainty, a margin or a verdict, here any more than on the production log, where the same rule shows up as a pass or a fail. This recipe’s own job starts after all of that: turning the figures and the notebook’s prose into paragraphs a reviewer reads in under a minute.

That job is one model call, level 1, order zero composed with prompt engineering. A fixed template could produce a report too, but a session’s story does not have a fixed shape: which board is thinnest, whether a line regulation reading lands on a pass, a fail or a cannot-say, which open items still matter, all of it changes sweep to sweep, and a template that covers every combination is a maze of conditionals nobody checks with the rigor the arithmetic underneath it gets.

The reasons to climb are reasons this page deliberately has none of. There is nothing to retrieve that is not already in the notebook and the figures, so RAG buys nothing. The output is prose a person reads directly, not a payload another program parses, so there is no schema for structured output to hold it to. Nothing here decides what happens next or calls a tool, so there is no case for an agent loop.

Build it

Implementation details and code

unsupported_numbers is the whole argument for why this one call is safe to make unattended up to the point a person reads the result. It scans the draft for every numeric token and returns whichever ones are not, character for character, one of the figures code produced:

View code: unsupported numbers
examples/bench_measurement_writeup/run.py · lines 260–279
def unsupported_numbers(draft: str, figures: Sequence[Figure]) -> tuple[str, ...]:
    """Every numeric token in `draft` that is not, character for character, one of `figures`'
    own text. In reading order, and it may repeat a token, so a report that leans on one bad
    number three times shows all three.

    Dates are matched first and checked whole, so 09/14/2026 is one token and not the three
    numbers 09, 14 and 2026. Identifiers are then blanked, so a serial number is not read as a
    quoted measurement. What is left is scanned for numbers.

    This is the whole safety argument for making one model call write a report nobody re-derives
    by hand: it costs two regular-expression scans and a set lookup, and it is a pass or a fail,
    never a judgment call. What it cannot do is check that a real figure is sitting next to the
    claim it belongs to, notice a caveat the draft dropped, or see a digit buried inside a word;
    this recipe's page names each of those and says what catches it instead.
    """
    allowed = {figure.text for figure in figures}
    scanned = _blank_identifiers(DATE_RE.sub(lambda m: " " * len(m.group()), draft))
    hits = [(m.start(), m.group()) for m in DATE_RE.finditer(draft)]
    hits += [(m.start(), m.group()) for m in NUMBER_RE.finditer(scanned)]
    return tuple(token for _, token in sorted(hits) if token not in allowed)

tests/test_example_bench_measurement_writeup.py runs it both ways: a clean draft, built only from the figures above, comes back with nothing unsupported, and the same draft with its uncertainty quietly rounded from 352.7 uV to a tidier 350 uV comes back naming "350", the exact token that was never one of the figures. The check does not retry and does not repair a draft it rejects; run hands the rejected text back with the failing tokens named, so a review starts from the specific sentence to look at rather than from “something in here is wrong.”

Most of the work in writing that check is deciding what counts as one token, and the cases that decide it are ordinary prose, not adversarial ones. A serial number has to be an identifier and not three figures, or every board’s ID would need pre-approving; that is what blanking identifier-shaped tokens does before the numeric scan, and SRB5030-2609-0003 never reaches it. A date has to be one token too, or 09/14/2026 reads as 09, 14 and 2026 and a report cannot say when the session ran, so dates are matched first and checked whole against the sweep’s own first and last day, which code takes off the CSV’s timestamps. Everything left has to be matched to its own edges: 352.7uV with no space is still the figure 352.7 and not the number 352, 350uV is still an unsupported 350 rather than nothing at all, 20.4-99.9 mV gives up both of its numbers and not just the first, and a figure followed by a comma is the figure. Each of those is a test in the file, and each of them was a way this check could have been quietly wrong while looking right.

What it still cannot catch is worth stating just as plainly:

  • A digit buried inside a word or hyphenated onto one. FIX99 and degC-70 are identifier shaped, so they are blanked with the serials. This is the price of not pre-approving every ID.
  • A figure spelled out. “three hundred fifty microvolts” has no digits in it to scan.
  • A right figure sitting next to the wrong claim, and a caveat the notebook states that the draft drops. Neither is a number, so neither is a token. Both are in the failure modes below, and both are why a person still reads the report.
Composition

Techniques this recipe uses

The highest level it needs is level 1.

When not to use a model

Sourced

How to tell when ordinary code, search or a form is enough.

Prompt engineering

Sourced

Writing instructions that get consistent results.

Same shape, other jobs

Turn one piece of text into another

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Summarize a meeting transcript
  • Explain a compiler error or a stack trace
  • Write release notes from a list of commits
  • Rewrite a test procedure for a less experienced operator
  • Write a characterization report around numbers that are already computed
  • Translate a supplier's datasheet excerpt
  • Turn bullet points into a status report

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page