# Turn a measurement session into a report somebody can review

_Recipe · needs level 1_

Turn computed measurements and notebook notes into a report. Code owns the figures, the model writes the prose, and a person checks the finished draft.


This is an engineering-test page. Five revision C prototypes of the Orbeck SRB-5030 spent three
days on the bench, swept over line, load and temperature, and the sweep is done: the margins are
computed, the uncertainty budget behind them is computed, the guardbanded verdicts are computed.
What is left is the afternoon's other job, turning a lab notebook
(`evals/bench/corpus/characterization-notebook.md`) and a table of numbers into a short
characterization report somebody else can read without re-deriving any of it. Code has already
decided every number in that report. What it has not decided is how to say it in sentences that
connect one finding to the next and do not quietly drop a caveat the notebook is honest about.

## The run, stepped

`compute_results` reads `evals/bench/data/characterization-2026-09.csv` and returns the corner
margins for all five boards, the uncertainty budget for the thinnest one, and board
SRB5030-2609-0005's line regulation figures and verdicts at all three ambients. The corner budget
is [the same one that page computes](/gradient_ascent/recipes/characterize-a-design/), for the
same five readings: board SRB5030-2609-0003 at 9.0 V in, 3.000 A out, 70 degC ambient, on the
10 V range, meter at 23 degC, one year calibration row.

| Contribution | Half-width or spread | Standard uncertainty |
| --- | --- | --- |
| Meter accuracy, 1 year, 10 V range, 23 degC | 222.2 uV | 128.3 uV |
| Resolution, 10 uV per count | 5.0 uV | 2.9 uV |
| Repeatability, 5 readings | s = 80.7 uV | 36.1 uV |
| Leads and connections | 200 uV | 115.5 uV |
| Combined | | 176.4 uV |
| Expanded, k = 2 | | 352.7 uV |

That number belongs to those five readings and to nothing else, which matters because two nearby
figures are in the documents this report is written from. The meter's manual works a budget for a
different reading and gets 349.5 uV, and the notebook works one for a generic reading of this
session and gets 351.6 uV. A report that quotes either of those for this corner is quoting the
wrong measurement, and the check below rejects it for that reason and not by accident.

Against that budget the corner margin is 20.4 mV, about fifty-eight times the expanded
uncertainty, so the measurement is not in doubt even though the margin is thin. Board
SRB5030-2609-0005's line regulation at 25 degC is a different case: 0.299 percent against a 0.300
percent limit, with 0.0075 percentage points of expanded uncertainty on the difference.
`guarded_verdict` returns `"cannot say"` there, `"pass"` at 0 degC and `"fail"` at 70 degC, and
all three verdicts are figures the model is handed, not conclusions it draws. The k of 2 is
stated with every one of those numbers because it is part of them: it covers roughly 95 percent
of where the true value could be, on the usual assumption that the contributions are independent
and roughly normal once they are combined, and it is not a guarantee that the value is inside.

Every figure is formatted once, as a `Figure(label, text)` pair, and that formatted string is the
only form the model is allowed to reproduce:

`examples/bench_measurement_writeup/run.py` (lines 174-239)

```python
def compute_results(csv_path: Path = CHARACTERIZATION_CSV) -> Results:
    """Read the characterization sweep and compute the corner margins, one uncertainty budget
    and the line-regulation verdicts this report needs. Nothing here is a model's arithmetic, and
    nothing the model is handed lets it redo this arithmetic differently.
    """
    rows = _read_rows(csv_path)
    serials = sorted({r["serial"] for r in rows})
    days = sorted({r["timestamp"][:10] for r in rows})

    figures: list[Figure] = [
        Figure("first day of the sweep", _mdy(days[0])),
        Figure("last day of the sweep", _mdy(days[-1])),
        Figure("output voltage minimum (V)", f"{VOUT_MIN_V:.3f}"),
        Figure("line regulation limit (%)", f"{LINE_REG_MAX_PCT:.3f}"),
        Figure("coverage factor (k)", f"{COVERAGE_FACTOR:.0f}"),
        Figure("corner ambient (degC)", CORNER_AMBIENT_TEXT),
        Figure("corner input voltage (V)", CORNER_VIN_V),
        Figure("corner load current (A)", CORNER_IOUT_A),
        Figure("uncertain block input voltage (V)", UNCERTAIN_BLOCK_VIN_TEXT),
        Figure("lead and connection uncertainty assumption (uV)", f"{LEAD_HALF_WIDTH_V * 1e6:.0f}"),
    ]

    corner_margins_mv: dict[str, float] = {}
    for serial in serials:
        mean = statistics.fmean(_readings(rows, serial, CORNER_TAMB_C, CORNER_VIN_V, CORNER_IOUT_A))
        margin_v = margin_to_limit(mean, VOUT_MIN_V, side="lower")
        corner_margins_mv[serial] = 1000.0 * margin_v
        figures.append(Figure(f"{serial} corner output voltage (V)", f"{mean:.5f}"))
        figures.append(Figure(f"{serial} corner margin (mV)", f"{1000.0 * margin_v:.1f}"))

    thin_serial = min(corner_margins_mv, key=corner_margins_mv.get)
    corner_readings = _readings(rows, thin_serial, CORNER_TAMB_C, CORNER_VIN_V, CORNER_IOUT_A)
    corner_expanded_v = expanded_uncertainty(
        combined_uncertainty(dc_voltage_budget(corner_readings, range_v=DC_RANGE_V))
    )
    corner_expanded_uv = 1e6 * corner_expanded_v
    corner_verdict = guarded_verdict(
        statistics.fmean(corner_readings), corner_expanded_v, lower=VOUT_MIN_V, upper=VOUT_MAX_V
    )
    figures.append(Figure("thin-margin board corner expanded uncertainty (uV)", f"{corner_expanded_uv:.1f}"))

    line_reg_pct = {tamb: line_regulation_pct(rows, MARGINAL_LINE_SERIAL, tamb) for tamb in ("0.0", "25.0", "70.0")}
    reg_readings = _readings(rows, MARGINAL_LINE_SERIAL, "25.0", "32.0", "1.000")
    per_reading = combined_uncertainty(
        dc_voltage_budget(reg_readings, range_v=DC_RANGE_V, lead_half_width_v=None)
    )
    reg_expanded_v = expanded_uncertainty(per_reading * math.sqrt(2.0))
    line_reg_uncertainty_pct = 100.0 * reg_expanded_v / VOUT_NOM_V
    figures.append(
        Figure("marginal-line board regulation uncertainty (percentage points)", f"{line_reg_uncertainty_pct:.4f}")
    )
    line_reg_verdicts: dict[str, str] = {}
    for tamb, pct in line_reg_pct.items():
        figures.append(Figure(f"marginal-line board regulation at {tamb} degC (%)", f"{pct:.3f}"))
        line_reg_verdicts[tamb] = guarded_verdict(pct, line_reg_uncertainty_pct, upper=LINE_REG_MAX_PCT)

    return Results(
        figures=tuple(figures),
        corner_margins_mv=corner_margins_mv,
        thin_serial=thin_serial,
        corner_expanded_uncertainty_uv=corner_expanded_uv,
        corner_verdict=corner_verdict,
        line_reg_pct=line_reg_pct,
        line_reg_uncertainty_pct=line_reg_uncertainty_pct,
        line_reg_verdicts=line_reg_verdicts,
    )
```

The notebook's own prose (`evals/bench/corpus/characterization-notebook.md`, loaded whole) and the
figures block go into one prompt. The system message is explicit that every number in the reply
has to be copied from the figures, character for character, and that anything else (a count, a
comparison between ambients) gets written in words, not digits. One call, and the model's only job
is the sentences:

`examples/bench_measurement_writeup/run.py` (lines 300-360)

```python
def run(
    notes: str,
    model: Model,
    tracer: Tracer,
    *,
    csv_path: Path = CHARACTERIZATION_CSV,
) -> Report:
    """One model call. `notes` is free text a requester may add (which finding to lead with, a
    house style note); it is appended to the prompt and never reaches the check, because it never
    contributes a number of its own.
    """
    results = compute_results(csv_path)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Compute the figures from the characterization data",
        detail=(
            f"{len(results.figures)} figures; thin-margin board {results.thin_serial} "
            f"({results.corner_margins_mv[results.thin_serial]:.1f} mV, {results.corner_verdict})"
        ),
    )

    notebook = load_bench_documents()["characterization-notebook"]
    tracer.record(kind="code", decided_by="code", title="Load the session notebook", detail=f"{len(notebook)} characters")

    user_parts = [
        f"Figures (use only these numbers, exactly as written):\n{_figures_block(results.figures)}",
        f"Notebook:\n{notebook}",
    ]
    if notes and notes.strip():
        user_parts.append(f"Additional guidance from the requester: {notes.strip()}")
    messages = [
        Message(role="system", content=SYSTEM_PROMPT),
        Message(role="user", content="\n\n".join(user_parts)),
    ]
    tracer.record(
        kind="code",
        decided_by="code",
        title="Build the prompt with the figures and the notebook",
        detail=f"{len(results.figures)} figures, {len(notebook)}-character notebook",
    )

    completion = model.complete(messages, max_tokens=700)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model to draft the report",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )

    unsupported = unsupported_numbers(completion.text, results.figures)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check the draft's numbers against the figures",
        detail="every number is supported" if not unsupported else f"unsupported: {', '.join(unsupported)}",
    )
    return Report(text=completion.text, figures=results.figures, unsupported=unsupported)
```

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls per session report:** 1
- **Tokens in / out:** 2,724 / 340
- **Figures computed and handed over:** 25
- **Cost per run, at typical rates:** a fraction of a cent

**Compared with characterize-a-design.** That page finds the same corner margin and the same cannot-say verdict for zero model calls: grouping, a subtraction, a root sum of squares. This page spends exactly one call on top of the same arithmetic, to turn the finding into prose a reviewer does not have to reassemble from a table themselves.

This is an engineering-test cost, counted per run and per session, not per unit: the notebook this
recipe reads took three days on the bench, and the report about it is written once. A tenth of a
cent a run is not a number worth optimizing; the number worth comparing it against is the afternoon
a person would otherwise spend writing the same paragraphs by hand, and the half hour lost if a
report ships with a number nobody checked, which is what the check above exists to catch before it
costs that much.

## How it fails on a real bench, specifically

### A right number attached to the wrong claim

- **How to notice it:** The draft says SRB5030-2609-0003 held 20.4 mV of margin, correctly, and then says SRB5030-2609-0002 held it, which is 58.2 mV in the figures. Both numbers are real figures code produced; the check only confirms each token appears somewhere among them, not which board or which claim it is sitting next to.
- **How to test for it:** A person checks the binding, not just the digits: for every figure the draft quotes, read the sentence around it against the label compute_results gave that figure. This is the one thing the automated check cannot do, and it is why a person still reads every report before it ships.

### A caveat from the notebook left out

- **How to notice it:** The notebook's own section 6 says the 200 uV lead and connection contribution is an assumption, not something measured on this harness, and that everything above it moves if that number is wrong. A draft that quotes every figure correctly but drops that sentence reads as more certain than the session actually was, and the check has nothing to compare an omission against: there is no missing number to flag.
- **How to test for it:** Keep a short list of the notebook's own open items (section 7) next to the draft and confirm each one that is still genuinely open appears somewhere in the report. A number-only check cannot do this; a person with the notebook open in the other window can, in under a minute.

### A short figure that is common enough to be a coincidence

- **How to notice it:** The coverage factor, 2, is one of this recipe's own figures, because docs/THE-BENCH.md requires stating it next to any expanded uncertainty. A bare "2" is also one of the most common tokens in ordinary prose, so the check would wave through an unrelated wrong number that happened to also be a bare 2. Nothing on this bench currently puts a fabricated single-digit figure next to a real finding, but the check's guarantee is about matching tokens, not about matching what a number means.
- **How to test for it:** Treat a short bare figure (one or two digits, no unit attached) as the one class of number the check is weakest on, and give it the same read-the-sentence attention as a claim-binding problem, not the confidence a five- or six-character figure earns.

### A model copies a real number from the notebook that code never produced

- **How to notice it:** The notebook goes into the prompt whole, and it carries its own uncertainty figure, 351.6 uV for a generic reading of this session rather than for the corner this report is about. It is a right number for the wrong reading, and it is only a microvolt away from the right one, so nothing about it looks wrong on the page.
- **How to test for it:** The check rejects it, because 351.6 is not one of the figures. When that happens, do not treat the check as the thing that is wrong: confirm which reading the report is about, and redraft from the figures rather than from numbers the source document worked out for itself.

## How to evaluate it

There is no 60-question set to grade this against: `scripts/eval_run.py` measures document
question answering over `evals/corpus/`, and this recipe reads a different document set and a
different data set entirely. What it measures instead is the share of drafts in which every figure
in the prose is one that code produced, which `unsupported_numbers` turns into a plain pass or a
fail rather than a judgment call: run a session's worth of drafts, count how many come back clean,
and a number under 100 percent means the prompt is losing the "copy the figure, do not compute
one" instruction, not that the arithmetic underneath is wrong.

That check is necessary and not sufficient, for the reasons the list above it gives. Before
trusting this recipe on a real session, keep a small set of past reports (five or six is enough at
this volume, since there is no golden run to score against and no second unit to catch a mistake
the way production volume would) and read each one against its own notebook, checking the claim
each number sits next to and the open items the report was supposed to carry forward.

## How to adapt it

Every figure above, and the notebook it is drawn from, is this bench's own. A real characterization
session has its own datasheet limits, its own uncertainty budget and its own notebook format; the
part that ports is the shape, not the numbers: compute every figure first, format each one exactly
once, hand the model the figures and the source text and nothing else it could compute a number
from, and check the draft's numbers against the figures before anyone reads it.

That shape is not specific to a bench. It is [turning
one piece of text into another](/gradient_ascent/techniques/prompt-engineering/) with the facts pinned down in advance: a test report written
from a results table a suite already produced, release notes drafted from a list of commits, a
status update written from what a tracker already says about a sprint. In every case the numbers or
the facts exist before the model is asked to write anything, so the same rule applies: a model
writes the sentences, code supplies and checks the facts, and nothing downstream trusts the prose
for a number it could have gotten from the table instead.

Where this stays level 1 and does not climb to the kind of pipeline
[accuracy-specs-from-the-manual](/gradient_ascent/recipes/accuracy-specs-from-the-manual/) needs,
a schema, a bounded retry and a person's recorded sign-off: one draft, read by one person, with
nothing else consuming it. It would move
to [write and check](/gradient_ascent/techniques/evaluator-optimizer/) the day a report started
generating unattended on a schedule with nobody reading each one before it went out, or the day
something downstream, a dashboard, another program, a second model, started parsing this report's
prose instead of a person reading it. Neither is true here, and a page that added a second model
pass anyway would be spending a call to guard against a risk this recipe does not have.

## Design choices

### Why this level, and when to use another approach

Code computes every figure this report can quote:
[the corner margin for all five boards](/gradient_ascent/recipes/characterize-a-design/), the
expanded uncertainty behind the thinnest one, and the three-ambient line regulation figures and
guardbanded verdict for the one board whose margin the measurement cannot quite decide. None of
that arithmetic is new to this page; it is the same `dc_voltage_budget`, `combined_uncertainty`,
`expanded_uncertainty`, `guarded_verdict` and `margin_to_limit` functions every other page on this
bench that touches uncertainty imports rather than reimplements. A model never produces a
reported measurement, an uncertainty, a margin or a verdict, here any more than on
[the production log](/gradient_ascent/recipes/limits-without-a-model/), where the same rule
shows up as a pass or a fail. This recipe's own job starts after all of that: turning the figures
and the notebook's prose into paragraphs a reviewer reads in under a minute.

That job is one model call, level 1, [order zero](/gradient_ascent/techniques/order-zero/)
composed with [prompt engineering](/gradient_ascent/techniques/prompt-engineering/). A fixed
template could produce a report too, but a session's story does not have a fixed shape: which
board is thinnest, whether a line regulation reading lands on a pass, a fail or a cannot-say,
which open items still matter, all of it changes sweep to sweep, and a template that covers every
combination is a maze of conditionals nobody checks with the rigor the arithmetic underneath it
gets.

The reasons to climb are reasons this page deliberately has none of. There is nothing to retrieve
that is not already in the notebook and the figures, so [RAG](/gradient_ascent/techniques/rag/)
buys nothing. The output is prose a person reads directly, not a payload another program parses,
so there is no schema for [structured output](/gradient_ascent/techniques/structured-output/) to
hold it to. Nothing here decides what happens next or calls a tool, so there is no case for an
agent loop.

## Build it

### Implementation details and code

`unsupported_numbers` is the whole argument for why this one call is safe to make unattended up to
the point a person reads the result. It scans the draft for every numeric token and returns
whichever ones are not, character for character, one of the figures code produced:

`examples/bench_measurement_writeup/run.py` (lines 260-279)

```python
def unsupported_numbers(draft: str, figures: Sequence[Figure]) -> tuple[str, ...]:
    """Every numeric token in `draft` that is not, character for character, one of `figures`'
    own text. In reading order, and it may repeat a token, so a report that leans on one bad
    number three times shows all three.

    Dates are matched first and checked whole, so 09/14/2026 is one token and not the three
    numbers 09, 14 and 2026. Identifiers are then blanked, so a serial number is not read as a
    quoted measurement. What is left is scanned for numbers.

    This is the whole safety argument for making one model call write a report nobody re-derives
    by hand: it costs two regular-expression scans and a set lookup, and it is a pass or a fail,
    never a judgment call. What it cannot do is check that a real figure is sitting next to the
    claim it belongs to, notice a caveat the draft dropped, or see a digit buried inside a word;
    this recipe's page names each of those and says what catches it instead.
    """
    allowed = {figure.text for figure in figures}
    scanned = _blank_identifiers(DATE_RE.sub(lambda m: " " * len(m.group()), draft))
    hits = [(m.start(), m.group()) for m in DATE_RE.finditer(draft)]
    hits += [(m.start(), m.group()) for m in NUMBER_RE.finditer(scanned)]
    return tuple(token for _, token in sorted(hits) if token not in allowed)
```

`tests/test_example_bench_measurement_writeup.py` runs it both ways: a clean draft, built only
from the figures above, comes back with nothing unsupported, and the same draft with its
uncertainty quietly rounded from 352.7 uV to a tidier 350 uV comes back naming `"350"`, the exact
token that was never one of the figures. The check does not retry and does not repair a draft it
rejects; `run` hands the rejected text back with the failing tokens named, so a review starts from
the specific sentence to look at rather than from "something in here is wrong."

Most of the work in writing that check is deciding what counts as one token, and the cases that
decide it are ordinary prose, not adversarial ones. A serial number has to be an identifier and
not three figures, or every board's ID would need pre-approving; that is what blanking
identifier-shaped tokens does before the numeric scan, and `SRB5030-2609-0003` never reaches it.
A date has to be one token too, or `09/14/2026` reads as 09, 14 and 2026 and a report cannot say
when the session ran, so dates are matched first and checked whole against the sweep's own first
and last day, which code takes off the CSV's timestamps. Everything left has to be matched to its
own edges: `352.7uV` with no space is still the figure 352.7 and not the number 352, `350uV` is
still an unsupported 350 rather than nothing at all, `20.4-99.9 mV` gives up both of its numbers
and not just the first, and a figure followed by a comma is the figure. Each of those is a test in
the file, and each of them was a way this check could have been quietly wrong while looking right.

What it still cannot catch is worth stating just as plainly:

- A digit buried inside a word or hyphenated onto one. `FIX99` and `degC-70` are identifier
  shaped, so they are blanked with the serials. This is the price of not pre-approving every ID.
- A figure spelled out. "three hundred fifty microvolts" has no digits in it to scan.
- A right figure sitting next to the wrong claim, and a caveat the notebook states that the draft
  drops. Neither is a number, so neither is a token. Both are in the failure modes below, and both
  are why a person still reads the report.



Last reviewed 2026-09-19.
