# Pull an instrument's accuracy table out of its manual

_Recipe · needs level 3_

Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.


The number going in the report is 0.299 percent against a 0.300 percent limit, and somebody is
going to ask how good the measurement is. That is
[board SRB5030-2609-0005 at 25 degC](/gradient_ascent/recipes/characterize-a-design/), and the
answer decides whether the row reads pass or cannot say. Getting to it means turning the
MDN-6100's accuracy specification, a table in section 2 of its programming manual, into something
code can compute from: one row per DC volts range per calibration interval, in parts per million
of reading plus parts per million of range, with a temperature band and a coefficient for every
degree outside it.
This is the precise-measurement setting: one number that has to be right, priced once and reused
for however many readings that meter takes until it is next calibrated or its manual is next
revised. `AccuracySpec` in `examples/common/bench.py` is that schema already; this recipe is what
fills it from the manual instead of from a person typing 15 rows into it by hand.

## The walkthrough, on the bench

No recorded run exists for this recipe yet (`docs/EVALS.md`), so what follows is worked by hand
from the same code and the same stub the tests exercise, not a played trace, and it is labeled
illustrated because it is one.

`run` reads `mdn6100-programming-manual.md` section 2 and asks for all 15 rows as one JSON array.
A clean reply validates on the first pass: every row is shaped right, all 15 combinations are
there once each, and every range's ppm numbers increase from the 24 hour row to the 1 year row.
Several kinds of broken reply are easy for validation to catch and are not this page's argument: a
missing row, a duplicate, a negative or non-numeric field, and two intervals of the same range
swapped for each other, which breaks the increasing sequence the same check already watches for.
`tests/test_example_bench_accuracy_specs_from_the_manual.py` puts each of those through
`_validate_table` and confirms it is reported, and scripts a reply that is not JSON at all
through `StubModel` to show the one retry `run` allows.

The reply this page is about validates too. The scripted mistake puts the 100 V range's numbers,
45 ppm of reading plus 6 ppm of range, under the 10 V range's own 1 year row, which the manual
prints as 35 plus 5. Nothing about that row is malformed: the range field still correctly says 10
V, the four numbers are all plausible, non-negative and present once, and the sequence for that
range still increases, 12, 25, 45, because a range's real accuracy generally gets looser with the
range too, in the same direction the check is already watching for. `_validate_table` returns no
problems for it. `run` holds the table anyway, the way it holds a clean one, because it never
returns anything else: a person still has to read section 2 and check this row before it goes
anywhere, and in the test that person finds the mismatch and calls `confirm_table` with
`approved=False`.

Once a table is confirmed, pricing a reading from it is arithmetic that never touches a model
again, and it is exactly as easy to get wrong a second way. Start with the meter alone, which is
the first line of any budget built from this table and the one the wrong row moves. The manual
works it out in section 2: a 4.9930 V reading on the 10 V range, one year specification, sits
inside an accuracy limit of 224.8 uV. The same reading has a limit of only 79.9 uV on the 24 hour
row, which is the row for a meter calibrated in the last day rather than eleven months ago, and
824.7 uV if it is priced against the 100 V range instead of the 10 V range it was taken on.

Those three are accuracy limits and not uncertainties, which is the distinction to keep hold of
when reading them next to a budget. A limit says nothing about where inside it the meter sits, so
it enters the budget as a rectangular contribution, divided by the square root of 3, alongside
the display's resolution, the repeatability of the readings and whatever the leads contribute.
`price_reading` combines those by root sum of squares and expands at k = 2, so its own answer for
that single reading is 259.6 uV on the 10 V row and 954.0 uV on the 100 V row: the same mistake,
carried through to the number that would go in a report.

All three rows are real rows of the identical, correctly confirmed table. `price_reading` reaches
the wrong one only when it is handed the wrong calibration age or the wrong range, which manual
section 7 calls out by name as settings rather than computed values, and squarely in the user's
hands.

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one extraction:** 1 (2 if the reply fails validation)
- **Rows extracted per call:** 15
- **Model calls once confirmed:** 0, ever, for this meter

**Compared with typing the table in by hand.** For one meter this is not obviously cheaper: reading the manual closely enough to check the draft is most of the work of typing the table directly. The case for drafting is several meters and several manuals, where the checking stays per-row and the retyping does not have to.

Once confirmed, a table is priced against as many readings as that meter ever takes before its
next calibration or its manual's next revision, at no further model cost: the setting's own unit,
per measurement, is what section 8 of the manual already works out by hand for one reading, and
`price_reading` is that same arithmetic run again for the next one and the one after that.

## How it fails on a real bench, specifically

### A row that borrowed the next range up

- **How to notice it:** A range's accuracy row is shaped correctly, is the only row for its slot, and increases across intervals the way every real row does, and it is still the wrong row: its two ppm numbers came from the range above it, which happens to make the same check pass.
- **How to test for it:** Run _validate_table on the table tests/test_example_bench_accuracy_specs_from_the_manual.py calls BAD_ROWS: it returns no problems. The recipe page exists because that test passes; the row-by-row read against mdn6100-programming-manual.md section 2 is what the confirm step is for, and confirm_table(result, False, tracer, note=...) is how the tests record it being caught.

### The calibration row for a meter that was not just calibrated

- **How to notice it:** A reading priced against the 24 hour row looks tighter than the same reading priced against the row for how long ago the meter was actually last calibrated, and nothing about the smaller number looks wrong on its own.
- **How to test for it:** Call price_reading with days_since_cal=0.5 against a meter actually calibrated eleven months ago and it returns the 24 hour row's budget, 92.5 uV expanded, where the honest one is 259.6 uV. test_assuming_a_meter_was_just_calibrated_understates_the_uncertainty checks that it comes out smaller and never that it raises: price_reading trusts the calibration age it is given.

### The range named is not the range the reading was taken on

- **How to notice it:** A reading is priced against a bigger range than it was actually measured on, most plausibly because the meter auto-ranged up briefly and nobody logged it, and the result overstates the uncertainty by a real row of the same table rather than understating it.
- **How to test for it:** Call price_reading with range_v=100.0 for a 4.9930 V reading actually taken on the 10 V range and the budget comes to 954.0 uV expanded instead of 259.6 uV, on a meter accuracy limit of 824.7 uV instead of 224.8 uV, which is the manual's own worked comparison in section 2. test_naming_the_wrong_range_overstates_the_uncertainty checks it.

## What to measure

This example runs on `evals/bench/`, a second document set with no question file of its own, so
`evals/questions.json` has nothing to grade it against (`docs/EVALS.md`). The right measurement is
row accuracy against the manual, and it counts two mistakes separately, because they are caught
differently. A value transcribed wrong inside an otherwise well-formed row, a digit or a decimal
place off, is often visible to validation too, if it breaks the increasing sequence or drifts
outside a plausible range. A right value taken from the wrong row never is; only a person
comparing the extracted table against `mdn6100-programming-manual.md` section 2, cell by cell,
catches it, which is why that comparison is the recipe and not a courtesy tacked onto the end of
it. Fifteen rows is not enough to trust a rate from for either kind of mistake; collect a table
per instrument as the lab's own manuals get read this way, and keep every mismatch a person finds,
not only the ones this bench happened to plant.

## How to adapt it

What ports: the schema (one row per range per calibration interval, `AccuracySpec`'s own fields),
the validation that checks shape, completeness and monotonicity without needing to already know
the answer, and the rule that nothing downstream may use a table before a person has read it
against the source. A reader's own instrument states accuracy in its own form, sometimes percent
of reading plus a fixed offset rather than ppm of reading plus ppm of range; `MDN4010_VOLT_READBACK`
in `examples/common/bench.py` shows that conversion, since a fixed offset is a ppm-of-range term
once the range is fixed.

What does not: every number on this page, the specific five ranges and three intervals, and the
assumption that a manual is Markdown with one table to read rather than a PDF whose table a real
extraction step would have to parse first. `mdn6100-programming-manual.md` was checked into this
bench already parsed; a reader's own manual was not.

The same shape, a document with a table in it turned into fixed-field records nothing downstream
may use unconfirmed, pulls a datasheet's key parameters into a parts database or a calibration
certificate's as-found and as-left readings into a drift record, the way
[document-extraction](/gradient_ascent/recipes/document-extraction/) pulls a form's fields into a
patient record. What is specific to this page is the two kinds of wrong row, one code can
generally catch and one it structurally cannot, and a reader adapting this to a parts database or
a drift record should ask which of theirs plays the harder part.

## Design choices

### Why this level, and when to use another approach

Level 0 checks the extraction: shape, that all five ranges times three intervals are present
once each, and that a range's own accuracy only gets looser from the 24 hour row through the 1
year row, never tighter. [Structured output](/gradient_ascent/techniques/structured-output/)
holds the model to a fixed schema and retries once on a validation error, the same
schema-then-retry pattern
[test-failure-triage](/gradient_ascent/recipes/test-failure-triage/) uses for a cause label.
[Human approval](/gradient_ascent/techniques/human-in-the-loop/) is what actually earns this
recipe its level: `run` never returns a table anyone may use, only a proposal, and
`confirm_table` is where a person's row-by-row read of the manual is recorded. That gate is not
loosened here the way human approval's own page allows when being wrong is cheap and easy to
notice after the fact. A wrong row is neither. It is a real number, shaped exactly like the row
next to it, and it stays wrong until somebody who has read the manual says so.

Nothing here climbs past level 3. There is one document, one schema and one bounded retry, and
nothing in it chooses which step happens next from what a step found: the retry is the same call
again with the validation error appended, and it happens at most once. Choosing the next step
from the last one's result is the case an agent loop has to be able to make for itself, and this
recipe cannot make it. Once the table is confirmed, every budget built from it is arithmetic:
`price_reading` calls `combined_uncertainty` and `expanded_uncertainty`, the same functions
[characterize-a-design](/gradient_ascent/recipes/characterize-a-design/) runs on the
characterization data, and no model runs again. A model never produces a reported measurement, an
uncertainty, a margin or a verdict, on this page or anywhere else on this bench, and
[limits without a model](/gradient_ascent/recipes/limits-without-a-model/) is the production-test
case of the same rule. Here a model only ever proposes what the manual's table says, and code
decides whether that proposal is even shaped like a table before a person is asked to read it
against the real thing.

Whether drafting is worth it at all is a question with an honest smaller answer: for one meter, a
person reading the manual closely enough to check a drafted table could have typed 15 rows into a
spreadsheet in about the time it takes to read this section, and typing is not a slower way to
make the identical mistake a wrong row makes. Where this earns its keep is scale a single meter
does not have: a lab with instruments from several vendors, each with its own manual and its own
row shape, a calibration house that reissues a table every time a meter comes back, a manual
revision that changes a number nobody re-reads. The checking discipline, every row against the
manual, has to stay exactly as strict per row in that world as it would for one meter; what
drafting buys is not a lighter check, only fewer hours spent retyping tables that turn over.

## Build it

### Implementation details and code

`examples/bench_accuracy_specs_from_the_manual/run.py` (lines 124-182)

```python
def _validate_table(rows: object) -> list[str]:
    """Shape, completeness and monotonicity: everything code can check without already knowing
    what the manual says. It cannot check that a row's numbers came from the cell they claim to,
    only that the 15 cells are all present once each, hold plausible numbers, and get no tighter
    from the 24 hour row to the 1 year row of the same range -- and a range's own numbers usually
    get looser with the range too, which is exactly why a row copied from the next range up still
    passes this last check.
    """
    if not isinstance(rows, list):
        return ["the extraction must be a JSON array of rows"]

    problems: list[str] = []
    clean: dict[tuple[str, float], dict] = {}
    for i, row in enumerate(rows):
        if not isinstance(row, dict):
            problems.append(f"row {i}: not an object")
            continue
        interval = row.get("interval")
        range_value = row.get("range_value")
        row_ok = True
        if interval not in INTERVALS:
            problems.append(f"row {i}: interval must be one of {INTERVALS}, got {interval!r}")
            row_ok = False
        if range_value not in RANGES_V:
            problems.append(f"row {i}: range_value must be one of {RANGES_V}, got {range_value!r}")
            row_ok = False
        for name in (
            "ppm_of_reading",
            "ppm_of_range",
            "tempco_ppm_of_reading_per_c",
            "tempco_ppm_of_range_per_c",
        ):
            value = row.get(name)
            if isinstance(value, bool) or not isinstance(value, (int, float)) or value < 0:
                problems.append(f"row {i}: {name} must be a non-negative number, got {value!r}")
                row_ok = False
        if not row_ok:
            continue
        key = (interval, float(range_value))
        if key in clean:
            problems.append(f"row {i}: duplicate row for interval={interval!r} range_value={range_value!r}")
            continue
        clean[key] = row

    expected = {(interval, range_value) for interval in INTERVALS for range_value in RANGES_V}
    missing = expected - set(clean)
    if missing:
        problems.append(f"missing rows: {sorted(missing)}")

    for range_value in RANGES_V:
        if not all((interval, range_value) in clean for interval in INTERVALS):
            continue  # already reported above
        for name in ("ppm_of_reading", "ppm_of_range"):
            values = [float(clean[(interval, range_value)][name]) for interval in INTERVALS]
            if any(a > b for a, b in zip(values, values[1:])):
                problems.append(
                    f"{range_value:g} V: {name} does not increase from 24 hour through 1 year: {values}"
                )
    return problems
```

`run` extracts, validates, retries once, and stops there. It returns a proposal, not a table:

`examples/bench_accuracy_specs_from_the_manual/run.py` (lines 205-283)

```python
def run(
    section: str,
    model: Model,
    tracer: Tracer,
    *,
    sections: dict | None = None,
) -> ExtractionResult:
    sections = sections if sections is not None else load_bench_sections()
    if section not in sections:
        raise ValueError(f"{section!r} is not a section of the bench corpus")
    manual_text = sections[section].text
    tracer.record(kind="code", decided_by="code", title="Read the manual section", detail=section)

    messages = [
        Message(role="system", content=EXTRACT_SYSTEM),
        Message(role="user", content=manual_text),
    ]
    rows: list = []
    problems: list[str] = ["no attempt made"]
    for attempt in range(MAX_RETRIES + 1):
        completion = model.complete(messages, schema=TABLE_SCHEMA, max_tokens=1400)
        tracer.record(
            kind="model",
            decided_by="code",
            title="Extract the accuracy table" if attempt == 0 else "Extract again with the validation error",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        try:
            rows = json.loads(completion.text)
            problems = _validate_table(rows)
        except json.JSONDecodeError as exc:
            rows, problems = [], [f"invalid JSON: {exc}"]
        tracer.record(
            kind="code",
            decided_by="code",
            title="Validate shape, completeness and monotonicity",
            detail=(
                "; ".join(problems)
                if problems
                else f"{len(rows)} rows, every range and interval present once, ppm increases 24 hour through 1 year"
            ),
        )
        if not problems:
            break
        if attempt < MAX_RETRIES:
            messages.append(
                Message(
                    role="user",
                    content=f"That did not validate: {'; '.join(problems)}. Reply again with the corrected JSON array only.",
                )
            )

    if problems:
        tracer.record(
            kind="code",
            decided_by="code",
            title="Stop: the table never validated",
            detail=(
                f"gave up after {MAX_RETRIES} retry(ies): {'; '.join(problems)}; nothing for a "
                f"person to check, so nothing is returned to price a reading from"
            ),
        )
        raw = tuple(rows) if isinstance(rows, list) else ()
        return ExtractionResult(section=section, rows={}, raw=raw)

    table = _rows_to_table(rows)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Hold the table for a person's row-by-row check against the manual",
        detail=(
            f"{len(table)} rows; validation cannot see a row whose numbers came from the wrong "
            f"range or the wrong interval, only a person reading {section} can"
        ),
    )
    return ExtractionResult(section=section, rows=table, raw=tuple(rows))
```

`confirm_table` is where a person's decision is actually recorded, the same shape as
`bench_test_failure_triage`'s `confirm` and the bench's own `GuardedSupply.output_on`:

`examples/bench_accuracy_specs_from_the_manual/run.py` (lines 286-304)

```python
def confirm_table(
    result: ExtractionResult,
    approved: bool,
    tracer: Tracer,
    *,
    note: str = "",
) -> dict[tuple[str, float], AccuracySpec]:
    """A person's decision on one proposed table. Never automatic, and never skipped: `run`
    returns a proposal, and this is where it becomes something `price_reading` may use, the same
    shape as `bench_test_failure_triage.confirm`."""
    tracer.record(
        kind="code",
        decided_by="code",
        title="Person confirms the table against the manual",
        detail=f"approved={approved} section={result.section}" + (f" note={note!r}" if note else ""),
    )
    if not approved:
        raise RowsNotConfirmed(note or "rejected: at least one row did not match the manual")
    return result.rows
```

Pricing a reading from the confirmed table calls no model. `interval_for_calibration` picks the
row by three comparisons; `price_reading` is `dc_voltage_budget` read from this extraction's own
table instead of the one already coded in `examples/common/bench.py`:

`examples/bench_accuracy_specs_from_the_manual/run.py` (lines 333-397)

```python
def price_reading(
    table: dict[tuple[str, float], AccuracySpec],
    readings_v: Sequence[float],
    *,
    range_v: float,
    days_since_cal: float,
    ambient_c: float = 23.0,
    lead_half_width_v: float | None = None,
) -> PricedReading:
    """The same four-line budget `examples.common.bench.dc_voltage_budget` computes, read from
    this extraction's own confirmed table instead of the one already coded in
    `examples/common/bench.py` -- a real instrument's table is not already coded anywhere until a
    run like this one puts it there. `range_v` and `days_since_cal` are facts about how the
    reading was actually taken, not choices this function makes: pass the wrong one and this
    returns a real number from a real row of the table, priced for a measurement that was not
    actually taken that way. Level 0 throughout; no model runs past `confirm_table`.
    """
    interval = interval_for_calibration(days_since_cal)
    try:
        spec = table[(interval, range_v)]
    except KeyError:
        raise ValueError(
            f"the confirmed table has no row for interval={interval!r} range_value={range_v!r}"
        ) from None

    values = [float(v) for v in readings_v]
    if not values:
        raise ValueError("price_reading needs at least one reading")
    mean_v = statistics.fmean(values)

    contributions = [
        Contribution(
            "meter accuracy",
            standard_uncertainty(spec.limit(mean_v, ambient_c)),
            f"{spec.interval} specification, {spec.range_value:g} V range, {ambient_c:g} degC",
        ),
        Contribution(
            "resolution",
            resolution_uncertainty(range_v),
            f"{reading_resolution(range_v) * 1e6:g} uV per count",
        ),
    ]
    if len(values) > 1:
        contributions.append(
            Contribution("repeatability", repeatability_uncertainty(values), f"{len(values)} readings")
        )
    if lead_half_width_v:
        contributions.append(
            Contribution(
                "leads and connections",
                standard_uncertainty(lead_half_width_v),
                f"+/-{lead_half_width_v * 1e6:g} uV, from the fixture record",
            )
        )

    combined = combined_uncertainty(contributions)
    expanded = expanded_uncertainty(combined)
    return PricedReading(
        interval=interval,
        range_v=range_v,
        mean_v=mean_v,
        contributions=tuple(contributions),
        combined_v=combined,
        expanded_v=expanded,
    )
```



Last reviewed 2026-09-19.
