Recipe

Pull an instrument's accuracy table out of its manual

Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.

SourcedNeeds level 3

The number going in the report is 0.299 percent against a 0.300 percent limit, and somebody is going to ask how good the measurement is. That is board SRB5030-2609-0005 at 25 degC, and the answer decides whether the row reads pass or cannot say. Getting to it means turning the MDN-6100’s accuracy specification, a table in section 2 of its programming manual, into something code can compute from: one row per DC volts range per calibration interval, in parts per million of reading plus parts per million of range, with a temperature band and a coefficient for every degree outside it. This is the precise-measurement setting: one number that has to be right, priced once and reused for however many readings that meter takes until it is next calibrated or its manual is next revised. AccuracySpec in examples/common/bench.py is that schema already; this recipe is what fills it from the manual instead of from a person typing 15 rows into it by hand.

The walkthrough, on the bench

No recorded run exists for this recipe yet (docs/EVALS.md), so what follows is worked by hand from the same code and the same stub the tests exercise, not a played trace, and it is labeled illustrated because it is one.

run reads mdn6100-programming-manual.md section 2 and asks for all 15 rows as one JSON array. A clean reply validates on the first pass: every row is shaped right, all 15 combinations are there once each, and every range’s ppm numbers increase from the 24 hour row to the 1 year row. Several kinds of broken reply are easy for validation to catch and are not this page’s argument: a missing row, a duplicate, a negative or non-numeric field, and two intervals of the same range swapped for each other, which breaks the increasing sequence the same check already watches for. tests/test_example_bench_accuracy_specs_from_the_manual.py puts each of those through _validate_table and confirms it is reported, and scripts a reply that is not JSON at all through StubModel to show the one retry run allows.

The reply this page is about validates too. The scripted mistake puts the 100 V range’s numbers, 45 ppm of reading plus 6 ppm of range, under the 10 V range’s own 1 year row, which the manual prints as 35 plus 5. Nothing about that row is malformed: the range field still correctly says 10 V, the four numbers are all plausible, non-negative and present once, and the sequence for that range still increases, 12, 25, 45, because a range’s real accuracy generally gets looser with the range too, in the same direction the check is already watching for. _validate_table returns no problems for it. run holds the table anyway, the way it holds a clean one, because it never returns anything else: a person still has to read section 2 and check this row before it goes anywhere, and in the test that person finds the mismatch and calls confirm_table with approved=False.

Once a table is confirmed, pricing a reading from it is arithmetic that never touches a model again, and it is exactly as easy to get wrong a second way. Start with the meter alone, which is the first line of any budget built from this table and the one the wrong row moves. The manual works it out in section 2: a 4.9930 V reading on the 10 V range, one year specification, sits inside an accuracy limit of 224.8 uV. The same reading has a limit of only 79.9 uV on the 24 hour row, which is the row for a meter calibrated in the last day rather than eleven months ago, and 824.7 uV if it is priced against the 100 V range instead of the 10 V range it was taken on.

Those three are accuracy limits and not uncertainties, which is the distinction to keep hold of when reading them next to a budget. A limit says nothing about where inside it the meter sits, so it enters the budget as a rectangular contribution, divided by the square root of 3, alongside the display’s resolution, the repeatability of the readings and whatever the leads contribute. price_reading combines those by root sum of squares and expands at k = 2, so its own answer for that single reading is 259.6 uV on the 10 V row and 954.0 uV on the 100 V row: the same mistake, carried through to the number that would go in a report.

All three rows are real rows of the identical, correctly confirmed table. price_reading reaches the wrong one only when it is handed the wrong calibration age or the wrong range, which manual section 7 calls out by name as settings rather than computed values, and squarely in the user’s hands.

What it costs

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1 (2 if the reply fails validation)Model calls, one extraction
15Rows extracted per call
0, ever, for this meterModel calls once confirmed
Compared with typing the table in by handFor one meter this is not obviously cheaper: reading the manual closely enough to check the draft is most of the work of typing the table directly. The case for drafting is several meters and several manuals, where the checking stays per-row and the retyping does not have to.

Once confirmed, a table is priced against as many readings as that meter ever takes before its next calibration or its manual’s next revision, at no further model cost: the setting’s own unit, per measurement, is what section 8 of the manual already works out by hand for one reading, and price_reading is that same arithmetic run again for the next one and the one after that.

How it fails on a real bench, specifically

A row that borrowed the next range up

How to notice it
A range's accuracy row is shaped correctly, is the only row for its slot, and increases across intervals the way every real row does, and it is still the wrong row: its two ppm numbers came from the range above it, which happens to make the same check pass.
How to test for it
Run _validate_table on the table tests/test_example_bench_accuracy_specs_from_the_manual.py calls BAD_ROWS: it returns no problems. The recipe page exists because that test passes; the row-by-row read against mdn6100-programming-manual.md section 2 is what the confirm step is for, and confirm_table(result, False, tracer, note=...) is how the tests record it being caught.

The calibration row for a meter that was not just calibrated

How to notice it
A reading priced against the 24 hour row looks tighter than the same reading priced against the row for how long ago the meter was actually last calibrated, and nothing about the smaller number looks wrong on its own.
How to test for it
Call price_reading with days_since_cal=0.5 against a meter actually calibrated eleven months ago and it returns the 24 hour row's budget, 92.5 uV expanded, where the honest one is 259.6 uV. test_assuming_a_meter_was_just_calibrated_understates_the_uncertainty checks that it comes out smaller and never that it raises: price_reading trusts the calibration age it is given.

The range named is not the range the reading was taken on

How to notice it
A reading is priced against a bigger range than it was actually measured on, most plausibly because the meter auto-ranged up briefly and nobody logged it, and the result overstates the uncertainty by a real row of the same table rather than understating it.
How to test for it
Call price_reading with range_v=100.0 for a 4.9930 V reading actually taken on the 10 V range and the budget comes to 954.0 uV expanded instead of 259.6 uV, on a meter accuracy limit of 824.7 uV instead of 224.8 uV, which is the manual's own worked comparison in section 2. test_naming_the_wrong_range_overstates_the_uncertainty checks it.

What to measure

This example runs on evals/bench/, a second document set with no question file of its own, so evals/questions.json has nothing to grade it against (docs/EVALS.md). The right measurement is row accuracy against the manual, and it counts two mistakes separately, because they are caught differently. A value transcribed wrong inside an otherwise well-formed row, a digit or a decimal place off, is often visible to validation too, if it breaks the increasing sequence or drifts outside a plausible range. A right value taken from the wrong row never is; only a person comparing the extracted table against mdn6100-programming-manual.md section 2, cell by cell, catches it, which is why that comparison is the recipe and not a courtesy tacked onto the end of it. Fifteen rows is not enough to trust a rate from for either kind of mistake; collect a table per instrument as the lab’s own manuals get read this way, and keep every mismatch a person finds, not only the ones this bench happened to plant.

How to adapt it

What ports: the schema (one row per range per calibration interval, AccuracySpec’s own fields), the validation that checks shape, completeness and monotonicity without needing to already know the answer, and the rule that nothing downstream may use a table before a person has read it against the source. A reader’s own instrument states accuracy in its own form, sometimes percent of reading plus a fixed offset rather than ppm of reading plus ppm of range; MDN4010_VOLT_READBACK in examples/common/bench.py shows that conversion, since a fixed offset is a ppm-of-range term once the range is fixed.

What does not: every number on this page, the specific five ranges and three intervals, and the assumption that a manual is Markdown with one table to read rather than a PDF whose table a real extraction step would have to parse first. mdn6100-programming-manual.md was checked into this bench already parsed; a reader’s own manual was not.

The same shape, a document with a table in it turned into fixed-field records nothing downstream may use unconfirmed, pulls a datasheet’s key parameters into a parts database or a calibration certificate’s as-found and as-left readings into a drift record, the way document-extraction pulls a form’s fields into a patient record. What is specific to this page is the two kinds of wrong row, one code can generally catch and one it structurally cannot, and a reader adapting this to a parts database or a drift record should ask which of theirs plays the harder part.

Design choices

Why this level, and when to use another approach

Level 0 checks the extraction: shape, that all five ranges times three intervals are present once each, and that a range’s own accuracy only gets looser from the 24 hour row through the 1 year row, never tighter. Structured output holds the model to a fixed schema and retries once on a validation error, the same schema-then-retry pattern test-failure-triage uses for a cause label. Human approval is what actually earns this recipe its level: run never returns a table anyone may use, only a proposal, and confirm_table is where a person’s row-by-row read of the manual is recorded. That gate is not loosened here the way human approval’s own page allows when being wrong is cheap and easy to notice after the fact. A wrong row is neither. It is a real number, shaped exactly like the row next to it, and it stays wrong until somebody who has read the manual says so.

Nothing here climbs past level 3. There is one document, one schema and one bounded retry, and nothing in it chooses which step happens next from what a step found: the retry is the same call again with the validation error appended, and it happens at most once. Choosing the next step from the last one’s result is the case an agent loop has to be able to make for itself, and this recipe cannot make it. Once the table is confirmed, every budget built from it is arithmetic: price_reading calls combined_uncertainty and expanded_uncertainty, the same functions characterize-a-design runs on the characterization data, and no model runs again. A model never produces a reported measurement, an uncertainty, a margin or a verdict, on this page or anywhere else on this bench, and limits without a model is the production-test case of the same rule. Here a model only ever proposes what the manual’s table says, and code decides whether that proposal is even shaped like a table before a person is asked to read it against the real thing.

Whether drafting is worth it at all is a question with an honest smaller answer: for one meter, a person reading the manual closely enough to check a drafted table could have typed 15 rows into a spreadsheet in about the time it takes to read this section, and typing is not a slower way to make the identical mistake a wrong row makes. Where this earns its keep is scale a single meter does not have: a lab with instruments from several vendors, each with its own manual and its own row shape, a calibration house that reissues a table every time a meter comes back, a manual revision that changes a number nobody re-reads. The checking discipline, every row against the manual, has to stay exactly as strict per row in that world as it would for one meter; what drafting buys is not a lighter check, only fewer hours spent retyping tables that turn over.

Build it

Implementation details and code
View code: validate table
examples/bench_accuracy_specs_from_the_manual/run.py · lines 124–182
def _validate_table(rows: object) -> list[str]:
    """Shape, completeness and monotonicity: everything code can check without already knowing
    what the manual says. It cannot check that a row's numbers came from the cell they claim to,
    only that the 15 cells are all present once each, hold plausible numbers, and get no tighter
    from the 24 hour row to the 1 year row of the same range -- and a range's own numbers usually
    get looser with the range too, which is exactly why a row copied from the next range up still
    passes this last check.
    """
    if not isinstance(rows, list):
        return ["the extraction must be a JSON array of rows"]

    problems: list[str] = []
    clean: dict[tuple[str, float], dict] = {}
    for i, row in enumerate(rows):
        if not isinstance(row, dict):
            problems.append(f"row {i}: not an object")
            continue
        interval = row.get("interval")
        range_value = row.get("range_value")
        row_ok = True
        if interval not in INTERVALS:
            problems.append(f"row {i}: interval must be one of {INTERVALS}, got {interval!r}")
            row_ok = False
        if range_value not in RANGES_V:
            problems.append(f"row {i}: range_value must be one of {RANGES_V}, got {range_value!r}")
            row_ok = False
        for name in (
            "ppm_of_reading",
            "ppm_of_range",
            "tempco_ppm_of_reading_per_c",
            "tempco_ppm_of_range_per_c",
        ):
            value = row.get(name)
            if isinstance(value, bool) or not isinstance(value, (int, float)) or value < 0:
                problems.append(f"row {i}: {name} must be a non-negative number, got {value!r}")
                row_ok = False
        if not row_ok:
            continue
        key = (interval, float(range_value))
        if key in clean:
            problems.append(f"row {i}: duplicate row for interval={interval!r} range_value={range_value!r}")
            continue
        clean[key] = row

    expected = {(interval, range_value) for interval in INTERVALS for range_value in RANGES_V}
    missing = expected - set(clean)
    if missing:
        problems.append(f"missing rows: {sorted(missing)}")

    for range_value in RANGES_V:
        if not all((interval, range_value) in clean for interval in INTERVALS):
            continue  # already reported above
        for name in ("ppm_of_reading", "ppm_of_range"):
            values = [float(clean[(interval, range_value)][name]) for interval in INTERVALS]
            if any(a > b for a, b in zip(values, values[1:])):
                problems.append(
                    f"{range_value:g} V: {name} does not increase from 24 hour through 1 year: {values}"
                )
    return problems

run extracts, validates, retries once, and stops there. It returns a proposal, not a table:

View code: run
examples/bench_accuracy_specs_from_the_manual/run.py · lines 205–283
def run(
    section: str,
    model: Model,
    tracer: Tracer,
    *,
    sections: dict | None = None,
) -> ExtractionResult:
    sections = sections if sections is not None else load_bench_sections()
    if section not in sections:
        raise ValueError(f"{section!r} is not a section of the bench corpus")
    manual_text = sections[section].text
    tracer.record(kind="code", decided_by="code", title="Read the manual section", detail=section)

    messages = [
        Message(role="system", content=EXTRACT_SYSTEM),
        Message(role="user", content=manual_text),
    ]
    rows: list = []
    problems: list[str] = ["no attempt made"]
    for attempt in range(MAX_RETRIES + 1):
        completion = model.complete(messages, schema=TABLE_SCHEMA, max_tokens=1400)
        tracer.record(
            kind="model",
            decided_by="code",
            title="Extract the accuracy table" if attempt == 0 else "Extract again with the validation error",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        try:
            rows = json.loads(completion.text)
            problems = _validate_table(rows)
        except json.JSONDecodeError as exc:
            rows, problems = [], [f"invalid JSON: {exc}"]
        tracer.record(
            kind="code",
            decided_by="code",
            title="Validate shape, completeness and monotonicity",
            detail=(
                "; ".join(problems)
                if problems
                else f"{len(rows)} rows, every range and interval present once, ppm increases 24 hour through 1 year"
            ),
        )
        if not problems:
            break
        if attempt < MAX_RETRIES:
            messages.append(
                Message(
                    role="user",
                    content=f"That did not validate: {'; '.join(problems)}. Reply again with the corrected JSON array only.",
                )
            )

    if problems:
        tracer.record(
            kind="code",
            decided_by="code",
            title="Stop: the table never validated",
            detail=(
                f"gave up after {MAX_RETRIES} retry(ies): {'; '.join(problems)}; nothing for a "
                f"person to check, so nothing is returned to price a reading from"
            ),
        )
        raw = tuple(rows) if isinstance(rows, list) else ()
        return ExtractionResult(section=section, rows={}, raw=raw)

    table = _rows_to_table(rows)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Hold the table for a person's row-by-row check against the manual",
        detail=(
            f"{len(table)} rows; validation cannot see a row whose numbers came from the wrong "
            f"range or the wrong interval, only a person reading {section} can"
        ),
    )
    return ExtractionResult(section=section, rows=table, raw=tuple(rows))

confirm_table is where a person’s decision is actually recorded, the same shape as bench_test_failure_triage’s confirm and the bench’s own GuardedSupply.output_on:

View code: confirm table
examples/bench_accuracy_specs_from_the_manual/run.py · lines 286–304
def confirm_table(
    result: ExtractionResult,
    approved: bool,
    tracer: Tracer,
    *,
    note: str = "",
) -> dict[tuple[str, float], AccuracySpec]:
    """A person's decision on one proposed table. Never automatic, and never skipped: `run`
    returns a proposal, and this is where it becomes something `price_reading` may use, the same
    shape as `bench_test_failure_triage.confirm`."""
    tracer.record(
        kind="code",
        decided_by="code",
        title="Person confirms the table against the manual",
        detail=f"approved={approved} section={result.section}" + (f" note={note!r}" if note else ""),
    )
    if not approved:
        raise RowsNotConfirmed(note or "rejected: at least one row did not match the manual")
    return result.rows

Pricing a reading from the confirmed table calls no model. interval_for_calibration picks the row by three comparisons; price_reading is dc_voltage_budget read from this extraction’s own table instead of the one already coded in examples/common/bench.py:

View code: price reading
examples/bench_accuracy_specs_from_the_manual/run.py · lines 333–397
def price_reading(
    table: dict[tuple[str, float], AccuracySpec],
    readings_v: Sequence[float],
    *,
    range_v: float,
    days_since_cal: float,
    ambient_c: float = 23.0,
    lead_half_width_v: float | None = None,
) -> PricedReading:
    """The same four-line budget `examples.common.bench.dc_voltage_budget` computes, read from
    this extraction's own confirmed table instead of the one already coded in
    `examples/common/bench.py` -- a real instrument's table is not already coded anywhere until a
    run like this one puts it there. `range_v` and `days_since_cal` are facts about how the
    reading was actually taken, not choices this function makes: pass the wrong one and this
    returns a real number from a real row of the table, priced for a measurement that was not
    actually taken that way. Level 0 throughout; no model runs past `confirm_table`.
    """
    interval = interval_for_calibration(days_since_cal)
    try:
        spec = table[(interval, range_v)]
    except KeyError:
        raise ValueError(
            f"the confirmed table has no row for interval={interval!r} range_value={range_v!r}"
        ) from None

    values = [float(v) for v in readings_v]
    if not values:
        raise ValueError("price_reading needs at least one reading")
    mean_v = statistics.fmean(values)

    contributions = [
        Contribution(
            "meter accuracy",
            standard_uncertainty(spec.limit(mean_v, ambient_c)),
            f"{spec.interval} specification, {spec.range_value:g} V range, {ambient_c:g} degC",
        ),
        Contribution(
            "resolution",
            resolution_uncertainty(range_v),
            f"{reading_resolution(range_v) * 1e6:g} uV per count",
        ),
    ]
    if len(values) > 1:
        contributions.append(
            Contribution("repeatability", repeatability_uncertainty(values), f"{len(values)} readings")
        )
    if lead_half_width_v:
        contributions.append(
            Contribution(
                "leads and connections",
                standard_uncertainty(lead_half_width_v),
                f"+/-{lead_half_width_v * 1e6:g} uV, from the fixture record",
            )
        )

    combined = combined_uncertainty(contributions)
    expanded = expanded_uncertainty(combined)
    return PricedReading(
        interval=interval,
        range_v=range_v,
        mean_v=mean_v,
        contributions=tuple(contributions),
        combined_v=combined,
        expanded_v=expanded,
    )
Composition

Techniques this recipe uses

The highest level it needs is level 3.

When not to use a model

Sourced

How to tell when ordinary code, search or a form is enough.

Structured output

Sourced

Getting answers in a fixed format such as JSON.

Human approval

Sourced

Pausing for a person to approve or correct.

Same shape, other jobs

Pull structured data out of something unstructured

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Invoices and receipts into an accounting system
  • Key parameters from a datasheet into a parts database
  • An instrument accuracy table into rows per range and per calibration interval
  • A calibration certificate into as-found and as-left readings for a drift record
  • Operator failure notes into cause, location and severity
  • Resumes into a candidate record
  • Lab reports into a results table
  • Log lines into typed events

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page