Pull an instrument's accuracy table out of its manual
Extract specification rows from a manual, validate their structure, and calculate uncertainty in code. A person verifies ranges, intervals, and conditions against the source.
SourcedNeeds level 3
The number going in the report is 0.299 percent against a 0.300 percent limit, and somebody is
going to ask how good the measurement is. That is
board SRB5030-2609-0005 at 25 degC, and the
answer decides whether the row reads pass or cannot say. Getting to it means turning the
MDN-6100’s accuracy specification, a table in section 2 of its programming manual, into something
code can compute from: one row per DC volts range per calibration interval, in parts per million
of reading plus parts per million of range, with a temperature band and a coefficient for every
degree outside it.
This is the precise-measurement setting: one number that has to be right, priced once and reused
for however many readings that meter takes until it is next calibrated or its manual is next
revised. AccuracySpec in examples/common/bench.py is that schema already; this recipe is what
fills it from the manual instead of from a person typing 15 rows into it by hand.
The walkthrough, on the bench
No recorded run exists for this recipe yet (docs/EVALS.md), so what follows is worked by hand
from the same code and the same stub the tests exercise, not a played trace, and it is labeled
illustrated because it is one.
run reads mdn6100-programming-manual.md section 2 and asks for all 15 rows as one JSON array.
A clean reply validates on the first pass: every row is shaped right, all 15 combinations are
there once each, and every range’s ppm numbers increase from the 24 hour row to the 1 year row.
Several kinds of broken reply are easy for validation to catch and are not this page’s argument: a
missing row, a duplicate, a negative or non-numeric field, and two intervals of the same range
swapped for each other, which breaks the increasing sequence the same check already watches for.
tests/test_example_bench_accuracy_specs_from_the_manual.py puts each of those through
_validate_table and confirms it is reported, and scripts a reply that is not JSON at all
through StubModel to show the one retry run allows.
The reply this page is about validates too. The scripted mistake puts the 100 V range’s numbers,
45 ppm of reading plus 6 ppm of range, under the 10 V range’s own 1 year row, which the manual
prints as 35 plus 5. Nothing about that row is malformed: the range field still correctly says 10
V, the four numbers are all plausible, non-negative and present once, and the sequence for that
range still increases, 12, 25, 45, because a range’s real accuracy generally gets looser with the
range too, in the same direction the check is already watching for. _validate_table returns no
problems for it. run holds the table anyway, the way it holds a clean one, because it never
returns anything else: a person still has to read section 2 and check this row before it goes
anywhere, and in the test that person finds the mismatch and calls confirm_table with
approved=False.
Once a table is confirmed, pricing a reading from it is arithmetic that never touches a model again, and it is exactly as easy to get wrong a second way. Start with the meter alone, which is the first line of any budget built from this table and the one the wrong row moves. The manual works it out in section 2: a 4.9930 V reading on the 10 V range, one year specification, sits inside an accuracy limit of 224.8 uV. The same reading has a limit of only 79.9 uV on the 24 hour row, which is the row for a meter calibrated in the last day rather than eleven months ago, and 824.7 uV if it is priced against the 100 V range instead of the 10 V range it was taken on.
Those three are accuracy limits and not uncertainties, which is the distinction to keep hold of
when reading them next to a budget. A limit says nothing about where inside it the meter sits, so
it enters the budget as a rectangular contribution, divided by the square root of 3, alongside
the display’s resolution, the repeatability of the readings and whatever the leads contribute.
price_reading combines those by root sum of squares and expands at k = 2, so its own answer for
that single reading is 259.6 uV on the 10 V row and 954.0 uV on the 100 V row: the same mistake,
carried through to the number that would go in a report.
All three rows are real rows of the identical, correctly confirmed table. price_reading reaches
the wrong one only when it is handed the wrong calibration age or the wrong range, which manual
section 7 calls out by name as settings rather than computed values, and squarely in the user’s
hands.
What it costs
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
Once confirmed, a table is priced against as many readings as that meter ever takes before its
next calibration or its manual’s next revision, at no further model cost: the setting’s own unit,
per measurement, is what section 8 of the manual already works out by hand for one reading, and
price_reading is that same arithmetic run again for the next one and the one after that.
How it fails on a real bench, specifically
A row that borrowed the next range up
- How to notice it
- A range's accuracy row is shaped correctly, is the only row for its slot, and increases across intervals the way every real row does, and it is still the wrong row: its two ppm numbers came from the range above it, which happens to make the same check pass.
- How to test for it
- Run _validate_table on the table tests/test_example_bench_accuracy_specs_from_the_manual.py calls BAD_ROWS: it returns no problems. The recipe page exists because that test passes; the row-by-row read against mdn6100-programming-manual.md section 2 is what the confirm step is for, and confirm_table(result, False, tracer, note=...) is how the tests record it being caught.
The calibration row for a meter that was not just calibrated
- How to notice it
- A reading priced against the 24 hour row looks tighter than the same reading priced against the row for how long ago the meter was actually last calibrated, and nothing about the smaller number looks wrong on its own.
- How to test for it
- Call price_reading with days_since_cal=0.5 against a meter actually calibrated eleven months ago and it returns the 24 hour row's budget, 92.5 uV expanded, where the honest one is 259.6 uV. test_assuming_a_meter_was_just_calibrated_understates_the_uncertainty checks that it comes out smaller and never that it raises: price_reading trusts the calibration age it is given.
The range named is not the range the reading was taken on
- How to notice it
- A reading is priced against a bigger range than it was actually measured on, most plausibly because the meter auto-ranged up briefly and nobody logged it, and the result overstates the uncertainty by a real row of the same table rather than understating it.
- How to test for it
- Call price_reading with range_v=100.0 for a 4.9930 V reading actually taken on the 10 V range and the budget comes to 954.0 uV expanded instead of 259.6 uV, on a meter accuracy limit of 824.7 uV instead of 224.8 uV, which is the manual's own worked comparison in section 2. test_naming_the_wrong_range_overstates_the_uncertainty checks it.
What to measure
This example runs on evals/bench/, a second document set with no question file of its own, so
evals/questions.json has nothing to grade it against (docs/EVALS.md). The right measurement is
row accuracy against the manual, and it counts two mistakes separately, because they are caught
differently. A value transcribed wrong inside an otherwise well-formed row, a digit or a decimal
place off, is often visible to validation too, if it breaks the increasing sequence or drifts
outside a plausible range. A right value taken from the wrong row never is; only a person
comparing the extracted table against mdn6100-programming-manual.md section 2, cell by cell,
catches it, which is why that comparison is the recipe and not a courtesy tacked onto the end of
it. Fifteen rows is not enough to trust a rate from for either kind of mistake; collect a table
per instrument as the lab’s own manuals get read this way, and keep every mismatch a person finds,
not only the ones this bench happened to plant.
How to adapt it
What ports: the schema (one row per range per calibration interval, AccuracySpec’s own fields),
the validation that checks shape, completeness and monotonicity without needing to already know
the answer, and the rule that nothing downstream may use a table before a person has read it
against the source. A reader’s own instrument states accuracy in its own form, sometimes percent
of reading plus a fixed offset rather than ppm of reading plus ppm of range; MDN4010_VOLT_READBACK
in examples/common/bench.py shows that conversion, since a fixed offset is a ppm-of-range term
once the range is fixed.
What does not: every number on this page, the specific five ranges and three intervals, and the
assumption that a manual is Markdown with one table to read rather than a PDF whose table a real
extraction step would have to parse first. mdn6100-programming-manual.md was checked into this
bench already parsed; a reader’s own manual was not.
The same shape, a document with a table in it turned into fixed-field records nothing downstream may use unconfirmed, pulls a datasheet’s key parameters into a parts database or a calibration certificate’s as-found and as-left readings into a drift record, the way document-extraction pulls a form’s fields into a patient record. What is specific to this page is the two kinds of wrong row, one code can generally catch and one it structurally cannot, and a reader adapting this to a parts database or a drift record should ask which of theirs plays the harder part.
Design choices
Why this level, and when to use another approach
Level 0 checks the extraction: shape, that all five ranges times three intervals are present
once each, and that a range’s own accuracy only gets looser from the 24 hour row through the 1
year row, never tighter. Structured output
holds the model to a fixed schema and retries once on a validation error, the same
schema-then-retry pattern
test-failure-triage uses for a cause label.
Human approval is what actually earns this
recipe its level: run never returns a table anyone may use, only a proposal, and
confirm_table is where a person’s row-by-row read of the manual is recorded. That gate is not
loosened here the way human approval’s own page allows when being wrong is cheap and easy to
notice after the fact. A wrong row is neither. It is a real number, shaped exactly like the row
next to it, and it stays wrong until somebody who has read the manual says so.
Nothing here climbs past level 3. There is one document, one schema and one bounded retry, and
nothing in it chooses which step happens next from what a step found: the retry is the same call
again with the validation error appended, and it happens at most once. Choosing the next step
from the last one’s result is the case an agent loop has to be able to make for itself, and this
recipe cannot make it. Once the table is confirmed, every budget built from it is arithmetic:
price_reading calls combined_uncertainty and expanded_uncertainty, the same functions
characterize-a-design runs on the
characterization data, and no model runs again. A model never produces a reported measurement, an
uncertainty, a margin or a verdict, on this page or anywhere else on this bench, and
limits without a model is the production-test
case of the same rule. Here a model only ever proposes what the manual’s table says, and code
decides whether that proposal is even shaped like a table before a person is asked to read it
against the real thing.
Whether drafting is worth it at all is a question with an honest smaller answer: for one meter, a person reading the manual closely enough to check a drafted table could have typed 15 rows into a spreadsheet in about the time it takes to read this section, and typing is not a slower way to make the identical mistake a wrong row makes. Where this earns its keep is scale a single meter does not have: a lab with instruments from several vendors, each with its own manual and its own row shape, a calibration house that reissues a table every time a meter comes back, a manual revision that changes a number nobody re-reads. The checking discipline, every row against the manual, has to stay exactly as strict per row in that world as it would for one meter; what drafting buys is not a lighter check, only fewer hours spent retyping tables that turn over.
Build it
Implementation details and code
View code: validate table
def _validate_table(rows: object) -> list[str]:
"""Shape, completeness and monotonicity: everything code can check without already knowing
what the manual says. It cannot check that a row's numbers came from the cell they claim to,
only that the 15 cells are all present once each, hold plausible numbers, and get no tighter
from the 24 hour row to the 1 year row of the same range -- and a range's own numbers usually
get looser with the range too, which is exactly why a row copied from the next range up still
passes this last check.
"""
if not isinstance(rows, list):
return ["the extraction must be a JSON array of rows"]
problems: list[str] = []
clean: dict[tuple[str, float], dict] = {}
for i, row in enumerate(rows):
if not isinstance(row, dict):
problems.append(f"row {i}: not an object")
continue
interval = row.get("interval")
range_value = row.get("range_value")
row_ok = True
if interval not in INTERVALS:
problems.append(f"row {i}: interval must be one of {INTERVALS}, got {interval!r}")
row_ok = False
if range_value not in RANGES_V:
problems.append(f"row {i}: range_value must be one of {RANGES_V}, got {range_value!r}")
row_ok = False
for name in (
"ppm_of_reading",
"ppm_of_range",
"tempco_ppm_of_reading_per_c",
"tempco_ppm_of_range_per_c",
):
value = row.get(name)
if isinstance(value, bool) or not isinstance(value, (int, float)) or value < 0:
problems.append(f"row {i}: {name} must be a non-negative number, got {value!r}")
row_ok = False
if not row_ok:
continue
key = (interval, float(range_value))
if key in clean:
problems.append(f"row {i}: duplicate row for interval={interval!r} range_value={range_value!r}")
continue
clean[key] = row
expected = {(interval, range_value) for interval in INTERVALS for range_value in RANGES_V}
missing = expected - set(clean)
if missing:
problems.append(f"missing rows: {sorted(missing)}")
for range_value in RANGES_V:
if not all((interval, range_value) in clean for interval in INTERVALS):
continue # already reported above
for name in ("ppm_of_reading", "ppm_of_range"):
values = [float(clean[(interval, range_value)][name]) for interval in INTERVALS]
if any(a > b for a, b in zip(values, values[1:])):
problems.append(
f"{range_value:g} V: {name} does not increase from 24 hour through 1 year: {values}"
)
return problemsrun extracts, validates, retries once, and stops there. It returns a proposal, not a table:
View code: run
def run(
section: str,
model: Model,
tracer: Tracer,
*,
sections: dict | None = None,
) -> ExtractionResult:
sections = sections if sections is not None else load_bench_sections()
if section not in sections:
raise ValueError(f"{section!r} is not a section of the bench corpus")
manual_text = sections[section].text
tracer.record(kind="code", decided_by="code", title="Read the manual section", detail=section)
messages = [
Message(role="system", content=EXTRACT_SYSTEM),
Message(role="user", content=manual_text),
]
rows: list = []
problems: list[str] = ["no attempt made"]
for attempt in range(MAX_RETRIES + 1):
completion = model.complete(messages, schema=TABLE_SCHEMA, max_tokens=1400)
tracer.record(
kind="model",
decided_by="code",
title="Extract the accuracy table" if attempt == 0 else "Extract again with the validation error",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
try:
rows = json.loads(completion.text)
problems = _validate_table(rows)
except json.JSONDecodeError as exc:
rows, problems = [], [f"invalid JSON: {exc}"]
tracer.record(
kind="code",
decided_by="code",
title="Validate shape, completeness and monotonicity",
detail=(
"; ".join(problems)
if problems
else f"{len(rows)} rows, every range and interval present once, ppm increases 24 hour through 1 year"
),
)
if not problems:
break
if attempt < MAX_RETRIES:
messages.append(
Message(
role="user",
content=f"That did not validate: {'; '.join(problems)}. Reply again with the corrected JSON array only.",
)
)
if problems:
tracer.record(
kind="code",
decided_by="code",
title="Stop: the table never validated",
detail=(
f"gave up after {MAX_RETRIES} retry(ies): {'; '.join(problems)}; nothing for a "
f"person to check, so nothing is returned to price a reading from"
),
)
raw = tuple(rows) if isinstance(rows, list) else ()
return ExtractionResult(section=section, rows={}, raw=raw)
table = _rows_to_table(rows)
tracer.record(
kind="code",
decided_by="code",
title="Hold the table for a person's row-by-row check against the manual",
detail=(
f"{len(table)} rows; validation cannot see a row whose numbers came from the wrong "
f"range or the wrong interval, only a person reading {section} can"
),
)
return ExtractionResult(section=section, rows=table, raw=tuple(rows))confirm_table is where a person’s decision is actually recorded, the same shape as
bench_test_failure_triage’s confirm and the bench’s own GuardedSupply.output_on:
View code: confirm table
def confirm_table(
result: ExtractionResult,
approved: bool,
tracer: Tracer,
*,
note: str = "",
) -> dict[tuple[str, float], AccuracySpec]:
"""A person's decision on one proposed table. Never automatic, and never skipped: `run`
returns a proposal, and this is where it becomes something `price_reading` may use, the same
shape as `bench_test_failure_triage.confirm`."""
tracer.record(
kind="code",
decided_by="code",
title="Person confirms the table against the manual",
detail=f"approved={approved} section={result.section}" + (f" note={note!r}" if note else ""),
)
if not approved:
raise RowsNotConfirmed(note or "rejected: at least one row did not match the manual")
return result.rowsPricing a reading from the confirmed table calls no model. interval_for_calibration picks the
row by three comparisons; price_reading is dc_voltage_budget read from this extraction’s own
table instead of the one already coded in examples/common/bench.py:
View code: price reading
def price_reading(
table: dict[tuple[str, float], AccuracySpec],
readings_v: Sequence[float],
*,
range_v: float,
days_since_cal: float,
ambient_c: float = 23.0,
lead_half_width_v: float | None = None,
) -> PricedReading:
"""The same four-line budget `examples.common.bench.dc_voltage_budget` computes, read from
this extraction's own confirmed table instead of the one already coded in
`examples/common/bench.py` -- a real instrument's table is not already coded anywhere until a
run like this one puts it there. `range_v` and `days_since_cal` are facts about how the
reading was actually taken, not choices this function makes: pass the wrong one and this
returns a real number from a real row of the table, priced for a measurement that was not
actually taken that way. Level 0 throughout; no model runs past `confirm_table`.
"""
interval = interval_for_calibration(days_since_cal)
try:
spec = table[(interval, range_v)]
except KeyError:
raise ValueError(
f"the confirmed table has no row for interval={interval!r} range_value={range_v!r}"
) from None
values = [float(v) for v in readings_v]
if not values:
raise ValueError("price_reading needs at least one reading")
mean_v = statistics.fmean(values)
contributions = [
Contribution(
"meter accuracy",
standard_uncertainty(spec.limit(mean_v, ambient_c)),
f"{spec.interval} specification, {spec.range_value:g} V range, {ambient_c:g} degC",
),
Contribution(
"resolution",
resolution_uncertainty(range_v),
f"{reading_resolution(range_v) * 1e6:g} uV per count",
),
]
if len(values) > 1:
contributions.append(
Contribution("repeatability", repeatability_uncertainty(values), f"{len(values)} readings")
)
if lead_half_width_v:
contributions.append(
Contribution(
"leads and connections",
standard_uncertainty(lead_half_width_v),
f"+/-{lead_half_width_v * 1e6:g} uV, from the fixture record",
)
)
combined = combined_uncertainty(contributions)
expanded = expanded_uncertainty(combined)
return PricedReading(
interval=interval,
range_v=range_v,
mean_v=mean_v,
contributions=tuple(contributions),
combined_v=combined,
expanded_v=expanded,
)Techniques this recipe uses
The highest level it needs is level 3.
Pull structured data out of something unstructured
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Invoices and receipts into an accounting system
- Key parameters from a datasheet into a parts database
- An instrument accuracy table into rows per range and per calibration interval
- A calibration certificate into as-found and as-left readings for a drift record
- Operator failure notes into cause, location and severity
- Resumes into a candidate record
- Lab reports into a results table
- Log lines into typed events
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page