Answer questions from a datasheet, a test spec and a change notice
Retrieval over the documents an engineer already has, answered with citations that can be checked. The case that matters is a change notice contradicting the datasheet on one number, where the right answer depends on the board revision. Level 2 is enough because one search finds the passage.
SourcedNeeds level 2
Before an SRB-5030 board goes on the bench, someone has to answer one question: how high can the input go. The number lives in the datasheet’s Recommended Operating Conditions table, section 3: 36.0 V. It also lives in ECN-2608-04, an engineering change notice that lowers it to 32.0 V for board revisions A and B and leaves 36.0 V standing only for revision C, a board that has not shipped yet. The datasheet itself has not been reissued to match. An engineer who opens the datasheet, the document everyone reaches for first, and stops at the number in the table gets 36.0 V, which is wrong for every revision A or B board actually in the building. The walkthrough below asks it the way a production line does, while a fixture is being set up for a revision. That is not the only setting it is asked in, and the second question this page works is the one a careful measurement asks.
This is what an engineer actually has: not one authoritative datasheet, but a stack of documents
that update each other. evals/bench/corpus/ holds the whole stack for the SRB-5030: the
datasheet, the test specification, four instrument programming manuals, a bill of materials,
design-review rules, two notebooks, a failure-analysis guide, a calibration procedure, and the
one change notice, 13 files and 87 numbered sections in total. Answering a question like this
means finding the right sections and saying which board revision the answer is for, with something
a person can go check. Nothing here produces a measurement, an uncertainty, a
margin or a verdict. This recipe answers what the documents say, which is a question that comes
up long before, and quite apart from, any single test;
limits without a model is where a measurement
meets a limit.
The walkthrough
No example on this site has called a live model yet (docs/EVALS.md), so nothing below is a
recorded trace. It is worked by hand from the same code and stub the tests exercise, and labeled
illustrated because it is one.
The question: “What is the maximum input voltage of the SRB-5030, revision B, per its
recommended operating conditions?” run() loads all 87 sections of evals/bench/corpus/,
embeds the question and every section, and keeps the top 8 by cosine similarity. Two of those
eight are the ones that matter: ecn-2608-04#1 (“Change,” the notice itself, ranked 1st of 8)
and srb5030-datasheet#3 (“Recommended Operating Conditions,” ranked 6th of 8). The other six
share some vocabulary but not the conflict. All eight go into one prompt, with a JSON schema
requiring an answer, an applies_to_revision, and a citations list.
A correctly scoped reply looks like this:
{
"answer": "The maximum input voltage is 32.0 V, not the datasheet's 36.0 V.",
"applies_to_revision": "A and B",
"citations": ["ecn-2608-04#1", "srb5030-datasheet#3"]
}
Code checks three things before this is handed back as the answer: every citation was actually
one of the eight retrieved sections, applies_to_revision is not blank, and, because
srb5030-datasheet#3 is cited, ecn-2608-04#1 is cited too, since the notice was retrieved and
does supersede that section. A reply missing any of those gets one retry with the specific
problem appended to the prompt; a reply still wrong after that comes back labeled unvalidated
rather than returned as an answer.
The second question is about the measurement rather than the board. The MDN-6100’s manual states DC volts accuracy per range and per calibration interval, so “how good is a 4.9930 V reading” has one answer per row: 79.9 uV on the 24 hour row, 224.8 uV on the one year row a meter calibrated eleven months ago is actually on, 824.7 uV if the reading was taken on the 100 V range instead of the 10 V. Retrieval is not the problem here, since one search returns the accuracy table first of 87 and the right row and the wrong one are inside it together. A figure quoted with no interval and no range named is the same wrong answer as 36.0 V quoted with no revision named: off a real row, wrong for the meter in the rack. The catch is the same field under another name, the range and the interval where this example requires a revision, and this example does not write it.
What it costs
One question through this example, scripted with the correctly scoped reply above, costs one
model call: 1,977 tokens in (the eight retrieved sections plus the schema and the question,
counted by count_tokens) and 42 tokens out, both pinned by a test so this page cannot drift
from the code. A reply that fails validation costs a second call of about the same size. Level 0
costs nothing: a keyword search runs in milliseconds over already-loaded text.
The unit here is not the board. This gets asked when a fixture is set up for a revision, when someone new joins the line, or when the question comes up in a design review: a shift where it is asked twenty times over is on the order of 40,000 tokens in and 840 out. That is small next to re-embedding all 87 sections, which this example does on every call; caching the corpus’s embeddings, which do not change between questions, is the obvious next optimization and is not done here.
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
How it fails on this bench
A number read off the datasheet, the notice unmentioned
- How to notice it
- The answer states 36.0 V for a revision A or B board and cites only the datasheet, naming no notice and no other document, even though the datasheet's own sentence next to the number says the figure has been superseded.
- How to test for it
- Script a reply that cites the Recommended Operating Conditions section without the change notice and confirm it is rejected whenever the notice was among the retrieved sections, rather than accepted because the citation it does have is real.
A search narrow enough to miss the notice entirely
- How to notice it
- The answer is confident and the citation list names only the datasheet, because retrieval asked for the top two or three sections rather than a wider set, and the notice never reached the candidates code could check.
- How to test for it
- Retrieve the top two sections for a phrasing close to the datasheet’s own wording and confirm the notice is not among them while the datasheet section is. Retrieval has to be wide enough, eight of 87 here, before the citation check has anything to check.
A number with no revision attached
- How to notice it
- The answer gives 32.0 V or 36.0 V with nothing saying which board revision it holds for: correct for one revision and silently wrong for the other.
- How to test for it
- Script a reply with applies_to_revision left blank and confirm it is rejected and retried rather than returned as the answer.
How to evaluate it
This example is not one scripts/eval_run.py scores (see docs/EVALS.md): that set is graded
against evals/corpus/, the Halvorsen appliance documents. A bench-specific eval would need its
own golden set, a dozen questions to start, over evals/bench/corpus/, each with a known right
value, a known right scope, and the citations a right answer has to include.
A right answer here is not just the correct number. It is the correct number, the scope it holds
for, and a citation to any retrieved notice that supersedes another retrieved passage. Grading
stays code: compare the parsed value and revision against the golden answer, and the citation set
against a must_cite list, the same citation hit rate document
Q&A’s eval section already tracks. The asymmetry worth watching is a reply that gets the
number right and stays silent on the scope, the revision on a board question and the range and
interval on a measurement one. A golden set with no such question in it never measures it.
How to adapt it
What ports: chunk by a corpus’s own numbered sections rather than writing a splitter, retrieve
broadly enough that a conflicting document is in the candidate set before code can check for it,
and require a reply to name what a value applies to rather than accepting a bare number.
SUPERSEDED_BY in _validate is the move worth keeping: a reply citing an older document without
the newer, retrieved one that supersedes it does not validate.
What does not: the corpus loader’s assumption that every document is Markdown with numbered ## N. Title headings (a real datasheet is usually a PDF, with tables that need their own parser),
the SUPERSEDED_BY map, which is hand knowledge about this one bench and not something code
discovered, and every number here. A reader’s own documents carry their own supersession history,
to hand-encode the same way or to replace with a general mechanism this example does not attempt.
The same shape shows up anywhere a newer document quietly changes what an older one says without the older one being reissued: an amendment against a contract, an errata sheet against a standard, one instrument’s calibration procedure revised without its programming manual catching up. Retrieve widely enough to see both documents in one pass, and never let an answer name a value without saying which version of the document it came from.
Design choices
Why this level, and when to use another approach
Level 0 here is a keyword search over the same 13 documents, BM25 over the section text, with a person reading the results. For the question below, that search puts the datasheet’s Recommended Operating Conditions section first and both of the change notice’s relevant sections second and third, of 87. The datasheet section a person reads first even states the problem outright: “The 36.0 V maximum in this table is superseded for revision A and revision B boards. See ecn-2608-04.md.” A reader who reads that sentence, not just the number in the table above it, is most of the way to the right answer with no model anywhere.
So the case for level 2 here is not that retrieval is hard: it plainly is not. It is that a person chasing one footnote at a time does not scale to a shop asking dozens of these questions a day, and a narrower lookup, “find the max input voltage in the datasheet”, can return the datasheet’s number and never reach the footnote at all. RAG, retrieving across the whole corpus rather than one document, and structured output, requiring the reply to name which revision it is for and rejecting a citation list that drops a retrieved notice, do that chasing in code every time instead of depending on whoever is reading that day.
Climbing to agentic RAG (level 5), where the model runs a second search depending on what the first one found, buys nothing extra here. That is worth saying plainly, because documents disagreeing about a revision-dependent number is normally a reason to climb. Three things make this corpus the exception. The two conflicting numbers use enough of the same words, “maximum,” “input voltage,” the revision letters, that one search returns both the datasheet section and the notice: 1st and 6th of the top 8 with the hashing stand-in embedder this repo tests against, and 1st and 3rd of 87 with BM25 and no embedder at all. The corpus is 87 sections, so a top-8 cut is a tenth of everything. And the notice names what it supersedes, in its own header and its first paragraph, which is what lets code check a citation list instead of hoping.
Change any of those three and the argument goes with it. A corpus of thousands of sections can leave the notice out of a candidate set with nothing to say so, and a notice that does not name what it supersedes gives the check below nothing to key on: finding it then means reading the first answer and searching again on what it said, which is the dependent second search level 5 exists for. On this corpus none of that holds, so the extra calls buy nothing.
Build it
Implementation details and code
The whole pipeline, retrieval through validation:
View code: run
def run(
question: str,
model: Model,
embedder: Embedder,
tracer: Tracer,
*,
corpus_dir: Path = BENCH_CORPUS_DIR,
top_k: int = TOP_K,
) -> Answer:
sections = load_sections(corpus_dir)
tracer.record(kind="code", decided_by="code", title="Chunk the bench documents", detail=f"{len(sections)} sections")
sources = _retrieve(question, sections, embedder, top_k)
retrieved = {s.cite for s in sources}
tracer.record(
kind="code",
decided_by="code",
title="Embed and retrieve top-k",
detail=", ".join(s.cite for s in sources),
)
blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
messages = [
Message(role="system", content=SYSTEM_PROMPT),
Message(role="user", content=f"Sources:\n\n{blocks}\n\nQuestion: {question}"),
]
tracer.record(kind="code", decided_by="code", title="Build prompt with sources and schema", detail=f"{len(sources)} sources")
record: dict = {}
for attempt in range(MAX_RETRIES + 1):
completion = model.complete(messages, schema=SCHEMA, max_tokens=400)
tracer.record(
kind="model",
decided_by="code",
title="Ask the model for a cited, revision-scoped answer" if attempt == 0 else "Ask again with the validation error",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
try:
record = json.loads(completion.text)
problems = _validate(record, retrieved)
except json.JSONDecodeError as exc:
record, problems = {}, [f"invalid JSON: {exc}"]
tracer.record(kind="code", decided_by="code", title="Validate the reply", detail="; ".join(problems) or "valid")
if not problems:
text = f"{record['answer']} (applies to: {record['applies_to_revision']})"
return Answer(text=text, citations=list(record["citations"]))
if attempt < MAX_RETRIES:
messages.append(
Message(role="user", content=f"That did not validate: {'; '.join(problems)}. Reply again with corrected JSON only.")
)
return Answer(text=json.dumps({"error": "did not validate after retry", "last": record}), citations=[])The check that makes the difference between this and plain RAG runs after every reply:
View code: validate
def _validate(record: dict, retrieved: set[str]) -> list[str]:
"""What a reply has to have before code will hand it back as the answer.
Two of these checks exist because of what this bench is for, not because of JSON schemas in
general: a citation has to be one of the sources code actually retrieved (a model cannot cite
a document it was never shown), and `applies_to_revision` has to be filled in, because a
number that is silently missing its revision is exactly the failure this recipe exists to
catch -- correct for one board revision and wrong for another, with nothing on the page to
tell them apart.
"""
problems = [f"missing field: {f}" for f in REQUIRED_FIELDS if f not in record]
if problems:
return problems
if not isinstance(record["answer"], str) or not record["answer"].strip():
problems.append("answer must be a non-empty string")
if not isinstance(record["applies_to_revision"], str) or not record["applies_to_revision"].strip():
problems.append("applies_to_revision must be a non-empty string")
if not isinstance(record["citations"], list) or not record["citations"]:
problems.append("citations must be a non-empty list")
else:
unknown = [c for c in record["citations"] if c not in retrieved]
if unknown:
problems.append(f"citation(s) not among the retrieved sources: {', '.join(unknown)}")
for superseded, superseding in SUPERSEDED_BY.items():
if superseded in record["citations"] and superseding in retrieved and superseding not in record["citations"]:
problems.append(f"cites {superseded} without {superseding}, which supersedes it")
return problemsEvery step above is decided_by: "code". Retrieval, the prompt, the schema and the retry count
are fixed before the model runs; the model composes the answer and names which of the eight
sections it used, and nothing it says chooses what code does next.
Techniques this recipe uses
The highest level it needs is level 2.
Retrieval-augmented generation (RAG)
MeasuredSearching your documents and giving the results to the model.
Answer questions from a body of documents
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- A policy handbook or a set of standard operating procedures
- Datasheets, errata and engineering change notices for the parts on a board
- Instrument programming manuals
- A calibration procedure and the records it requires
- Contracts and their amendments
- A codebase's design documents
- Product manuals for a support team
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page