# Answer questions from a datasheet, a test spec and a change notice

_Recipe · needs level 2_

Retrieval over the documents an engineer already has, answered with citations that can be checked. The case that matters is a change notice contradicting the datasheet on one number, where the right answer depends on the board revision. Level 2 is enough because one search finds the passage.


Before an SRB-5030 board goes on the bench, someone has to answer one question: how high can
the input go. The number lives in the datasheet's Recommended Operating Conditions table,
section 3: 36.0 V. It also lives in ECN-2608-04, an engineering change notice that lowers it to
32.0 V for board revisions A and B and leaves 36.0 V standing only for revision C, a board that
has not shipped yet. The datasheet itself has not been reissued to match. An engineer who opens
the datasheet, the document everyone reaches for first, and stops at the number in the table
gets 36.0 V, which is wrong for every revision A or B board actually in the building. The
walkthrough below asks it the way a production line does, while a fixture is being set up for a
revision. That is not the only setting it is asked in, and the second question this page works is
the one a careful measurement asks.

This is what an engineer actually has: not one authoritative datasheet, but a stack of documents
that update each other. `evals/bench/corpus/` holds the whole stack for the SRB-5030: the
datasheet, the test specification, four instrument programming manuals, a bill of materials,
design-review rules, two notebooks, a failure-analysis guide, a calibration procedure, and the
one change notice, 13 files and 87 numbered sections in total. Answering a question like this
means finding the right sections and saying which board revision the answer is for, with something
a person can go check. Nothing here produces a measurement, an uncertainty, a
margin or a verdict. This recipe answers what the documents say, which is a question that comes
up long before, and quite apart from, any single test;
[limits without a model](/gradient_ascent/recipes/limits-without-a-model/) is where a measurement
meets a limit.

## The walkthrough

No example on this site has called a live model yet (`docs/EVALS.md`), so nothing below is a
recorded trace. It is worked by hand from the same code and stub the tests exercise, and labeled
illustrated because it is one.

The question: "What is the maximum input voltage of the SRB-5030, revision B, per its
recommended operating conditions?" `run()` loads all 87 sections of `evals/bench/corpus/`,
embeds the question and every section, and keeps the top 8 by cosine similarity. Two of those
eight are the ones that matter: `ecn-2608-04#1` ("Change," the notice itself, ranked 1st of 8)
and `srb5030-datasheet#3` ("Recommended Operating Conditions," ranked 6th of 8). The other six
share some vocabulary but not the conflict. All eight go into one prompt, with a JSON schema
requiring an `answer`, an `applies_to_revision`, and a `citations` list.

A correctly scoped reply looks like this:

```
{
  "answer": "The maximum input voltage is 32.0 V, not the datasheet's 36.0 V.",
  "applies_to_revision": "A and B",
  "citations": ["ecn-2608-04#1", "srb5030-datasheet#3"]
}
```

Code checks three things before this is handed back as the answer: every citation was actually
one of the eight retrieved sections, `applies_to_revision` is not blank, and, because
`srb5030-datasheet#3` is cited, `ecn-2608-04#1` is cited too, since the notice was retrieved and
does supersede that section. A reply missing any of those gets one retry with the specific
problem appended to the prompt; a reply still wrong after that comes back labeled unvalidated
rather than returned as an answer.

The second question is about the measurement rather than the board. The MDN-6100's manual states
DC volts accuracy per range and per calibration interval, so "how good is a 4.9930 V reading" has
one answer per row: 79.9 uV on the 24 hour row, 224.8 uV on the one year row a meter calibrated
eleven months ago is actually on, 824.7 uV if the reading was taken on the 100 V range instead of
the 10 V. Retrieval is not the problem here, since one search returns the accuracy table first of
87 and the right row and the wrong one are inside it together. A figure quoted with no interval
and no range named is the same wrong answer as 36.0 V quoted with no revision named: off a real
row, wrong for the meter in the rack. The catch is the same field under another name, the range
and the interval where this example requires a revision, and this example does not write it.

## What it costs

One question through this example, scripted with the correctly scoped reply above, costs one
model call: 1,977 tokens in (the eight retrieved sections plus the schema and the question,
counted by `count_tokens`) and 42 tokens out, both pinned by a test so this page cannot drift
from the code. A reply that fails validation costs a second call of about the same size. Level 0
costs nothing: a keyword search runs in milliseconds over already-loaded text.

The unit here is not the board. This gets asked when a fixture is set up for a revision, when
someone new joins the line, or when the question comes up in a design review: a shift where it is
asked twenty times over is on the order of 40,000 tokens in and 840 out. That is small next to
re-embedding all 87 sections, which this example does on every call; caching the corpus's
embeddings, which do not change between questions, is the obvious next optimization and is not
done here.

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 1 (2 if the reply fails validation)
- **Tokens in:** ~1,980
- **Tokens out:** ~42
- **Sections retrieved:** 8 of 87

**Compared with level 0, keyword search.** No model call and no tokens: BM25 over the same 87 sections, with a person reading the result instead of a schema checking it.

## How it fails on this bench

### A number read off the datasheet, the notice unmentioned

- **How to notice it:** The answer states 36.0 V for a revision A or B board and cites only the datasheet, naming no notice and no other document, even though the datasheet's own sentence next to the number says the figure has been superseded.
- **How to test for it:** Script a reply that cites the Recommended Operating Conditions section without the change notice and confirm it is rejected whenever the notice was among the retrieved sections, rather than accepted because the citation it does have is real.

### A search narrow enough to miss the notice entirely

- **How to notice it:** The answer is confident and the citation list names only the datasheet, because retrieval asked for the top two or three sections rather than a wider set, and the notice never reached the candidates code could check.
- **How to test for it:** Retrieve the top two sections for a phrasing close to the datasheet’s own wording and confirm the notice is not among them while the datasheet section is. Retrieval has to be wide enough, eight of 87 here, before the citation check has anything to check.

### A number with no revision attached

- **How to notice it:** The answer gives 32.0 V or 36.0 V with nothing saying which board revision it holds for: correct for one revision and silently wrong for the other.
- **How to test for it:** Script a reply with applies_to_revision left blank and confirm it is rejected and retried rather than returned as the answer.

## How to evaluate it

This example is not one `scripts/eval_run.py` scores (see `docs/EVALS.md`): that set is graded
against `evals/corpus/`, the Halvorsen appliance documents. A bench-specific eval would need its
own golden set, a dozen questions to start, over `evals/bench/corpus/`, each with a known right
value, a known right scope, and the citations a right answer has to include.

A right answer here is not just the correct number. It is the correct number, the scope it holds
for, and a citation to any retrieved notice that supersedes another retrieved passage. Grading
stays code: compare the parsed value and revision against the golden answer, and the citation set
against a `must_cite` list, the same citation hit rate [document
Q&A](/gradient_ascent/recipes/document-qa/)'s eval section already tracks. The asymmetry worth watching is a reply that gets the
number right and stays silent on the scope, the revision on a board question and the range and
interval on a measurement one. A golden set with no such question in it never measures it.

## How to adapt it

What ports: chunk by a corpus's own numbered sections rather than writing a splitter, retrieve
broadly enough that a conflicting document is in the candidate set before code can check for it,
and require a reply to name what a value applies to rather than accepting a bare number.
`SUPERSEDED_BY` in `_validate` is the move worth keeping: a reply citing an older document without
the newer, retrieved one that supersedes it does not validate.

What does not: the corpus loader's assumption that every document is Markdown with numbered `##
N. Title` headings (a real datasheet is usually a PDF, with tables that need their own parser),
the `SUPERSEDED_BY` map, which is hand knowledge about this one bench and not something code
discovered, and every number here. A reader's own documents carry their own supersession history,
to hand-encode the same way or to replace with a general mechanism this example does not attempt.

The same shape shows up anywhere a newer document quietly changes what an older one says without
the older one being reissued: an amendment against a contract, an errata sheet against a
standard, one instrument's calibration procedure revised without its programming manual catching
up. Retrieve widely enough to see both documents in one pass, and never let an answer name a
value without saying which version of the document it came from.
## Design choices

### Why this level, and when to use another approach

Level 0 here is a keyword search over the same 13 documents, BM25 over the section text, with a
person reading the results. For the question below, that search puts the datasheet's Recommended
Operating Conditions section first and both of the change notice's relevant sections second and
third, of 87. The datasheet section a person reads first even states the problem outright:
"The 36.0 V maximum in this table is superseded for revision A and revision B boards. See
ecn-2608-04.md." A reader who reads that sentence, not just the number in the table above it, is
most of the way to the right answer with no model anywhere.

## When you do not need this

Try [order zero](/gradient_ascent/techniques/order-zero/) keyword search first if a
person is going to read the top few hits themselves. It already ranks the datasheet's
Recommended Operating Conditions section and the change notice inside the top three of 87
sections for the question this page walks through.

So the case for level 2 here is not that retrieval is hard: it plainly is not. It is that a
person chasing one footnote at a time does not scale to a shop asking dozens of these questions
a day, and a narrower lookup, "find the max input voltage in the datasheet", can return the
datasheet's number and never reach the footnote at all. [RAG](/gradient_ascent/techniques/rag/), retrieving across the whole corpus rather than one document,
and [structured output](/gradient_ascent/techniques/structured-output/), requiring the reply to
name which revision it is for and rejecting a citation list that drops a retrieved notice, do that
chasing in code every time instead of depending on whoever is reading that day.

Climbing to [agentic RAG](/gradient_ascent/techniques/agentic-rag/) (level 5), where the model
runs a second search depending on what the first one found, buys nothing extra here. That is worth
saying plainly, because documents disagreeing about a revision-dependent number is normally a
reason to climb. Three things make this corpus the exception. The two conflicting numbers use
enough of the same words, "maximum," "input voltage," the revision letters, that one search
returns both the datasheet section and the notice: 1st and 6th of the top 8 with the hashing
stand-in embedder this repo tests against, and 1st and 3rd of 87 with BM25 and no embedder at all.
The corpus is 87 sections, so a top-8 cut is a tenth of everything. And the notice names what it
supersedes, in its own header and its first paragraph, which is what lets code check a citation
list instead of hoping.

Change any of those three and the argument goes with it. A corpus of thousands of sections can
leave the notice out of a candidate set with nothing to say so, and a notice that does not name
what it supersedes gives the check below nothing to key on: finding it then means reading the
first answer and searching again on what it said, which is the dependent second search level 5
exists for. On this corpus none of that holds, so the extra calls buy nothing.

## Build it

### Implementation details and code

The whole pipeline, retrieval through validation:

`examples/bench_ask_the_datasheet/run.py` (lines 107-157)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder,
    tracer: Tracer,
    *,
    corpus_dir: Path = BENCH_CORPUS_DIR,
    top_k: int = TOP_K,
) -> Answer:
    sections = load_sections(corpus_dir)
    tracer.record(kind="code", decided_by="code", title="Chunk the bench documents", detail=f"{len(sections)} sections")
    sources = _retrieve(question, sections, embedder, top_k)
    retrieved = {s.cite for s in sources}
    tracer.record(
        kind="code",
        decided_by="code",
        title="Embed and retrieve top-k",
        detail=", ".join(s.cite for s in sources),
    )
    blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
    messages = [
        Message(role="system", content=SYSTEM_PROMPT),
        Message(role="user", content=f"Sources:\n\n{blocks}\n\nQuestion: {question}"),
    ]
    tracer.record(kind="code", decided_by="code", title="Build prompt with sources and schema", detail=f"{len(sources)} sources")
    record: dict = {}
    for attempt in range(MAX_RETRIES + 1):
        completion = model.complete(messages, schema=SCHEMA, max_tokens=400)
        tracer.record(
            kind="model",
            decided_by="code",
            title="Ask the model for a cited, revision-scoped answer" if attempt == 0 else "Ask again with the validation error",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        try:
            record = json.loads(completion.text)
            problems = _validate(record, retrieved)
        except json.JSONDecodeError as exc:
            record, problems = {}, [f"invalid JSON: {exc}"]
        tracer.record(kind="code", decided_by="code", title="Validate the reply", detail="; ".join(problems) or "valid")
        if not problems:
            text = f"{record['answer']} (applies to: {record['applies_to_revision']})"
            return Answer(text=text, citations=list(record["citations"]))
        if attempt < MAX_RETRIES:
            messages.append(
                Message(role="user", content=f"That did not validate: {'; '.join(problems)}. Reply again with corrected JSON only.")
            )
    return Answer(text=json.dumps({"error": "did not validate after retry", "last": record}), citations=[])
```

The check that makes the difference between this and plain RAG runs after every reply:

`examples/bench_ask_the_datasheet/run.py` (lines 78-104)

```python
def _validate(record: dict, retrieved: set[str]) -> list[str]:
    """What a reply has to have before code will hand it back as the answer.

    Two of these checks exist because of what this bench is for, not because of JSON schemas in
    general: a citation has to be one of the sources code actually retrieved (a model cannot cite
    a document it was never shown), and `applies_to_revision` has to be filled in, because a
    number that is silently missing its revision is exactly the failure this recipe exists to
    catch -- correct for one board revision and wrong for another, with nothing on the page to
    tell them apart.
    """
    problems = [f"missing field: {f}" for f in REQUIRED_FIELDS if f not in record]
    if problems:
        return problems
    if not isinstance(record["answer"], str) or not record["answer"].strip():
        problems.append("answer must be a non-empty string")
    if not isinstance(record["applies_to_revision"], str) or not record["applies_to_revision"].strip():
        problems.append("applies_to_revision must be a non-empty string")
    if not isinstance(record["citations"], list) or not record["citations"]:
        problems.append("citations must be a non-empty list")
    else:
        unknown = [c for c in record["citations"] if c not in retrieved]
        if unknown:
            problems.append(f"citation(s) not among the retrieved sources: {', '.join(unknown)}")
        for superseded, superseding in SUPERSEDED_BY.items():
            if superseded in record["citations"] and superseding in retrieved and superseding not in record["citations"]:
                problems.append(f"cites {superseded} without {superseding}, which supersedes it")
    return problems
```

Every step above is `decided_by: "code"`. Retrieval, the prompt, the schema and the retry count
are fixed before the model runs; the model composes the `answer` and names which of the eight
sections it used, and nothing it says chooses what code does next.



Last reviewed 2026-09-19.
