# Turn a requirements list into a test plan

_Recipe · needs level 3_

A fixed chain: read the requirements, propose a test for each, build the traceability table, then check that every requirement has a test and every test names a requirement. A person approves before any of it is adopted. The order of the steps is known in advance, which is what keeps this at level 3.


Orbeck's SRB-5030 datasheet states eight numbers a board has to meet: an input voltage range, an
output voltage window, two regulation percentages, a ripple ceiling, a no-load current draw, a
switching frequency band and a current limit. Two different plans get drawn from that one list,
and this page builds the production one. A design verification plan takes each requirement to its
corners, reports the margin there with the uncertainty on it, and is read once by a design review;
a production test plan checks each requirement against a limit, fast, on every unit that goes down
the line. Same requirements, different documents, and a step written for one is usually wrong for
the other. [Characterize a design](/gradient_ascent/recipes/characterize-a-design/) is the
verification side of the same list.

Each of those numbers needs a step in the production test spec that actually measures it, on the
right instrument, against the right limit. Drop one and a board can ship with a real requirement
nobody tests for. Copy a number a later engineering change notice has overridden and a good board
can fail against a limit that is no longer the real one, or the test can quietly stop catching
what it was written to catch. This recipe reads the datasheet's requirements, drafts one test per
requirement, and builds a table proving every requirement traces to a test and every test back to
a requirement, before a person signs off on the plan.

## The run on the bench, stepped

Illustrated, not measured: no recorded trace exists for this example yet (see
`docs/EVALS.md`), so the run below is `run()` called with a scripted stand-in for the model's
replies, the same eight requirements and stand-in `tests/test_example_bench_requirements_to_test_plan.py`
checks against.

Which revision the plan is for comes out of the request itself: "B", "rev B" and "a test plan for
revision C boards" all work, a revision this board does not have is refused rather than guessed
at, and a request naming none gets revision A, since the ECN's 32.0 V ceiling is the stricter of
the two ways to be wrong.

Reading the requirements for a revision B board pulls eight rows from `srb5030-datasheet.md`
sections 3 and 4, and the same notice moves two of them. REQ-VIN's 36.0 V ceiling is superseded
to 32.0 V by `ecn-2608-04#1`. REQ-LINEREG's sweep goes with it, by section 4 of the notice:
production may not apply 36.0 V to a revision A or B board, even during test, so the line
regulation those revisions are tested for is measured 9.0 V to 32.0 V. The limit itself, 0.30%,
does not change; only the conditions it is measured under do. A plan drafted from the datasheet's
own sweep would tell production to do something the notice forbids, so the superseded sweep never
reaches the model at all.

The model is then asked about each requirement in turn, eight separate calls. Five proposals land
clean: REQ-VOUT to the MDN-6100 at 4.900 to 5.100 V, REQ-LINEREG and REQ-LOADREG to the MDN-6100
at 0.30% and 0.80%, REQ-RIPPLE to the TRN-1102 (not the MDN-6100, whose AC volts function stops
at 300 kHz, well under what a 500 kHz switcher's ripple needs) at 50.0 mV, and REQ-IQNL to the
MDN-4010 at 25.0 mA. Three do not:

| Requirement | Proposal | Coverage check |
| --- | --- | --- |
| REQ-VIN | MDN-4010, upper 36.0 V | fails: 36.0 V is `srb5030-datasheet#3`'s figure; `ecn-2608-04#1` supersedes it to 32.0 V |
| REQ-FSW | (declined) | fails: no test was proposed for this requirement |
| REQ-ILIM | TRN-2500, 3.70 to 5.00 A | fails: names an instrument this bench does not have (the load here is a TRN-2400) |

`check_coverage` returns three problems, one per row above, and `run` returns a `Blocked` result
rather than a checkpoint: there is nothing yet for a person to approve. Once REQ-FSW gets a
proposal, REQ-ILIM's instrument is corrected to TRN-2400, and REQ-VIN's upper limit is corrected
to 32.0 V, the same eight calls clear the check and `run` returns a `PendingApproval` holding all
eight rows. A person reads it and calls `resume(pending, "approve", tracer)`; only then is the
plan adopted, as `Answer.text`, a plain comma-separated table with every source cited.

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls per revision:** 8
- **Tokens in:** 1,608
- **Tokens out:** 283
- **Coverage check cost:** $0, no model

The unit here is the plan, not the unit tested: eight calls buy one board revision's plan, and a
ninth requirement adds one more call and nothing else. The 1,608 tokens in and 283 out above are
`tracer.tokens_in_total()` and `tracer.tokens_out_total()` from a scripted stub run of all eight
requirements for revision B, counted by `count_tokens` over the prompts the example builds and the
replies scripted into the test. That is an estimate of what a backend would bill, not a
measurement of one, and the test pins both numbers so this page cannot drift from the code.
Building the traceability table and checking it costs nothing in tokens: both are loops over eight
rows, the same level-0 arithmetic [limits without a
model](/gradient_ascent/recipes/limits-without-a-model/) argues most of test automation already is.

Compare that against the mistake the check exists to catch. A stale 36.0 V limit on a revision B
line either fails a good board on a rail the ECN says is fine at that voltage, or, worse, never
exercises the fixture near the 32.0 V ceiling that is the board's real one.

## How it fails on a real bench, specifically

### A requirement the model never answers

- **How to notice it:** The traceability table has a row with no proposed test at all, easy to miss by eye in a table of dozens of rows.
- **How to test for it:** check_coverage reports the requirement by id when its row has no proposal; test_coverage_catches_a_requirement_silently_dropped plants exactly that (REQ-FSW, which the eight-step production spec never tests either) and checks the report names it.

### A test naming an instrument this bench does not have

- **How to notice it:** The plan reads fine until someone tries to run a TRN-2500’s command set against a fixture that has a TRN-2400, and there is no TRN-2500 on the bench or in any manual.
- **How to test for it:** check_coverage compares every proposed instrument against the four this bench actually has; test_coverage_catches_an_instrument_not_on_this_bench plants TRN-2500 for REQ-ILIM and checks the report names it.

### A limit copied from a datasheet number an ECN superseded

- **How to notice it:** The plan looks complete and the test even runs; it checks a rail against 36.0 V on a board a later notice has capped at 32.0 V for the revisions actually in the field.
- **How to test for it:** _stale_limit flags a proposed upper limit that matches the datasheet’s own figure for a requirement a notice supersedes; one test plants the stale figure on a revision B board and checks it is caught, another plants the identical figure on a revision C board, where it is correct, and checks it is not.

### A test that answers a different requirement than the one it was asked about

- **How to notice it:** Two rows end up pointing at the same test and one requirement is left with none, the kind of copy and paste error a hand built spreadsheet makes too.
- **How to test for it:** check_coverage compares the requirement id the proposal itself states against the requirement it was actually asked about; test_coverage_catches_a_test_that_names_the_wrong_requirement constructs exactly that mismatch and checks it is caught.

One thing the check does not catch: an instrument that exists on this bench but is the wrong one
for the measurement, the MDN-6100 proposed for REQ-RIPPLE instead of the TRN-1102.
`check_coverage` only checks that a named instrument is one of the four here, not that it suits
what it is measuring; catching that needs a second, requirement-specific rule (a ripple test that
does not name the scope fails, specifically), which this recipe leaves for a
reader to add once real proposals show whether the model needs it.

## How to evaluate it

There is no right or wrong answer to grade here the way a question over a document set has one.
`check_coverage`'s verdict is already a pass or fail, and a test proves what it catches, so what
is worth measuring is how often the model's proposals need it. A plan is drafted once a revision,
so the set is small by nature: collect the plans for a handful of board revisions, ten to twenty
requirements in total, and track two counts separately. False drops, a real requirement that gets
no usable proposal, are the expensive kind: a missing test is a missing test until somebody
notices. False instruments and stale limits `check_coverage` catches for free, so what to watch
there is whether the model converges on the same mistake (always sending ripple to the meter, say),
which is worth fixing in the prompt or the schema rather than re-catching in every plan. No result
file exists for this example yet (see `docs/EVALS.md`), so none of this is a score, only what to
start counting.

## How to adapt it to your own bench

Nothing in this recipe sends a command. It reads two documents and writes a third, and an
instrument appears in it only as a name a proposed test has to match against the equipment list.
What ports is the check: your own equipment list in place of `BENCH_INSTRUMENTS`, and your own
requirement ids. When one of these plans becomes a script that actually drives an instrument, that
is [drafting a script from the
manual](/gradient_ascent/recipes/instrument-script-from-the-manual/)'s job, and it carries the rule that matters there: the model drafts, code checks
every command against that instrument's own manual, the script runs on a simulated instrument
first, and a person bench-checks it with the current limit set low before it touches hardware.
Nothing in this repo has been run against real hardware.

What does not port is everything specific to this board: the SRB-5030's eight requirements, its
`REQUIREMENTS` tuple, `BENCH_INSTRUMENTS`'s four names, and the ECN that supersedes two of them.
Your own datasheet has its own list, your own test spec has its own gaps, and a notice that has
overridden one of your own numbers is the document to check a proposed limit against.
[Ask the datasheet](/gradient_ascent/recipes/ask-the-datasheet/) is the same conflict read from
the other end.

The shape carries further than electronics. Anywhere a written list of requirements needs a
matching, provably complete set of checks before anyone signs off, the same four steps apply: a
project brief turned into tasks with an owner each, an incident report turned into a runbook step
per contributing cause, a compliance document checked clause by clause. What moves between domains
is the content of the check; what stays fixed is reading the requirements, drafting one check per
requirement, building the table, and verifying coverage in code before a person signs off.

## Design choices

### Why this level, and when to use another approach

Two of the four steps are level 0. Building the traceability table is a `zip` of two equal-length
lists. Checking it is four comparisons per row: does a proposal exist, does it name the right
requirement, is its instrument one of the four this bench has, does it carry a limit and a unit.
`check_coverage` never calls a model and never treats what a model wrote as anything but a string
or a number to compare. The one step that needs a model is turning a requirement's free text
("output stays within plus or minus two percent over line, load and temperature") into a
structured proposal naming an instrument, a measurement and a limit with its unit, the same
paraphrase into a fixed shape [structured
output](/gradient_ascent/techniques/structured-output/)'s own page walks through.

A model could also be handed the whole eight-row datasheet table at once and asked to draft the
plan in one call. That would still be level 3: the same fixed steps run in the same order
regardless of what comes back. What asking once per requirement buys,
[prompt chaining](/gradient_ascent/techniques/prompt-chaining/)'s own shape, is that one mangled
reply drops one row instead of the table, and the trace shows which call a bad proposal came from.

Nothing here needs level 4 or level 5. Level 4 would let the model decide whether to look something
up before answering, but every run reads the same two documents for the same board. Level 5 would
let one proposal's content change what happens next, and it does not: a proposal that copies a
superseded limit changes only what the coverage check reports about that row. The order is
fixed by code before the first model call; the coverage check is code after the last one; the
model only ever drafts what sits between them. A drafted limit is data for the coverage check and
not an answer, compared in code against the requirement it claims to answer and against the notice
that supersedes it. Nothing a model writes here becomes a measurement, an uncertainty, a margin or
a verdict, and a person, not the check and not the model, decides whether the plan is adopted.

## Build it

### Implementation details and code

`examples/bench_requirements_to_test_plan/run.py` (lines 377-403)

```python
def run(revision: str, model: Model, tracer: Tracer) -> Blocked | PendingApproval:
    revision = _revision_from(revision)
    requirements = _requirements_for_revision(revision)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Read the requirements",
        detail=f"{len(requirements)} requirements, board revision {revision.strip().upper()}",
    )
    proposals = [_propose_test(r, model, tracer) for r in requirements]
    rows = _build_traceability(requirements, proposals, tracer)
    coverage = check_coverage(rows)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Check coverage",
        detail="pass" if coverage.ok else f"fail: {len(coverage.problems)} problem(s)",
    )
    if not coverage.ok:
        return Blocked(rows=rows, coverage=coverage)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Hold for a person's approval",
        detail="the plan cleared the coverage check; nothing is adopted before resume() records a decision",
    )
    return PendingApproval(rows=rows, coverage=coverage)
```

`check_coverage` is the part that decides pass or fail:

`examples/bench_requirements_to_test_plan/run.py` (lines 297-330)

```python
def check_coverage(rows: list[TraceRow]) -> CoverageResult:
    """Pass or fail. Every requirement needs at least one proposed test, and every proposed test
    has to name the requirement it answers, an instrument this bench actually has, and a limit
    with a unit. This function never calls a model and never reads one's output as anything but
    data to check: the decision is arithmetic and string comparison, the same as every pass/fail
    decision this site makes."""
    problems: list[CoverageProblem] = []
    for row in rows:
        requirement, proposal = row.requirement, row.proposal
        if proposal is None:
            problems.append(CoverageProblem(requirement.id, "no test was proposed for this requirement"))
            continue
        if proposal.requirement_id != requirement.id:
            problems.append(
                CoverageProblem(
                    requirement.id,
                    f"the proposed test names {proposal.requirement_id!r}, not this requirement",
                )
            )
        if proposal.instrument not in BENCH_INSTRUMENTS:
            problems.append(
                CoverageProblem(
                    requirement.id,
                    f"names {proposal.instrument!r}, which is not one of the four instruments on this bench",
                )
            )
        if proposal.lower is None and proposal.upper is None:
            problems.append(CoverageProblem(requirement.id, "the proposed test carries no limit"))
        if not proposal.unit:
            problems.append(CoverageProblem(requirement.id, "the proposed test carries no unit"))
        stale = _stale_limit(requirement, proposal)
        if stale is not None:
            problems.append(CoverageProblem(requirement.id, stale))
    return CoverageResult(ok=not problems, problems=problems)
```



Last reviewed 2026-09-19.
