Turn a requirements list into a test plan
A fixed chain: read the requirements, propose a test for each, build the traceability table, then check that every requirement has a test and every test names a requirement. A person approves before any of it is adopted. The order of the steps is known in advance, which is what keeps this at level 3.
SourcedNeeds level 3
Orbeck’s SRB-5030 datasheet states eight numbers a board has to meet: an input voltage range, an output voltage window, two regulation percentages, a ripple ceiling, a no-load current draw, a switching frequency band and a current limit. Two different plans get drawn from that one list, and this page builds the production one. A design verification plan takes each requirement to its corners, reports the margin there with the uncertainty on it, and is read once by a design review; a production test plan checks each requirement against a limit, fast, on every unit that goes down the line. Same requirements, different documents, and a step written for one is usually wrong for the other. Characterize a design is the verification side of the same list.
Each of those numbers needs a step in the production test spec that actually measures it, on the right instrument, against the right limit. Drop one and a board can ship with a real requirement nobody tests for. Copy a number a later engineering change notice has overridden and a good board can fail against a limit that is no longer the real one, or the test can quietly stop catching what it was written to catch. This recipe reads the datasheet’s requirements, drafts one test per requirement, and builds a table proving every requirement traces to a test and every test back to a requirement, before a person signs off on the plan.
The run on the bench, stepped
Illustrated, not measured: no recorded trace exists for this example yet (see
docs/EVALS.md), so the run below is run() called with a scripted stand-in for the model’s
replies, the same eight requirements and stand-in tests/test_example_bench_requirements_to_test_plan.py
checks against.
Which revision the plan is for comes out of the request itself: “B”, “rev B” and “a test plan for revision C boards” all work, a revision this board does not have is refused rather than guessed at, and a request naming none gets revision A, since the ECN’s 32.0 V ceiling is the stricter of the two ways to be wrong.
Reading the requirements for a revision B board pulls eight rows from srb5030-datasheet.md
sections 3 and 4, and the same notice moves two of them. REQ-VIN’s 36.0 V ceiling is superseded
to 32.0 V by ecn-2608-04#1. REQ-LINEREG’s sweep goes with it, by section 4 of the notice:
production may not apply 36.0 V to a revision A or B board, even during test, so the line
regulation those revisions are tested for is measured 9.0 V to 32.0 V. The limit itself, 0.30%,
does not change; only the conditions it is measured under do. A plan drafted from the datasheet’s
own sweep would tell production to do something the notice forbids, so the superseded sweep never
reaches the model at all.
The model is then asked about each requirement in turn, eight separate calls. Five proposals land clean: REQ-VOUT to the MDN-6100 at 4.900 to 5.100 V, REQ-LINEREG and REQ-LOADREG to the MDN-6100 at 0.30% and 0.80%, REQ-RIPPLE to the TRN-1102 (not the MDN-6100, whose AC volts function stops at 300 kHz, well under what a 500 kHz switcher’s ripple needs) at 50.0 mV, and REQ-IQNL to the MDN-4010 at 25.0 mA. Three do not:
| Requirement | Proposal | Coverage check |
|---|---|---|
| REQ-VIN | MDN-4010, upper 36.0 V | fails: 36.0 V is srb5030-datasheet#3’s figure; ecn-2608-04#1 supersedes it to 32.0 V |
| REQ-FSW | (declined) | fails: no test was proposed for this requirement |
| REQ-ILIM | TRN-2500, 3.70 to 5.00 A | fails: names an instrument this bench does not have (the load here is a TRN-2400) |
check_coverage returns three problems, one per row above, and run returns a Blocked result
rather than a checkpoint: there is nothing yet for a person to approve. Once REQ-FSW gets a
proposal, REQ-ILIM’s instrument is corrected to TRN-2400, and REQ-VIN’s upper limit is corrected
to 32.0 V, the same eight calls clear the check and run returns a PendingApproval holding all
eight rows. A person reads it and calls resume(pending, "approve", tracer); only then is the
plan adopted, as Answer.text, a plain comma-separated table with every source cited.
What it costs
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
The unit here is the plan, not the unit tested: eight calls buy one board revision’s plan, and a
ninth requirement adds one more call and nothing else. The 1,608 tokens in and 283 out above are
tracer.tokens_in_total() and tracer.tokens_out_total() from a scripted stub run of all eight
requirements for revision B, counted by count_tokens over the prompts the example builds and the
replies scripted into the test. That is an estimate of what a backend would bill, not a
measurement of one, and the test pins both numbers so this page cannot drift from the code.
Building the traceability table and checking it costs nothing in tokens: both are loops over eight
rows, the same level-0 arithmetic limits without a
model argues most of test automation already is.
Compare that against the mistake the check exists to catch. A stale 36.0 V limit on a revision B line either fails a good board on a rail the ECN says is fine at that voltage, or, worse, never exercises the fixture near the 32.0 V ceiling that is the board’s real one.
How it fails on a real bench, specifically
A requirement the model never answers
- How to notice it
- The traceability table has a row with no proposed test at all, easy to miss by eye in a table of dozens of rows.
- How to test for it
- check_coverage reports the requirement by id when its row has no proposal; test_coverage_catches_a_requirement_silently_dropped plants exactly that (REQ-FSW, which the eight-step production spec never tests either) and checks the report names it.
A test naming an instrument this bench does not have
- How to notice it
- The plan reads fine until someone tries to run a TRN-2500’s command set against a fixture that has a TRN-2400, and there is no TRN-2500 on the bench or in any manual.
- How to test for it
- check_coverage compares every proposed instrument against the four this bench actually has; test_coverage_catches_an_instrument_not_on_this_bench plants TRN-2500 for REQ-ILIM and checks the report names it.
A limit copied from a datasheet number an ECN superseded
- How to notice it
- The plan looks complete and the test even runs; it checks a rail against 36.0 V on a board a later notice has capped at 32.0 V for the revisions actually in the field.
- How to test for it
- _stale_limit flags a proposed upper limit that matches the datasheet’s own figure for a requirement a notice supersedes; one test plants the stale figure on a revision B board and checks it is caught, another plants the identical figure on a revision C board, where it is correct, and checks it is not.
A test that answers a different requirement than the one it was asked about
- How to notice it
- Two rows end up pointing at the same test and one requirement is left with none, the kind of copy and paste error a hand built spreadsheet makes too.
- How to test for it
- check_coverage compares the requirement id the proposal itself states against the requirement it was actually asked about; test_coverage_catches_a_test_that_names_the_wrong_requirement constructs exactly that mismatch and checks it is caught.
One thing the check does not catch: an instrument that exists on this bench but is the wrong one
for the measurement, the MDN-6100 proposed for REQ-RIPPLE instead of the TRN-1102.
check_coverage only checks that a named instrument is one of the four here, not that it suits
what it is measuring; catching that needs a second, requirement-specific rule (a ripple test that
does not name the scope fails, specifically), which this recipe leaves for a
reader to add once real proposals show whether the model needs it.
How to evaluate it
There is no right or wrong answer to grade here the way a question over a document set has one.
check_coverage’s verdict is already a pass or fail, and a test proves what it catches, so what
is worth measuring is how often the model’s proposals need it. A plan is drafted once a revision,
so the set is small by nature: collect the plans for a handful of board revisions, ten to twenty
requirements in total, and track two counts separately. False drops, a real requirement that gets
no usable proposal, are the expensive kind: a missing test is a missing test until somebody
notices. False instruments and stale limits check_coverage catches for free, so what to watch
there is whether the model converges on the same mistake (always sending ripple to the meter, say),
which is worth fixing in the prompt or the schema rather than re-catching in every plan. No result
file exists for this example yet (see docs/EVALS.md), so none of this is a score, only what to
start counting.
How to adapt it to your own bench
Nothing in this recipe sends a command. It reads two documents and writes a third, and an
instrument appears in it only as a name a proposed test has to match against the equipment list.
What ports is the check: your own equipment list in place of BENCH_INSTRUMENTS, and your own
requirement ids. When one of these plans becomes a script that actually drives an instrument, that
is drafting a script from the
manual’s job, and it carries the rule that matters there: the model drafts, code checks
every command against that instrument’s own manual, the script runs on a simulated instrument
first, and a person bench-checks it with the current limit set low before it touches hardware.
Nothing in this repo has been run against real hardware.
What does not port is everything specific to this board: the SRB-5030’s eight requirements, its
REQUIREMENTS tuple, BENCH_INSTRUMENTS’s four names, and the ECN that supersedes two of them.
Your own datasheet has its own list, your own test spec has its own gaps, and a notice that has
overridden one of your own numbers is the document to check a proposed limit against.
Ask the datasheet is the same conflict read from
the other end.
The shape carries further than electronics. Anywhere a written list of requirements needs a matching, provably complete set of checks before anyone signs off, the same four steps apply: a project brief turned into tasks with an owner each, an incident report turned into a runbook step per contributing cause, a compliance document checked clause by clause. What moves between domains is the content of the check; what stays fixed is reading the requirements, drafting one check per requirement, building the table, and verifying coverage in code before a person signs off.
Design choices
Why this level, and when to use another approach
Two of the four steps are level 0. Building the traceability table is a zip of two equal-length
lists. Checking it is four comparisons per row: does a proposal exist, does it name the right
requirement, is its instrument one of the four this bench has, does it carry a limit and a unit.
check_coverage never calls a model and never treats what a model wrote as anything but a string
or a number to compare. The one step that needs a model is turning a requirement’s free text
(“output stays within plus or minus two percent over line, load and temperature”) into a
structured proposal naming an instrument, a measurement and a limit with its unit, the same
paraphrase into a fixed shape structured
output’s own page walks through.
A model could also be handed the whole eight-row datasheet table at once and asked to draft the plan in one call. That would still be level 3: the same fixed steps run in the same order regardless of what comes back. What asking once per requirement buys, prompt chaining’s own shape, is that one mangled reply drops one row instead of the table, and the trace shows which call a bad proposal came from.
Nothing here needs level 4 or level 5. Level 4 would let the model decide whether to look something up before answering, but every run reads the same two documents for the same board. Level 5 would let one proposal’s content change what happens next, and it does not: a proposal that copies a superseded limit changes only what the coverage check reports about that row. The order is fixed by code before the first model call; the coverage check is code after the last one; the model only ever drafts what sits between them. A drafted limit is data for the coverage check and not an answer, compared in code against the requirement it claims to answer and against the notice that supersedes it. Nothing a model writes here becomes a measurement, an uncertainty, a margin or a verdict, and a person, not the check and not the model, decides whether the plan is adopted.
Build it
Implementation details and code
View code: run
def run(revision: str, model: Model, tracer: Tracer) -> Blocked | PendingApproval:
revision = _revision_from(revision)
requirements = _requirements_for_revision(revision)
tracer.record(
kind="code",
decided_by="code",
title="Read the requirements",
detail=f"{len(requirements)} requirements, board revision {revision.strip().upper()}",
)
proposals = [_propose_test(r, model, tracer) for r in requirements]
rows = _build_traceability(requirements, proposals, tracer)
coverage = check_coverage(rows)
tracer.record(
kind="code",
decided_by="code",
title="Check coverage",
detail="pass" if coverage.ok else f"fail: {len(coverage.problems)} problem(s)",
)
if not coverage.ok:
return Blocked(rows=rows, coverage=coverage)
tracer.record(
kind="code",
decided_by="code",
title="Hold for a person's approval",
detail="the plan cleared the coverage check; nothing is adopted before resume() records a decision",
)
return PendingApproval(rows=rows, coverage=coverage)check_coverage is the part that decides pass or fail:
View code: check coverage
def check_coverage(rows: list[TraceRow]) -> CoverageResult:
"""Pass or fail. Every requirement needs at least one proposed test, and every proposed test
has to name the requirement it answers, an instrument this bench actually has, and a limit
with a unit. This function never calls a model and never reads one's output as anything but
data to check: the decision is arithmetic and string comparison, the same as every pass/fail
decision this site makes."""
problems: list[CoverageProblem] = []
for row in rows:
requirement, proposal = row.requirement, row.proposal
if proposal is None:
problems.append(CoverageProblem(requirement.id, "no test was proposed for this requirement"))
continue
if proposal.requirement_id != requirement.id:
problems.append(
CoverageProblem(
requirement.id,
f"the proposed test names {proposal.requirement_id!r}, not this requirement",
)
)
if proposal.instrument not in BENCH_INSTRUMENTS:
problems.append(
CoverageProblem(
requirement.id,
f"names {proposal.instrument!r}, which is not one of the four instruments on this bench",
)
)
if proposal.lower is None and proposal.upper is None:
problems.append(CoverageProblem(requirement.id, "the proposed test carries no limit"))
if not proposal.unit:
problems.append(CoverageProblem(requirement.id, "the proposed test carries no unit"))
stale = _stale_limit(requirement, proposal)
if stale is not None:
problems.append(CoverageProblem(requirement.id, stale))
return CoverageResult(ok=not problems, problems=problems)Techniques this recipe uses
The highest level it needs is level 3.
Turn a goal or a set of requirements into a structured plan
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Requirements into a test plan with a traceability table
- Every datasheet parameter into the corners a design verification plan measures it at
- A project brief into tasks and owners
- An incident report into a runbook
- A learning goal into a syllabus
- A customer request into a statement of work
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page