# Check a board against the design rules document

_Recipe · needs level 3_

A bill of materials and a netlist summary are checked rule by rule against the written design rules. One pass drafts findings, a second checks each finding against the rule text it cites and drops the ones that cite nothing. Level 3, because code decides every step and the rules do not change between boards. This is the rule check that happens before a review meeting, not the design review report itself: for the report, and the characterization data behind it, see the two recipes this page links in its first paragraph.


Two different jobs are called a design review, and this is the narrower one: checking a board
against a written rules document, rule by rule, before anyone meets about it. If what you need is
the report itself, the one built from measured data, the numbers come from [sweeping the design over its corners](/gradient_ascent/recipes/characterize-a-design/) and the writing
up is [turning a measurement session into a report](/gradient_ascent/recipes/measurement-writeup/).
Neither of those is this page.

This is design work rather than test work, and it happens before there is a board to power at all:
the review reads documents against documents, and nothing in it goes near an instrument. It sits
in front of all three of the settings this bench covers, because the production line, the
characterization sweep and the one careful measurement all run on a board some review like this
one released. Before Orbeck releases a board, someone checks it against DR-0100, the seven-rule
design review document: a bill of materials, a netlist summary, and every rule in
`design-review-rules.md` checked one at a time, met, not met or not applicable, with the numbers
that support each answer. DR-0100 is explicit about what counts: "A finding that does not name a
rule is a comment, not a finding, and does not hold a release." A reviewer who cannot tell from
the documents whether a rule is met records that too, and it counts as not met until the documents
say otherwise.

Nothing below decides whether the board ships, and no model here produces a measurement, an
uncertainty or a margin: where a rule is a number against a threshold, code computes it. A rule
marked not met still needs a person to fix the design or waive it in writing, by name, with a
reason; no output here marks anything waived. This checklist produces the record a person then
acts on.

## Walkthrough

No trace has been recorded for this example (see `docs/EVALS.md`), so this is an illustrated run
on `StubModel`: the numeric steps below are the actual arithmetic the code runs, and the two model
responses are scripted stub text chosen to demonstrate one catch. Two inputs are written into the
example rather than read from the bench corpus, because the corpus has neither: the netlist
summary's placement figures, and the evidence submitted for the two rules that need reading. The
bill of materials, DR-0100 and the change notice are real files in `evals/bench/corpus/`.

Checking SRB-5030 revision B. A request naming a revision this board does not have is refused
rather than reviewed against the wrong bill of materials, and one naming none gets revision B, the
revision in production:

1. Code loads DR-0100 and keys its seven rules by id: DR-10, DR-12, DR-14, DR-16, DR-20, DR-24,
   DR-30.
2. Code computes the five numeric rules against revision B's bill of materials and netlist
   summary. DR-14: met, 33.3 V supported against the 32.0 V ECN ceiling. DR-10: met, a margin of
   1.33 against the 1.3x the rule requires, the same narrow margin `docs/THE-BENCH.md` shows.
   DR-20: met, every decoupling cap within 3 mm of its pin. DR-12: not applicable, no discrete
   semiconductor on a DC rail is in this bill of materials. DR-16: not met, the bill of materials
   carries no power or ripple-current rating for the parts this rule covers, and DR-0100 rule 1
   says an unanswerable rule counts as not met.
3. A model drafts findings for DR-24 and DR-30 from the rule text and the evidence submitted for
   each. For DR-24 it drafts "met," reasoning that "thermal shutdown at 145 degC protects the
   design, so the junction temperature requirement is satisfied." That is the argument DR-24's own
   text disclaims.
4. A second, separate model call checks each drafted finding against that rule's full text. It
   rejects the DR-24 finding: "DR-24 says a design whose junction temperature reaches the shutdown
   threshold in any rated operating condition does not meet this rule, whatever the protection
   does; citing the shutdown as the reason it is met is the opposite of what the rule says." It
   confirms DR-30, whose evidence addresses every clause the rule names.
5. Code merges the two passes. DR-30 ships as drafted, "met," attributed to the model. DR-24 is
   recorded "not met," with the checker's reason attached, so a person redoes it against the
   evidence that would actually support "met," the 110.3 degC figure the submission computed but
   did not cite for this rule.

The final record covers all seven rules for revision B, five checked by code and two by the two
model passes, none of them a release decision.

## What it costs

Two model calls happen no matter how many rules DR-0100 has, because both passes read every
judgment rule at once rather than one rule at a time. It does not scale with rule count the way a
per-rule call would: an eighth rule that needs reading adds tokens to the same two calls, not a
third and fourth call.

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one review:** 2
- **Tokens in, counted on the stub:** ~923
- **Tokens out, counted on the stub:** ~212
- **Rules checked for free:** 5 of 7

The unit here is the review, not the board. A design review runs once per board revision, or once
per engineering change on one, not once per board built: if the SRB-5030 sees roughly one ECN a
quarter, that is 2 model calls and under 1,200 tokens a quarter. The five arithmetic rules cost
nothing per review, the same way a limit check costs nothing per unit on
[limits without a model](/gradient_ascent/recipes/limits-without-a-model/).

## How it fails, specifically

### A finding cites a rule that does not say what it claims

- **How to notice it:** The finding names a real rule number and sounds plausible, but the rule's own text argues the opposite, or something the rule never says at all.
- **How to test for it:** Read the full text of the cited rule against the finding's stated reasoning, not just its status. That is the second pass's whole job, and DR-24's shutdown reasoning above is a scripted example of exactly this failure.

### The wrong document governs the number

- **How to notice it:** A numeric rule is checked against a superseded figure (the datasheet's 36.0 V maximum input) instead of the one that actually governs the board revision in hand (the ECN's 32.0 V for revisions A and B), or against a ceiling from one revision and a part from another.
- **How to test for it:** Check the ceiling and the bill of materials the code picked against the revision it was asked about, the way test_revision_a_and_b_use_the_ecn_ceiling_revision_c_uses_the_datasheet does: revision B's 50 V part against 32.0 V and revision C's 63 V part against 36.0 V both read met, while the same 50 V rating checked against 36.0 V, the arithmetic in the notice itself, reads not met.

### Missing evidence reads as not applicable instead of not met

- **How to notice it:** A rule with a real subject on the board (R1 through R3 and C1 through C8 are real parts DR-16 covers) gets marked not applicable because the bill of materials happens not to carry the rating the rule needs, when DR-0100 rule 1 says an unanswerable rule counts as not met.
- **How to test for it:** Check that a rule whose subject exists but whose evidence is missing comes back not met, not not applicable; only a rule with no matching subject at all, like DR-12 on this board, should read not applicable.

### A drafted finding ships uncorrected

- **How to notice it:** The first pass's status and evidence reach the review record even after the second pass rejects the reasoning behind them, because the merge step trusted the draft instead of the verdict.
- **How to test for it:** Assert on the merged report, not the draft: a rejected finding's final status must be not met with the checker's reason attached, never the drafted met carried through.

## How to evaluate it

This recipe's task is not one the site's shared 60-question set measures (see `docs/EVALS.md`):
it never answers a question about a document set, it produces findings from documents it is
handed. A right answer here is a finding whose status and cited rule agree with what a person
checking DR-0100 by hand would write for the same bill of materials and netlist summary.

Build a labeled set before tuning anything: DR-0100's seven rules against a handful of bills of
materials and netlist summaries with known right answers, including one board that should fail
each rule and one drafted finding you know cites its rule wrong, the way DR-24's does here. A
review runs once a revision, so a few dozen rule and finding pairs is both enough to start and
about all a reader will have.

The confusion that matters most is a false "met": it is what a wrong pass ships. Watch it
separately from a false "not met," which only costs a person's time re-checking something that
was fine, and separately again from how often the second pass rejects the first, since a checker
that rejects everything or nothing is not checking.

## How to adapt it

The instrument-porting story other recipes on this bench carry does not apply here. What does port
is the split and the two-pass shape: read every rule once, decide by a fixed test in code whether
a rule is arithmetic or judgment, compute the arithmetic ones directly, and for the rest draft a
finding and check it against the rule's own full text before a person sees it. What does not port
is DR-0100's rule numbers and thresholds, the SRB-5030's part numbers and ratings, and every figure
in `docs/THE-BENCH.md`; a reader's own design review document and bill of materials are what a
real port checks a finding against. Which document governs a number, when a notice has changed one
and the datasheet still prints the old figure, is [ask the
datasheet](/gradient_ascent/recipes/ask-the-datasheet/)'s subject.

The same shape, work already done checked against rules already written down, fits a pull request
against a style and security guide, a contract against a negotiation playbook, or a test plan
against its requirements just as well as it fits a circuit board.

## Design choices

### Why this level, and when to use another approach

This is level 3. Two model calls happen every run, always in the same order, and code always
does the same thing with whatever comes back: draft, then check, then merge. The model never
picks the next action, so every step below is `decided_by: "code"`, the same as any other fixed
pipeline on this site.

Five of DR-0100's seven rules need no model, and code settles all five. DR-14 is the clean case:
"A ceramic capacitor on a DC rail shall be rated at least 1.5 times the maximum steady-state rail
voltage stated in the product's own datasheet." C1 and C2, the SRB-5030's input capacitors on
revisions A and B, are 50 V parts. 50 / 1.5 = 33.3 V, so a 50 V part supports a 33.3 V rail and
not a 36.0 V one: that arithmetic is exactly what ECN-2608-04 cites to justify lowering the
board's maximum input to 32.0 V for revisions A and B. Code does the division and compares it with
whichever ceiling governs the revision under review, using the part that revision actually
carries: 50 V parts against the ECN's 32.0 V on revisions A and B, and revision C's 63 V parts,
where 63 / 1.5 = 42 V, against the datasheet's 36.0 V. A model asked to do that division is only
an added way to get 33.3 wrong.

DR-10 (inductor saturation margin) and DR-20's placement thresholds are the same shape: a number
off the bill of materials or the netlist summary, a factor or a distance the rule names, one
comparison. The last two are not comparisons at all, and still need no model: DR-12 has no subject
on this board, since the bill of materials lists no discrete MOSFET or diode on a DC rail, and
DR-16's evidence is missing outright, which rule 1 records as not met. Code states both directly
rather than dressing them up as calculations.

Two rules cannot be settled that way. DR-24 sets a numeric limit (junction temperature at or below
125 degC) but also requires the calculation shown "term by term, not as a single number", and
disclaims one specific wrong argument by name: a design that reaches thermal shutdown in normal
use does not meet the rule, "whatever the protection does". Telling whether submitted evidence
argues from the computed number, rather than from the shutdown being there, needs reading. DR-30
names several things at once (test points sized for a probe, a switch-node test point marked for
scope use and excluded as a fixture contact, a silkscreen character that matches the assembly
number), and checking a netlist summary against all of them is a reading task too. Those two, and
only those two, go to a model, through
[structured output](/gradient_ascent/techniques/structured-output/) so each finding comes back
as a rule, a status and evidence, and [evaluator
optimizer](/gradient_ascent/techniques/evaluator-optimizer/) so a second, separate pass checks each drafted finding against the rule's own
full text before it reaches a person.

Climbing to level 4 would let a model decide which rule needs a second look instead of running
both passes on every rule; that buys nothing here, because DR-0100 does not change between boards
and every rule is checked every time regardless. [Review
and debate](/gradient_ascent/techniques/debate-review/) (level 6) would add a third, independent reader where a wrong finding is
expensive enough to want two model opinions to agree first. The DR-24 catch below is that kind of
case, and a reader whose false findings hold up a real release should read that page next.

## Build it

### Implementation details and code

`examples/bench_design_review_checklist/run.py` (lines 386-412)

```python
def run(board_revision: str, model: Model, tracer: Tracer) -> ReviewReport:
    revision = _revision_from(board_revision)
    rules = _rule_sections()
    tracer.record(kind="code", decided_by="code", title="Load DR-0100", detail=f"{len(rules)} numbered rules")

    numeric = _numeric_findings(revision)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Compute the numeric rules",
        detail=", ".join(f"{f.rule}: {f.status}" for f in numeric),
    )

    rule_texts = {rid: rules[rid] for rid in JUDGMENT_RULES if rid in rules}
    evidence = {"DR-24": DR24_EVIDENCE, "DR-30": DR30_EVIDENCE}
    drafts = _draft_findings(rule_texts, evidence, model, tracer)
    verdicts = _check_findings(drafts, rule_texts, model, tracer)
    judged = _merge_judgment_findings(drafts, verdicts)

    findings = tuple(numeric + judged)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Assemble the review record",
        detail=f"{len(findings)} findings for SRB-5030 revision {revision}",
    )
    return ReviewReport(board=f"SRB-5030 revision {revision}", findings=findings)
```

`_numeric_findings` (not shown) calls five small functions, one per rule; here is the one behind
the walkthrough's DR-14 line, doing exactly the division above:

`examples/bench_design_review_checklist/run.py` (lines 215-226)

```python
def _check_capacitor_derating(rating_v: float, max_rail_v: float, *, ref: str) -> Finding:
    """DR-14: a ceramic on a DC rail must be rated at least 1.5 times the maximum steady-state
    rail voltage. This is the same arithmetic ECN-2608-04 uses to justify lowering the SRB-5030's
    input ceiling: a 50 V part supports 50 / 1.5 = 33.3 V, not the datasheet's superseded 36.0 V.
    """
    supported_v = rating_v / DERATE_FACTOR
    met = supported_v >= max_rail_v
    evidence = (
        f"{ref} is rated {rating_v:.1f} V; at the {DERATE_FACTOR}x factor DR-14 requires that "
        f"supports up to {supported_v:.1f} V, against a {max_rail_v:.1f} V maximum rail."
    )
    return Finding(rule="DR-14", status="met" if met else "not met", evidence=evidence, checked_by="code")
```

And here is the merge that turns a rejected draft into a recorded finding rather than a corrected
one:

`examples/bench_design_review_checklist/run.py` (lines 363-383)

```python
def _merge_judgment_findings(drafts: list[dict], verdicts: dict[str, dict]) -> list[Finding]:
    """A confirmed finding ships as drafted. A rejected or unverified one does not ship as a
    finding at all: it is recorded as not met, per DR-0100 rule 1 (unclear counts as not met
    until the documents say otherwise), with the checker's own reason attached, so a person
    redoes it instead of a wrong 'met' quietly reaching the review record."""
    findings = []
    for d in drafts:
        verdict = verdicts.get(d.get("rule", ""))
        if verdict is not None and verdict.get("verdict") == "confirm":
            findings.append(Finding(rule=d["rule"], status=d["status"], evidence=d["evidence"], checked_by="model"))
            continue
        reason = verdict["reason"] if verdict is not None else "the second pass returned no verdict for this rule"
        findings.append(
            Finding(
                rule=d.get("rule", "?"),
                status="not met",
                evidence=f"Pass 2 rejected the drafted citation: {reason}",
                checked_by="model",
            )
        )
    return findings
```

Run it: `python -m examples.bench_design_review_checklist --model stub:scripted`. All seven rules
come back judged, five by code and two by the model, with DR-24's drafted "met" rejected by the
second pass in the words the walkthrough describes. The same command with `--model stub` returns
free text where the two model passes ask for JSON, so it prints the five numeric findings and
stops there, which is worth seeing on its own: the code half of this checklist needs no model.



Last reviewed 2026-09-19.
