# Sweep a design over its corners and report the margins

_Recipe · needs level 0_

Sweep prototype boards across line, load, and temperature. Code calculates margins, uncertainty, and guardbanded verdicts; no model is needed.


Five hand-built Orbeck SRB-5030 revision C boards came back from the board house. The design
review is Thursday, and nobody knows yet what the regulator does at 9 V in, 3 A out and 70 degC,
the corner where low line, full load and heat all push the output the same way.
This is engineering test, not production: five boards, not two hundred, and a report at the end
rather than a per-unit pass or fail. The artifacts on hand are
`evals/bench/data/characterization-2026-09.csv` (four input voltages, three load currents, three
ambients, five readings a point) and `evals/bench/corpus/characterization-notebook.md`, the
engineer's working notes.

## The sweep, stepped

Nothing here talks to an instrument: the MDN-6100 already wrote every reading to the file. A row
is a reading, not a verdict, since the limit a margin is taken against lives in the datasheet:
the 4.900 to 5.100 V window `srb5030-datasheet.md` section 4 holds over the full line, load and
temperature range, not the tighter window production checks use at one fixed condition.

Grouping every corner by board and keeping the smallest margin finds the same corner for all five
boards:

| Serial | VOUT at 9.0 V, 3.000 A, 70 degC | Margin to 4.900 V |
| --- | --- | --- |
| SRB5030-2609-0001 | 4.97391 V | 73.9 mV |
| SRB5030-2609-0002 | 4.95822 V | 58.2 mV |
| SRB5030-2609-0003 | 4.92038 V | 20.4 mV |
| SRB5030-2609-0004 | 4.97849 V | 78.5 mV |
| SRB5030-2609-0005 | 4.96470 V | 64.7 mV |

Board SRB5030-2609-0003 is not a failing board. It passes every corner in the file, and its line
regulation is 0.236 percent against a 0.300 percent limit: no single parameter is out of
specification. What it lacks is margin left over once low line, full load and heat all push the
same way. A sample of one typical board at 25 degC would have said nothing about it.

That margin holds against the measurement. A reading from this session prices out as follows,
from the meter's own accuracy specification (`mdn6100-programming-manual.md` sections 2 and 7):

`examples/bench_characterize_a_design/run.py` (lines 216-233)

```python
def budget_for_point(readings: list[Reading], *, range_v: float | None = None) -> UncertaintyBudget:
    """The uncertainty budget behind one output-voltage reading, at the range it was actually
    read on: meter accuracy, the display's resolution, the repeatability of these five readings,
    and the lead and connection contribution `characterization-notebook.md` section 6 carries over
    from `calibration-procedure.md`. `dc_voltage_budget` is the one function every engineering page
    on this bench imports for this; nothing here writes a second root sum of squares.

    `range_v` defaults to the range these readings carry in the file rather than to the range the
    sweep script meant to use. Pass it only to price the same readings on a range they were not
    taken on, which is what `range_cost` does deliberately."""
    values = [r.vout_v for r in readings]
    contributions = dc_voltage_budget(
        values, range_v=point_range_v(readings) if range_v is None else range_v
    )
    combined_v = combined_uncertainty(contributions)
    return UncertaintyBudget(
        contributions=contributions, combined_v=combined_v, expanded_v=expanded_uncertainty(combined_v)
    )
```

For board 3's five readings at its worst corner, 9.0 V in, 3.000 A out, 70 degC ambient, on the
10 V range the sweep script selected:

| Contribution | Half-width or spread | Standard uncertainty |
| --- | --- | --- |
| Meter accuracy, 1 year, 10 V range, 23 degC | 222.2 uV | 128.3 uV |
| Resolution, 10 uV per count | 5.0 uV | 2.9 uV |
| Repeatability, 5 readings | s = 80.7 uV | 36.1 uV |
| Leads and connections | 200 uV | 115.5 uV |
| Combined | | 176.4 uV |
| Expanded, k = 2 | | 352.7 uV |

The first column is what the specification or the record allows; the second reduces that to a
standard uncertainty. An accuracy specification is a limit and says nothing about where inside it
the meter sits, so it is treated as a rectangular distribution and divided by the square root of
3, as is the display's own quantization, whose half-width is half a count. Repeatability is the
one line measured on the day, not looked up: the sample standard deviation of the five readings
over the square root of five. The four combine by root sum of squares, and k = 2 covers roughly
95 percent of where the true value could be, a coverage statement rather than a guarantee that
the value is inside.

The 23 degC in the first line is the meter's own ambient, not the chamber's: the boards sat at
70 degC while the meter sat next to the operator (`characterization-notebook.md` section 6 names
both). The 352.7 uV belongs to those five readings and to no others: the meter's manual works a
budget for ten readings of a different rail and gets 349.5 uV, and the notebook works one for a
generic reading of this session and gets 351.6 uV. All three are honest, and none is
interchangeable with another.

Against 20.4 mV of margin, 352.7 uV is nothing: the margin runs about 58 times the measurement,
a design finding and not a coincidence of where the meter sat. `guarded_verdict` calls it a
pass, and a thin margin is a design question for Thursday's review, not a test failure.

Running the same check over every board's line regulation at every ambient finds a second board
with nothing wrong at its nominal corner:

| Serial | Ambient | Line regulation | Uncertainty | Verdict |
| --- | --- | --- | --- | --- |
| SRB5030-2609-0005 | 0 degC | 0.267% | +/-0.0075 points | pass |
| SRB5030-2609-0005 | 25 degC | 0.299% | +/-0.0075 points | cannot say |
| SRB5030-2609-0005 | 70 degC | 0.334% | +/-0.0076 points | fail |

A bare comparison against the 0.300 percent limit says 0.299 is a pass, and that is not wrong
about the arithmetic; it is wrong about what the measurement can support. Line regulation is a
difference of two readings through the same leads, so the lead-and-connection line cancels and
the uncertainty is smaller than a single reading's: 0.0075 percentage points, expanded at k = 2,
against a margin of 0.0006 points. Guardbanded, the acceptance limit is 0.2925 percent, and
0.299 is over it: the honest statement is that this measurement does not say whether that board
meets its line regulation limit at 25 degC. It is not a pass and it is not a fail. The 70 degC
row is not ambiguous: 0.334 against 0.300, far more than the measurement is worth, so board 5
raises a real design question whatever the 25 degC row can say.

The other twelve checks are passes, and not all the same kind, which is why the scan reports the
distance and not only the word: boards 1, 2 and 4 clear by about 25 times the uncertainty, board
3 by nine, board 5's 0 degC row by four. Two of them, both at 25 degC, carry four times the usual
uncertainty for a reason that is not the boards.

Grouping the output voltage's spread by the `meter_range_v` column every row already carries,
rather than by board or corner, finds sixty readings that scatter about seven times more than the
rest:

| Meter range | Points | Mean spread | Boards |
| --- | --- | --- | --- |
| 10.0 V | 168 | 58.6 uV | 5 |
| 100.0 V | 12 | 413.9 uV | 2 |

Nothing about those sixty readings is wrong: the numbers are right, just worth less, since the
meter's accuracy specification is per range and the 100 V range's term does not scale down for a
5 V reading. Pricing the same five readings both ways puts a number on it: 2.4 times the expanded
uncertainty. The window crosses a board boundary, touching two boards only partly, which rules
the boards out: a repeatability problem would show at every corner.

That window is also why two regulation checks carry 0.031 points instead of 0.0075: one end of
each was read on the 100 V range, so `scan_line_regulation` prices the pair on the coarser of the
two. A difference is no better than the weaker half of it.

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls:** 0
- **Tokens in / out:** 0 / 0
- **Compute, whole sweep:** ~10 ms
- **Cost per run:** $0

**Compared with measurement-writeup.** measurement-writeup spends one model call turning this sweep into a report, checked afterward against these numbers; this page's arithmetic never reaches that call.

This is an engineering-test cost, counted per run: nothing here is a per-item call, so the cost
is the same whether five boards came back or fifty.

## How it fails on a real bench, specifically

### A corner never visited

- **How to notice it:** This margin table is only as good as the corners it visits. A review that checks only 24 V in, 1 A out and 25 degC never sees board 3's thinnest margin, since that corner sits outside nominal testing.
- **How to test for it:** Check that the sweep's voltages, currents and ambients include the datasheet's minimum and maximum on each axis, not only its typical point: worst_corner_per_board reports only the worst of the corners it was given.

### A condition recorded that was not the condition taken

- **How to notice it:** One block in this file is labeled 12.0 V but was taken at 24.0 V. The output voltage gives nothing away, since holding it steady is the part's entire job. The input current does: at the labeled voltage the board would be putting out more power than it took in.
- **How to test for it:** tests/test_bench_characterization.py proves it: a power balance finds the one point where input power is below output power. No uncertainty budget would catch it, since every reading there is a good reading of a condition nobody asked for.

### A block that scatters for a reason that is not the board

- **How to notice it:** Sixty readings here scatter for a reason that has nothing to do with any board. Blamed on the nearest board, it reads as a repeatability problem that is not there.
- **How to test for it:** spread_by_meter_range groups spread by meter_range_v first, and range_cost prices the same readings both ways; budget_for_point reads the range off the readings rather than being told it.

### A margin declared a pass when it is smaller than the measurement

- **How to notice it:** A bare comparison against the limit reads 0.299 as a pass. The measurement does not support that: the expanded uncertainty is 0.0075 points, twelve times the margin. Calling it a pass ships a design decision the data does not back; calling it a fail is just as wrong.
- **How to test for it:** guarded_verdict, called with the figure's own expanded uncertainty instead of 0.0, returns cannot say here; scan_line_regulation runs it for every board at every ambient, so a case like this is found, not assumed away.

### A temperature coefficient applied to the wrong ambient

- **How to notice it:** The meter's accuracy specification has its own calibration band, centered on the meter, not the board. This bench's meter sits at about 23 degC while the boards sit in a chamber at 0, 25 or 70 degC; pricing it at the chamber's ambient moves the budget for no reason connected to the board.
- **How to test for it:** budget_for_point calls dc_voltage_budget with the meter's own ambient, not the sweep's tamb_c: a budget that moved with tamb_c while the meter's temperature did not would be the bug this guards against.

## How to evaluate it

There is no free-text answer here for a person to grade, so there is no confusion matrix: a
margin, a budget and a guardbanded verdict are closed-form arithmetic, and either they reproduce
the numbers `docs/THE-BENCH.md` and the notebook state or there is a bug.
`tests/test_example_bench_characterize_a_design.py` checks every number this page states two
ways: once by calling the module, once by recomputing it independently from the raw CSV. What a
person checks by hand is what those tests check: that the worst corner really is the minimum over
every corner in the file, that a cannot say is not a pass with a softer name, and that every
budget was priced on the range and calibration interval its readings carry.
`scripts/eval_run.py`'s 60-question set grades a drafted answer; nothing here drafts one, so it
says as much rather than pretending to score it.

## How to adapt it

Every number above is this bench's own. A real sweep has its own window and corners: the same
code takes them from the datasheet and the file, not from a list typed into the page.

The lead-and-connection contribution is a number from `calibration-procedure.md`, carried over
from a fixture path this bench harness is not: `characterization-notebook.md` section 7 says the
line is assumed, not measured.

This shape is not specific to electronics. Any job with a small number of prototypes swept over
several conditions, where the answer is a margin rather than a verdict and there are too few
units for a population statistic to mean anything, follows the same reasoning: a material's
strength across a batch of coupons at several temperatures, a mechanical tolerance stack checked
across a small first-article build. None of those needs a model either.

## Design choices

### Why this level, and when to use another approach

A sweep over line, load and temperature is a nested loop. The margin to a datasheet limit at a
corner is a subtraction, and an uncertainty budget for one reading is a root sum of squares over
four named numbers. Finding the corner where a board's margin is thinnest, of the 36 this sweep
visits for each of the five boards, is a minimum. None of it needs a model, for the same reason
[limits without a model](/gradient_ascent/recipes/limits-without-a-model/) needs none: a margin
nobody can reproduce is worse than no margin, and an uncertainty a model assembled looks exactly
like one a budget produced. This page has no model call in it.

Low volume changes what the arithmetic is for, not whether it is arithmetic. There is no golden
run and no population to run statistics over, so no Cpk and no control chart; a margin and an
uncertainty budget do that work instead, one board at a time.

A model earns a place once the job stops being arithmetic over numbers this sweep already
produced: turning the notebook and these results into a report is
[measurement writeup](/gradient_ascent/recipes/measurement-writeup/), and getting the meter's
accuracy table out of its programming manual is
[accuracy specs from the manual](/gradient_ascent/recipes/accuracy-specs-from-the-manual/).

## Build it

### Implementation details and code

`examples/bench_characterize_a_design/run.py` (lines 194-206)

```python
def worst_corner_per_board(points: Points) -> list[CornerMargin]:
    """One `CornerMargin` per board: the corner, of every one this sweep visited, where the
    margin to the datasheet's window is smallest. `points` already holds every corner; this keeps
    the minimum per serial and says nothing about which corner that will turn out to be.
    Sorted worst first, so `corners[0]` is the board this sweep should worry about."""
    best: dict[str, CornerMargin] = {}
    for (serial, tamb_c, vin_v, iout_a), readings in points.items():
        mean_v = statistics.fmean(r.vout_v for r in readings)
        margin_v, side = window_margin(mean_v)
        candidate = CornerMargin(serial, tamb_c, vin_v, iout_a, mean_v, margin_v, side)
        if serial not in best or candidate.margin_v < best[serial].margin_v:
            best[serial] = candidate
    return sorted(best.values(), key=lambda c: c.margin_v)
```

`examples/bench_characterize_a_design/run.py` (lines 267-282)

```python
def scan_line_regulation(points: Points) -> list[RegulationCheck]:
    """Line regulation, guardbanded, for every board at every ambient this sweep visited, not only
    the one board and one ambient `docs/THE-BENCH.md` happens to name. A verdict that is not a
    clean pass is a finding this scan makes, not one it was handed."""
    ambients = sorted({tamb for (_, tamb, _, _) in points})
    serials = sorted({serial for (serial, _, _, _) in points})
    checks = []
    for serial in serials:
        for tamb_c in ambients:
            value_pct = line_regulation_pct(points, serial, tamb_c)
            uncertainty_pct = regulation_uncertainty_pct(
                points[(serial, tamb_c, 32.0, 1.000)], points[(serial, tamb_c, 9.0, 1.000)]
            )
            verdict = guarded_verdict(value_pct, uncertainty_pct, upper=LINE_REG_MAX_PCT)
            checks.append(RegulationCheck(serial, tamb_c, value_pct, uncertainty_pct, verdict))
    return checks
```

`examples/bench_characterize_a_design/run.py` (lines 285-311)

```python
def spread_by_meter_range(points: Points) -> list[RangeSpread]:
    """The output voltage's point-to-point spread, grouped by the meter range each point was
    actually read on, not by board and not by corner. Every point in this file was read entirely
    on one range (the sweep script picks a range once per board; characterization-notebook.md
    section 3 records the one exception), so this is one `GROUP BY` on a column the file already
    carries. Sorted by how many points used that range, most first."""
    spreads_by_range: dict[float, list[float]] = defaultdict(list)
    boards_by_range: dict[float, set[str]] = defaultdict(set)
    for (serial, _tamb_c, _vin_v, _iout_a), readings in points.items():
        ranges_here = {r.meter_range_v for r in readings}
        if len(ranges_here) != 1:
            raise ValueError("a point spans more than one meter range")
        range_v = ranges_here.pop()
        spreads_by_range[range_v].append(statistics.stdev(r.vout_v for r in readings))
        boards_by_range[range_v].add(serial)
    return sorted(
        (
            RangeSpread(
                range_v=range_v,
                n_points=len(spreads),
                mean_stdev_v=statistics.fmean(spreads),
                boards=frozenset(boards_by_range[range_v]),
            )
            for range_v, spreads in spreads_by_range.items()
        ),
        key=lambda s: -s.n_points,
    )
```

Every step here is `decided_by: "code"`, and no budget is assembled by hand:
`dc_voltage_budget`, `combined_uncertainty` and `expanded_uncertainty` come from
`examples/common/bench.py`, checked against the meter's manual in `tests/test_bench.py`. Run it
yourself:

`examples/bench_characterize_a_design/README.md` (lines 13-14)

```text
python -m examples.bench_characterize_a_design --serial SRB5030-2609-0003
python -m examples.bench_characterize_a_design --serial SRB5030-2609-0005
```



Last reviewed 2026-09-19.
