Check measurements against limits, and chart what drifts
Use code to calculate limits, yield, process capability, and trends across lots and fixtures. The pass/fail decision stays deterministic; no model is involved.
SourcedNeeds level 0
A month of Orbeck SRB-5030 regulator boards has gone through final test: 200 units, eight
measurements each, on four fixtures across two shifts a day. First-pass yield came back at 92.0
percent. Someone wants to know why the other 8 percent failed, and whether it is one bad batch of
parts, one fixture that needs recalibrating, or the ordinary spread of a process that is in fact
fine. The artifacts on hand are evals/bench/data/production-run-2026-08.csv (one row per
measurement: serial, lot, fixture, shift, timestamp, value, unit, its own limits, and the verdict)
and evals/bench/corpus/srb5030-test-spec.md, which says what each of the eight steps measures
and where its limits came from. This is production test: many units, a fixed sequence, a verdict
per step. It is one of the three ways this bench gets used, and its counterpart is
sweeping five prototypes and reporting the
margins, which is the same arithmetic asked a different question.
Reading the log, stepped
Nothing here talks to an instrument. The four instruments that produced this log
(docs/THE-BENCH.md describes them) already wrote their readings to
production-run-2026-08.csv; this recipe starts after that, reading the file the way
evals.bench.PRODUCTION_CSV names it.
The first unit in the log, SRB5030-2608-0001, has a RIPPLE row of 22.7 mV against an upper limit
of 50.0 mV. within_limits(22.7, None, 50.0) compares once and passes. Further down, unit
SRB5030-2608-0008’s RIPPLE row reads 52.2 mV against the same 50.0 mV limit and fails by 2.2 mV.
Its operator note reads “ripple 52mv”, which repeats the number and adds nothing the row did not
already say.
Grouping every RIPPLE reading by lot finds the first story in the data:
| Lot | n | Mean (mV) | SD (mV) | Cpk |
|---|---|---|---|---|
| L2608A | 52 | 21.21 | 3.40 | 2.82 |
| L2608B | 51 | 44.41 | 3.76 | 0.50 |
| L2608C | 47 | 21.50 | 1.36 | 6.97 |
| L2608D | 48 | 21.67 | 3.54 | 2.67 |
Ripple has one limit, so Cpk here is the one-sided form against the upper limit,
(50.0 - mean) / (3 * sd). Lot L2608B’s ripple mean is more than double the other three lots’,
and its Cpk of 0.50 says the process is not capable of holding the 50 mV limit, against 2.67 or
higher everywhere else. Grouped
by lot, first-pass yield tells almost none of this: 96.2 percent, 90.4 percent, 89.6 percent and
91.7 percent for the four lots in order, and L2608B is not even the worst of the four. A mean that
doubles and a Cpk that drops from 2.82 to 0.50 is a different lot of parts; a yield that moves by a
few points is ordinary. That is the argument for charting the measurement rather than counting the
failures.
A Shewhart individuals chart over the same 198 RIPPLE readings, in the order the log recorded them, makes the same point without being told which lot is which. Its center line is the overall mean, 27.37 mV; its control limits come from the average difference between one reading and the next (the moving range), not from the 50 mV spec limit and not from the sample standard deviation: center 27.37 mV, upper control limit 42.11 mV, lower control limit 12.62 mV. Thirty-eight of the 198 readings fall outside those limits, and thirty-six of the thirty-eight are L2608B units: the chart finds the same lot the grouped table did, from run order alone, before anyone thinks to group by anything. The other two are the boards with no output at all, 0.0 mV and 0.4 mV, under the lower control limit for a reason that has nothing to do with ripple. A chart says a reading does not belong with the others, and never says why.
Grouping VOUT by fixture (two boards with no output at all set aside first; more on that below) finds a second, unrelated story. VOUT is bounded on both sides, 4.9500 V to 5.0500 V, so its Cpk is the two-sided form, the smaller of the two one-sided numbers:
| Fixture | n | Mean (V) | Cpk |
|---|---|---|---|
| FIX-01 | 49 | 4.9892 | 1.09 |
| FIX-02 | 48 | 4.9922 | 1.09 |
| FIX-03 | 49 | 4.9640 | 0.44 |
| FIX-04 | 50 | 4.9906 | 1.09 |
FIX-03’s mean sits about 27 mV below the other three, and nowhere else does it move: its LINE_REG, LOAD_REG and EFF_FL means, which are differences of two readings rather than one absolute reading, do not, because a fixed offset in the measurement path cancels in a difference. That is a stale calibration constant on one fixture, not a board problem, and grouping by fixture is what tells the two apart. Confirming it takes one reading of the same node through a path that does not go through that channel, which is where working one of those boards at the bench picks the story up and measures the offset itself at 30 mV.
Grouping either measurement by day or by shift finds nothing worth following up. RIPPLE’s four daily means span 1.55 mV and its two shift means span 2.01 mV, against the 23.2 mV the lot grouping moves the mean by; VOUT’s daily means span about 5 mV and its shift means about 3 mV, against the roughly 27 mV the fixture grouping moves the mean by. Being able to say a grouping found nothing is as much the point of this recipe as finding the lot and the fixture were.
What it costs
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
The 6 ms is one machine timing one pass over 198 rows, not a claim about any other hardware. The shape is what matters, not the digits: nothing here costs a model call per item, so a log of 200 rows and a log of 200,000 cost the same per row, and so does a sweep over five prototypes.
How it fails on a real bench, specifically
Trusting the log's own result column
- How to notice it
- Every row already carries a PASS or FAIL. It will usually agree with a limit check written against the current test spec; an outdated limit table, a boundary tested strictly instead of inclusively, or one hand-edited row would not, and nothing about reading the column instead of the value tells you which case you're in.
- How to test for it
- Recompute PASS/FAIL from value, lower_limit and upper_limit for every row and diff against the logged column. within_limits is that recomputation; tests/test_bench_data.py proves the production log itself was written the same way, so a real mismatch means something changed.
A control limit read as the spec limit
- How to notice it
- RIPPLE's control chart upper limit comes out near 42 mV, well under the 50 mV spec limit. Reported as one unlabeled number, a reading between the two either ships a board that is inside its process's normal spread and gets flagged, or the reverse.
- How to test for it
- Print the control limits and the spec limits on the same line, from two different sources in the code (control_chart's own arithmetic, and the lower_limit/upper_limit columns), and confirm by construction that they are never the same computation.
A dead board folded into a capability number
- How to notice it
- A board with no output at all reads near 0 V on a rail with a 5 V nominal. Left in an unfiltered Cpk or a mean, one or two of these can make a genuinely capable process read as barely capable, for a reason that has nothing to do with capability.
- How to test for it
- Compare a group's Cpk with and without excluding a reading whose magnitude is implausible for the measurement (VOUT near 0 V, not just low). A Cpk that swings by more than the group's own spread when one or two rows move says those rows never belonged in the population being charted.
A measurement taken a different way landing in the same column
- How to notice it
- The RIPPLE column assumes every reading came from the same scope setup, AC coupled with the 20 MHz bandwidth limit on. Leave that limit off on one fixture and every RIPPLE row from it roughly doubles for a reason that has nothing to do with the board, and grouping by fixture calls it a fixture problem, correctly, for the wrong reason.
- How to test for it
- When a group's mean moves, check what produced the number before charting it further: same instrument setting, same probe point, same settle delay. A grouped table cannot see the difference between a fixture that measures differently and a fixture that runs different boards.
A units mismatch that still parses
- How to notice it
- A value column in millivolts under a header that says volts, or the reverse, still parses as a number and still compares as a number. Nothing raises.
- How to test for it
- Check a measurement's computed mean against its datasheet typical (docs/THE-BENCH.md states one for every step) before trusting a Cpk built from it. A mean off by a factor of 1,000 from the typical is a units bug, not a process finding.
How to evaluate it
There is no free-text answer here for a person to grade, so there is no confusion matrix and no
sample of graded examples to collect: within_limits, cpk and control_chart are closed-form
arithmetic, and either they reproduce the numbers a hand calculation gets or they have a bug.
tests/test_example_bench_limits_without_a_model.py checks every number this page states against
two independent computations: one call through run(), and one written straight from the CSV
with nothing but csv and statistics, the same way tests/test_bench_data.py checks the bench
data itself. Before trusting this recipe against a real line’s log, keep one hand-checked example
for each limit shape it will see (upper only, lower only, both) and one for the case that is easy
to get backwards: a group with too few readings to have a standard deviation at all. scripts/ eval_run.py’s 60-question set grades a drafted answer against a document set; nothing here drafts
an answer, so that runner does not apply to this recipe, and no page on this site should point a
reader at it as if it did.
How to adapt it
Every number in the two tables above, the 50 mV ripple limit, the four lots and fixtures, the
fixture offset itself, is this bench’s own and specific to docs/THE-BENCH.md. A real line has its
own test specification, its own lots and fixtures (or panels, reels or work orders, whichever your
line actually calls them), and its own answer for which limits are one-sided and which are two.
Read the real limits from the real spec the same way load_readings reads them from the log’s own
lower_limit and upper_limit columns, rather than typing them in twice where they can drift out
of sync with each other.
The individuals chart’s constant, 1.128, is specific to a moving range taken between single consecutive readings. A line that already groups its own readings into small subgroups (five boards from the same reel, say) uses a different, larger subgroup and a different table constant for that subgroup size; the shape of the chart does not change, only which number divides the moving range.
This shape is not specific to electronics test. Flagging an invoice over an approval threshold, finding a scheduling conflict in a calendar, reordering stock once a count falls below a minimum, and checking a bill of materials against a supplier’s end-of-life list are the same job: the input is already structured, the right answer is fixed by a rule or a subtraction, and two people given the same numbers get the same answer every time. None of those needs a model either.
Design choices
Why this level, and when to use another approach
Every question in the paragraph above is a comparison, a mean, a standard deviation or a
GROUP BY. Did this unit pass: two comparisons against its own row’s limits. What fraction of
units passed: a count. Is the ripple step capable of holding its 50 mV limit: a mean, a standard
deviation and a subtraction, which is the whole of Cpk. Did the process drift at some point: a
mean and a standard deviation plotted against run order, which is the whole of a control chart.
Which lot, fixture, day or shift looks different from the rest: GROUP BY each of the four in
turn. None of it needs a model, and none of it may have one: a model never produces a reported
measurement, an uncertainty, a margin or a verdict. Here the reason is the arithmetic of a defect
rate, since a wrong pass ships a bad unit and a model right 99 times in 100 adds a one percent
defect rate to a line that measures its own in parts per million. On five prototypes the reason
is different and the rule is the same: a margin nobody can reproduce is what a design decision
gets made on.
A model earns a place once the job stops being arithmetic over numbers that are already there.
Sorting failing units and operator notes into known
causes reads free text a GROUP BY cannot parse.
Answering a question from a datasheet, a test spec and a change notice resolves two
documents that disagree. Asking questions of a
production log starts only once the fixed charts below have been built and a question comes
in that none of them answers. All three cost a model call somewhere this recipe costs nothing.
This page and its engineering-test
counterpart are the two every other engineering recipe here points back to for the part of
the week a GROUP BY already does, one over a production run and one over a sweep.
Build it
Implementation details and code
The whole pass or fail decision is two comparisons:
View code: within limits
def within_limits(value: float, lower: float | None, upper: float | None) -> bool:
"""The whole pass or fail decision: two comparisons, no judgment, nothing a model could get
right 99 times out of 100 and wrong the hundredth."""
if lower is not None and value < lower:
return False
if upper is not None and value > upper:
return False
return TrueCpk is a mean, a standard deviation and a subtraction, one-sided against whichever single limit a step has and the smaller of both one-sided numbers when a step has two, which is the standard two-sided form:
View code: cpk
def cpk(values: list[float], lower: float | None, upper: float | None) -> float | None:
"""The standard capability index. One-sided against whichever single limit exists
(`(USL - mean) / (3 * sigma)` or `(mean - LSL) / (3 * sigma)`); the smaller of the two
one-sided numbers when both limits exist, which is the standard two-sided Cpk. `None` when
there is no limit to be capable against, or fewer than two readings to take a spread from."""
if len(values) < 2 or (lower is None and upper is None):
return None
mean = statistics.fmean(values)
sd = statistics.stdev(values)
if sd == 0.0:
return None
candidates = []
if upper is not None:
candidates.append((upper - mean) / (3.0 * sd))
if lower is not None:
candidates.append((mean - lower) / (3.0 * sd))
return min(candidates)GROUP BY is one function, called once each for lot, fixture, day and shift:
View code: group stats
def group_stats(readings: list[Reading], key: Callable[[Reading], str]) -> list[GroupStats]:
"""`GROUP BY key`, then mean, standard deviation, Cpk and yield within each group."""
groups: dict[str, list[Reading]] = defaultdict(list)
for r in readings:
groups[key(r)].append(r)
out = []
for name in sorted(groups):
rows = groups[name]
values = [r.value for r in rows]
failed = sum(1 for r in rows if not r.passed)
out.append(
GroupStats(
key=name,
n=len(rows),
mean=statistics.fmean(values),
sd=statistics.stdev(values) if len(values) > 1 else None,
cpk=cpk(values, rows[0].lower, rows[0].upper),
yield_pct=100.0 * (len(rows) - failed) / len(rows),
)
)
return outThe control chart estimates its sigma from the average moving range between consecutive readings rather than from the sample standard deviation, which is the textbook individuals-chart construction and not an arbitrary alternative: a standard deviation taken over a run that already contains a shifted lot is inflated by that shift, and the moving range is not.
View code: control chart
def control_chart(readings: list[Reading]) -> ControlChart:
"""A Shewhart individuals chart: center line is the mean, sigma is estimated from the
average moving range between consecutive readings (`mRbar / 1.128`), not from the sample
standard deviation. A sample standard deviation over a run that includes a shifted subgroup
is inflated by the shift itself; the moving-range estimate is not, which is why it is the
textbook choice for an individuals chart and not an arbitrary alternative to `statistics.
stdev`. Points beyond the resulting control limits are flagged without being told which lot,
fixture, day or shift they belong to -- that grouping is `group_stats`' job, not this one's."""
values = [r.value for r in readings]
if len(values) < 2:
mean = values[0] if values else 0.0
return ControlChart(center=mean, sigma=0.0, ucl=mean, lcl=mean)
moving_ranges = [abs(values[i] - values[i - 1]) for i in range(1, len(values))]
mr_bar = statistics.fmean(moving_ranges)
sigma = mr_bar / D2_MOVING_RANGE_N2
mean = statistics.fmean(values)
ucl = mean + 3.0 * sigma
lcl = max(0.0, mean - 3.0 * sigma)
flagged = [r for r in readings if r.value > ucl or r.value < lcl]
return ControlChart(center=mean, sigma=sigma, ucl=ucl, lcl=lcl, flagged=flagged)Every step this code takes is decided_by: "code" (examples/common/trace.py is the one
definition of that field on this site), because there is no model anywhere in the call for one to
decide anything. Run it yourself:
View code: README.md
python -m examples.bench_limits_without_a_model --measurement RIPPLE
python -m examples.bench_limits_without_a_model --measurement VOUTTechniques this recipe uses
The highest level it needs is level 0.
Look something up, or work it out from numbers you already have
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Pass or fail a measurement against its limits, and compute yield and Cpk
- Work out the margin to a specification at every corner of a sweep
- Build an uncertainty budget and guardband a limit by it
- Flag invoices over an approval threshold
- Find scheduling conflicts in a calendar
- Reorder stock when a count falls below a minimum
- Convert units or currencies
- Roll a week of work up into the counts, dates and totals a status report quotes
- Check a bill of materials for end-of-life parts against a supplier list
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page