Sort failing units and operator notes into causes
Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.
SourcedNeeds level 3
This is a production test page, and production is the setting it is for. Every day’s failing units
land in a queue with two things attached to each one: which limit it missed, and whatever the
operator happened to type while the unit was on the bench. Orbeck’s own failure analysis guide,
evals/bench/corpus/failure-analysis-guide.md, already lists the causes and a routing table by
step, so this is not an investigation that starts from nothing. It is sorting failures into causes
the guide already names, and sending each one where the guide already says it goes.
The low-volume counterpart of this job is a person reading the notes. An engineer with five prototypes on the bench wrote the notes themselves and can read all of them back faster than anyone can check a classifier’s work. Even the month on this bench is close to that line: 200 units produce eight VOUT failures, and eight notes is an afternoon. This recipe starts paying when there are more notes than there are people to read them.
Two causes are visible before anyone reads a note at all: a fixture with a stale calibration
offset, and a lot of output capacitors that measure short at bias, both found by grouping the
day’s failures the same way
limits-without-a-model groups a whole month of
measurements. What is left after that grouping is a genuine one-off, and the operator’s free-text
note is the only signal left for it, unevenly written and often blank. This recipe reads that note
and sorts it into a cause; it never produces a measurement, a margin or a verdict. Pass or fail
was decided by the test executive, at the fixture, against the limits in
evals/bench/corpus/srb5030-test-spec.md, before any of this runs.
The walkthrough, on the bench
No recorded run exists for this recipe yet, so what follows is a stepped walkthrough of the code
against the real production data, not a played trace: the site has no measured results. Two units
from evals/bench/data/production-run-2026-08.csv carry the whole argument.
SRB5030-2608-0011 failed VOUT at 4.9497 V against a 4.9500 V lower limit, on FIX-03, with the
note “low again on fix3, thats 3 today.” Grouping the run’s VOUT failures finds FIX-03 behind six
of the eight, so _group_signature already returns "fixture" before the note is read. The note
agrees, and the route comes out hold_check_fixture: check the calibration record, not the board.
SRB5030-2608-0063 also failed VOUT, on the same fixture, with the note “dead. no vout at all,
u1 not switching.” The group signature is identical: FIX-03, flagged. A rule that only asked which
fixture a failure landed on would send this unit to the same calibration check, wrongly. This
board is not offset by a few tens of millivolts; it reads -0.0300 V, which a calibration offset of
-0.030 V cannot produce from a working board, only from one with no output at all. The classifier
reads the note and returns dead_board, with the note quoted as evidence, and the route rule
sends this one to failure_analysis instead, overriding the group signature.
View code: route
def _route(group_signature: str, cause: str) -> str:
"""Code, always: the note's label never decides where a unit goes by itself.
`dead_board` outranks a group signature on purpose: a fixture offset of a few tens of
millivolts cannot produce zero output, so a dead-board claim is never explained by the
fixture, whatever the count says. Below that, the group signature outranks the note, because
a board that merely reads low is exactly what a fixture offset also produces (guide section
3). An uncorroborated `fixture_signature` -- one note, no group pattern behind it -- is not
enough on its own to skip failure analysis, for the same reason in reverse: trusting it wrongly
lets a real defect through on a guess, where the group check would have caught a real fixture
fault for free. See the recipe page for the cost each direction of that mistake carries.
"""
if cause == "dead_board":
return "failure_analysis"
if group_signature == "fixture":
return "hold_check_fixture"
if group_signature == "lot":
return "hold_check_lot"
if cause == "board_low_general":
return "failure_analysis"
return "retest_other_fixture"That override is where the asymmetric cost lives. Routing a real fixture problem to failure
analysis by mistake costs an engineer an afternoon confirming a good board is good. Routing a real
defect to a fixture check by mistake lets that board get retested, possibly pass, and ship. The
guide’s own default for an ungrouped VOUT failure, retest once and only escalate on a second
failure, already leans toward catching the costlier mistake, and _route keeps that lean: a
fixture claim skips failure analysis only when the day’s own data backs it, and a dead-board claim
goes through at once because nothing else explains that reading.
The whole thing is a checkpoint, never a disposition on its own:
View code: run
def run(
serial: str,
model: Model,
tracer: Tracer,
*,
measurement: str = DEFAULT_MEASUREMENT,
production_csv: Path = PRODUCTION_CSV,
) -> Disposition:
failures = load_failures(measurement, production_csv=production_csv)
tracer.record(kind="code", decided_by="code", title=f"Load the run's {measurement} failures", detail=f"{len(failures)} rows")
row = next((f for f in failures if f.serial == serial), None)
if row is None:
raise ValueError(f"{serial!r} did not fail {measurement} in {production_csv}")
signature = _group_signature(failures, row)
tracer.record(kind="code", decided_by="code", title="Group failures by fixture and by lot", detail=f"signature={signature}")
cause, evidence = _classify_note(row.note, model, tracer)
route = _route(signature, cause)
tracer.record(kind="code", decided_by="code", title="Route from the cause and the group signature", detail=f"route={route}")
disposition = Disposition(row=row, group_signature=signature, cause=cause, evidence=evidence, route=route)
tracer.record(
kind="code",
decided_by="code",
title="Hold for a person's confirmation",
detail="every disposition needs one before anything is scrapped, reworked, or held",
)
return dispositionconfirm is where a person’s decision is actually recorded, the same shape as the bench’s own
OUTP ON: nothing acts on a proposed route until it is confirmed, whatever the route was.
View code: confirm
def confirm(disposition: Disposition, approved: bool, tracer: Tracer, *, note: str = "") -> str:
"""A person's decision on one proposed disposition. Never automatic, and never skipped: see
`run`'s last step. Returns the route that was actually acted on."""
tracer.record(
kind="code",
decided_by="code",
title="Person confirms the disposition",
detail=f"approved={approved} route={disposition.route}" + (f" note={note!r}" if note else ""),
)
if approved:
return disposition.route
return f"overridden_by_person: {note or 'no reason given'}"What it costs
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
The 190 units that passed VOUT cost nothing here; they never reach this code at all. Neither do the two failures with no note, and the group-flagged failures whose notes agree with the group signature pay for a model call that changes no route. Only the genuine one-offs, like the dead board on FIX-03, are where the tokens buy something the grouping alone could not.
How it fails on a real bench, specifically
A note typed on the wrong row
- How to notice it
- A person reads the disposition for a failing row and the note field is empty or unrelated, while a different row for the same unit turns out to hold the sentence that actually explains it, filed under whichever step the operator happened to be looking at.
- How to test for it
- Pull every note a serial has, across all eight of its rows, not just the note on the row that failed, and check by hand whether a person reading the full set would have reached a different cause than the code did reading one row alone. This recipe does not stitch notes across rows for that unit; a person confirming the disposition is what is meant to catch it.
A group signature that is right about the fixture and wrong about the unit
- How to notice it
- A unit on the flagged fixture is not merely offset, it is dead, and a rule that only asked which fixture a failure landed on would send it to a calibration check that explains nothing about it.
- How to test for it
- Run the recipe on SRB5030-2608-0063: FIX-03 is flagged as a fixture signature, and the note says the board never switches. The route has to come out failure_analysis, not hold_check_fixture, and tests/test_example_bench_test_failure_triage.py checks exactly this case by name.
A cause asserted with nothing behind it
- How to notice it
- A disposition names a cause that a person, reading the note themselves, cannot find any support for.
- How to test for it
- Feed the classifier a note like "FAIL" or "?". The schema requires a quoted evidence string for every cause except no_information, so a reply that names dead_board or board_low_general with empty evidence fails validation, gets one retry, and falls back to no_information rather than being accepted on a guess.
What to measure
Build the labeled set from the answer key: every VOUT failure in
evals/bench/data/production-run-2026-08.csv, with the cause docs/THE-BENCH.md and the failure
analysis guide assign it once someone has checked the fixture record and, where needed, the board.
That is eight rows this month, too few to trust a rate from; collect several months before tuning
the group thresholds or the prompt, and add every disagreement between the recipe’s route and the
confirmed one as it happens.
The confusion that matters is not accuracy on four labels evenly. It is the two ways a fixture
call and a dead_board or board_low_general call can be confused, and the two directions cost
differently. Calling a real defect a fixture problem sends a bad board toward a calibration check
instead of failure analysis, and a board that is only marginally bad can pass a retest and ship.
Calling a real fixture problem a defect sends a good board to failure analysis for nothing, which
costs an engineer’s time and not a shipped unit. Score the first direction, a false fixture or
hold call on a unit the confirmed record says was actually bad, as the number to drive toward
zero, even at the cost of more of the second, cheaper kind.
Variations
- Swap the measurement from VOUT to RIPPLE and the causes to the ones section 4 of the guide
names,
lot_signaturein place offixture_signature: the group check, the schema-and-retry pattern and the confirmation gate carry over unchanged; only the cause list and the routing table’s targets are this measurement’s own. - Retune
_group_signature’smin_countandmin_shareonce a few months of confirmed dispositions exist, from what the confusion matrix above actually shows, not from a guess. - The same shape sorts a support ticket by product area or an incoming lead by fit: a small fixed set of categories, a lookup from label to handler your code already wrote, and a model reading only the free text a rule cannot.
- Move to human approval’s own threshold pattern instead of always pausing, once confirmed data shows most dispositions agree with the proposed route.
Design choices
Why this level, and when to use another approach
Three techniques compose this recipe: routing reads the operator’s note and picks one of four causes the failure analysis guide already lists; structured output keeps that answer in a fixed shape, a cause plus the exact words that support it; and human approval holds every result for a person before it becomes a disposition. Level 3 is enough because the categories are fixed in advance and so is what happens once a unit lands in one, the same argument the routing page makes for choosing a rule or a classifier: try a rule first, and reach for a classifier once the wording varies too much for one to catch reliably. An operator’s note is exactly that kind of wording.
Most of the job is level 0 and never reaches a model. Which unit failed, and on which limit, is a
column read out of evals/bench/data/production-run-2026-08.csv: 200 units, eight steps each,
1,586 rows, already scored PASS or FAIL by the test executive. The routing table in the guide’s
section 1 is a lookup by step number, no judgment involved. And grouping the run’s VOUT failures
by fixture, or by lot, is a count and a share: _group_signature finds that FIX-03 carries six of
the run’s eight VOUT failures, which is limits-without-a-model’s whole argument playing out
inside one recipe rather than across it.
View code: group signature
def _group_signature(failures: list[FailureRow], row: FailureRow, *, min_count: int = 3, min_share: float = 0.5) -> str:
"""Level 0: does this row's fixture, or its lot, already account for most of the run's
failures on this measurement? A `GROUP BY` and a count, nothing else -- see
`evals/bench/corpus/failure-analysis-guide.md` section 7 (fixture) and section 4 (lot), and
`docs/THE-BENCH.md`'s Story 1 and Story 2 for the numbers this threshold is checked against.
`min_count` keeps one or two coincidental failures on the same fixture from reading as a
pattern; `min_share` requires that fixture or lot to be most of the failures, not merely more
than any other single one.
"""
total = len(failures)
if total == 0:
return "none"
fixture_count = sum(1 for f in failures if f.fixture == row.fixture)
if fixture_count >= min_count and fixture_count / total > min_share:
return "fixture"
lot_count = sum(1 for f in failures if f.lot == row.lot)
if lot_count >= min_count and lot_count / total > min_share:
return "lot"
return "none"Only the note is left, and only reading it needs a model. Section 8 of the guide says plainly how uneven that source is: most failures carry no note at all, and the ones that do range from a full diagnosis to a question mark. Nothing the model writes becomes a measurement, a margin or a verdict. A wrong cause routes a unit to the wrong first step at the bench; it does not change what the bench already measured.
Techniques this recipe uses
The highest level it needs is level 3.
Sort incoming items and send each where it belongs
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Tenant, customer or patient messages by urgency
- Support tickets by product area
- Failed units by likely cause: fixture, lot, handling or design
- Bug reports by component and severity
- Monitoring alerts by who should be paged
- Incoming leads by fit
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page