# Sort failing units and operator notes into causes

_Recipe · needs level 3_

Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.


This is a production test page, and production is the setting it is for. Every day's failing units
land in a queue with two things attached to each one: which limit it missed, and whatever the
operator happened to type while the unit was on the bench. Orbeck's own failure analysis guide,
`evals/bench/corpus/failure-analysis-guide.md`, already lists the causes and a routing table by
step, so this is not an investigation that starts from nothing. It is sorting failures into causes
the guide already names, and sending each one where the guide already says it goes.

The low-volume counterpart of this job is a person reading the notes. An engineer with five
prototypes on the bench wrote the notes themselves and can read all of them back faster than
anyone can check a classifier's work. Even the month on this bench is close to that line: 200
units produce eight VOUT failures, and eight notes is an afternoon. This recipe starts paying when
there are more notes than there are people to read them.

Two causes are visible before anyone reads a note at all: a fixture with a stale calibration
offset, and a lot of output capacitors that measure short at bias, both found by grouping the
day's failures the same way
[limits-without-a-model](/gradient_ascent/recipes/limits-without-a-model/) groups a whole month of
measurements. What is left after that grouping is a genuine one-off, and the operator's free-text
note is the only signal left for it, unevenly written and often blank. This recipe reads that note
and sorts it into a cause; it never produces a measurement, a margin or a verdict. Pass or fail
was decided by the test executive, at the fixture, against the limits in
`evals/bench/corpus/srb5030-test-spec.md`, before any of this runs.

## The walkthrough, on the bench

No recorded run exists for this recipe yet, so what follows is a stepped walkthrough of the code
against the real production data, not a played trace: the site has no measured results. Two units
from `evals/bench/data/production-run-2026-08.csv` carry the whole argument.

**SRB5030-2608-0011** failed VOUT at 4.9497 V against a 4.9500 V lower limit, on FIX-03, with the
note "low again on fix3, thats 3 today." Grouping the run's VOUT failures finds FIX-03 behind six
of the eight, so `_group_signature` already returns `"fixture"` before the note is read. The note
agrees, and the route comes out `hold_check_fixture`: check the calibration record, not the board.

**SRB5030-2608-0063** also failed VOUT, on the same fixture, with the note "dead. no vout at all,
u1 not switching." The group signature is identical: FIX-03, flagged. A rule that only asked which
fixture a failure landed on would send this unit to the same calibration check, wrongly. This
board is not offset by a few tens of millivolts; it reads -0.0300 V, which a calibration offset of
-0.030 V cannot produce from a working board, only from one with no output at all. The classifier
reads the note and returns `dead_board`, with the note quoted as evidence, and the route rule
sends this one to `failure_analysis` instead, overriding the group signature.

`examples/bench_test_failure_triage/run.py` (lines 192-212)

```python
def _route(group_signature: str, cause: str) -> str:
    """Code, always: the note's label never decides where a unit goes by itself.

    `dead_board` outranks a group signature on purpose: a fixture offset of a few tens of
    millivolts cannot produce zero output, so a dead-board claim is never explained by the
    fixture, whatever the count says. Below that, the group signature outranks the note, because
    a board that merely reads low is exactly what a fixture offset also produces (guide section
    3). An uncorroborated `fixture_signature` -- one note, no group pattern behind it -- is not
    enough on its own to skip failure analysis, for the same reason in reverse: trusting it wrongly
    lets a real defect through on a guess, where the group check would have caught a real fixture
    fault for free. See the recipe page for the cost each direction of that mistake carries.
    """
    if cause == "dead_board":
        return "failure_analysis"
    if group_signature == "fixture":
        return "hold_check_fixture"
    if group_signature == "lot":
        return "hold_check_lot"
    if cause == "board_low_general":
        return "failure_analysis"
    return "retest_other_fixture"
```

That override is where the asymmetric cost lives. Routing a real fixture problem to failure
analysis by mistake costs an engineer an afternoon confirming a good board is good. Routing a real
defect to a fixture check by mistake lets that board get retested, possibly pass, and ship. The
guide's own default for an ungrouped VOUT failure, retest once and only escalate on a second
failure, already leans toward catching the costlier mistake, and `_route` keeps that lean: a
fixture claim skips failure analysis only when the day's own data backs it, and a dead-board claim
goes through at once because nothing else explains that reading.

The whole thing is a checkpoint, never a disposition on its own:

`examples/bench_test_failure_triage/run.py` (lines 215-245)

```python
def run(
    serial: str,
    model: Model,
    tracer: Tracer,
    *,
    measurement: str = DEFAULT_MEASUREMENT,
    production_csv: Path = PRODUCTION_CSV,
) -> Disposition:
    failures = load_failures(measurement, production_csv=production_csv)
    tracer.record(kind="code", decided_by="code", title=f"Load the run's {measurement} failures", detail=f"{len(failures)} rows")

    row = next((f for f in failures if f.serial == serial), None)
    if row is None:
        raise ValueError(f"{serial!r} did not fail {measurement} in {production_csv}")

    signature = _group_signature(failures, row)
    tracer.record(kind="code", decided_by="code", title="Group failures by fixture and by lot", detail=f"signature={signature}")

    cause, evidence = _classify_note(row.note, model, tracer)

    route = _route(signature, cause)
    tracer.record(kind="code", decided_by="code", title="Route from the cause and the group signature", detail=f"route={route}")

    disposition = Disposition(row=row, group_signature=signature, cause=cause, evidence=evidence, route=route)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Hold for a person's confirmation",
        detail="every disposition needs one before anything is scrapped, reworked, or held",
    )
    return disposition
```

`confirm` is where a person's decision is actually recorded, the same shape as the bench's own
`OUTP ON`: nothing acts on a proposed route until it is confirmed, whatever the route was.

`examples/bench_test_failure_triage/run.py` (lines 248-259)

```python
def confirm(disposition: Disposition, approved: bool, tracer: Tracer, *, note: str = "") -> str:
    """A person's decision on one proposed disposition. Never automatic, and never skipped: see
    `run`'s last step. Returns the route that was actually acted on."""
    tracer.record(
        kind="code",
        decided_by="code",
        title="Person confirms the disposition",
        detail=f"approved={approved} route={disposition.route}" + (f" note={note!r}" if note else ""),
    )
    if approved:
        return disposition.route
    return f"overridden_by_person: {note or 'no reason given'}"
```

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Tokens in, one note:** 244
- **Tokens out, one note:** 18
- **Model calls, worst case:** 2 (one retry)
- **VOUT failures reaching the model, this run:** 6 of 8

**Compared with classifying every row instead of only the failures.** load_failures filters the file’s 1,586 rows to 8 before any note is read, and two of those eight carry no note and never reach the model either. Skipping that filter and asking the model about every row would spend roughly 262 tokens in and out per row across 1,586 rows, about 416,000 tokens a run, to answer the same 6 questions this recipe answers for about 1,570.

The 190 units that passed VOUT cost nothing here; they never reach this code at all. Neither do
the two failures with no note, and the group-flagged failures whose notes agree with the group
signature pay for a model call that changes no route. Only the genuine one-offs, like the dead
board on FIX-03, are where the tokens buy something the grouping alone could not.

## How it fails on a real bench, specifically

### A note typed on the wrong row

- **How to notice it:** A person reads the disposition for a failing row and the note field is empty or unrelated, while a different row for the same unit turns out to hold the sentence that actually explains it, filed under whichever step the operator happened to be looking at.
- **How to test for it:** Pull every note a serial has, across all eight of its rows, not just the note on the row that failed, and check by hand whether a person reading the full set would have reached a different cause than the code did reading one row alone. This recipe does not stitch notes across rows for that unit; a person confirming the disposition is what is meant to catch it.

### A group signature that is right about the fixture and wrong about the unit

- **How to notice it:** A unit on the flagged fixture is not merely offset, it is dead, and a rule that only asked which fixture a failure landed on would send it to a calibration check that explains nothing about it.
- **How to test for it:** Run the recipe on SRB5030-2608-0063: FIX-03 is flagged as a fixture signature, and the note says the board never switches. The route has to come out failure_analysis, not hold_check_fixture, and tests/test_example_bench_test_failure_triage.py checks exactly this case by name.

### A cause asserted with nothing behind it

- **How to notice it:** A disposition names a cause that a person, reading the note themselves, cannot find any support for.
- **How to test for it:** Feed the classifier a note like "FAIL" or "?". The schema requires a quoted evidence string for every cause except no_information, so a reply that names dead_board or board_low_general with empty evidence fails validation, gets one retry, and falls back to no_information rather than being accepted on a guess.

## What to measure

Build the labeled set from the answer key: every VOUT failure in
`evals/bench/data/production-run-2026-08.csv`, with the cause `docs/THE-BENCH.md` and the failure
analysis guide assign it once someone has checked the fixture record and, where needed, the board.
That is eight rows this month, too few to trust a rate from; collect several months before tuning
the group thresholds or the prompt, and add every disagreement between the recipe's route and the
confirmed one as it happens.

The confusion that matters is not accuracy on four labels evenly. It is the two ways a `fixture`
call and a `dead_board` or `board_low_general` call can be confused, and the two directions cost
differently. Calling a real defect a fixture problem sends a bad board toward a calibration check
instead of failure analysis, and a board that is only marginally bad can pass a retest and ship.
Calling a real fixture problem a defect sends a good board to failure analysis for nothing, which
costs an engineer's time and not a shipped unit. Score the first direction, a false `fixture` or
`hold` call on a unit the confirmed record says was actually bad, as the number to drive toward
zero, even at the cost of more of the second, cheaper kind.

## Variations

- Swap the measurement from VOUT to RIPPLE and the causes to the ones section 4 of the guide
  names, `lot_signature` in place of `fixture_signature`: the group check, the schema-and-retry
  pattern and the confirmation gate carry over unchanged; only the cause list and the routing
  table's targets are this measurement's own.
- Retune `_group_signature`'s `min_count` and `min_share` once a few months of confirmed
  dispositions exist, from what the confusion matrix above actually shows, not from a guess.
- The same shape sorts a support ticket by product area or an incoming lead by fit: a small fixed
  set of categories, a lookup from label to handler your code already wrote, and a model reading
  only the free text a rule cannot.
- Move to [human approval](/gradient_ascent/techniques/human-in-the-loop/)'s own threshold
  pattern instead of always pausing, once confirmed data shows most dispositions agree with the
  proposed route.

## Design choices

### Why this level, and when to use another approach

Three techniques compose this recipe: [routing](/gradient_ascent/techniques/routing/) reads the
operator's note and picks one of four causes the failure analysis guide already lists;
[structured output](/gradient_ascent/techniques/structured-output/) keeps that answer in a fixed
shape, a cause plus the exact words that support it; and
[human approval](/gradient_ascent/techniques/human-in-the-loop/) holds every result for a person
before it becomes a disposition. Level 3 is enough because the categories are fixed in advance and
so is what happens once a unit lands in one, the same argument the routing page makes for choosing
a rule or a classifier: try a rule first, and reach for a classifier once the wording varies too
much for one to catch reliably. An operator's note is exactly that kind of wording.

Most of the job is level 0 and never reaches a model. Which unit failed, and on which limit, is a
column read out of `evals/bench/data/production-run-2026-08.csv`: 200 units, eight steps each,
1,586 rows, already scored PASS or FAIL by the test executive. The routing table in the guide's
section 1 is a lookup by step number, no judgment involved. And grouping the run's VOUT failures
by fixture, or by lot, is a count and a share: `_group_signature` finds that FIX-03 carries six of
the run's eight VOUT failures, which is `limits-without-a-model`'s whole argument playing out
inside one recipe rather than across it.

`examples/bench_test_failure_triage/run.py` (lines 125-143)

```python
def _group_signature(failures: list[FailureRow], row: FailureRow, *, min_count: int = 3, min_share: float = 0.5) -> str:
    """Level 0: does this row's fixture, or its lot, already account for most of the run's
    failures on this measurement? A `GROUP BY` and a count, nothing else -- see
    `evals/bench/corpus/failure-analysis-guide.md` section 7 (fixture) and section 4 (lot), and
    `docs/THE-BENCH.md`'s Story 1 and Story 2 for the numbers this threshold is checked against.
    `min_count` keeps one or two coincidental failures on the same fixture from reading as a
    pattern; `min_share` requires that fixture or lot to be most of the failures, not merely more
    than any other single one.
    """
    total = len(failures)
    if total == 0:
        return "none"
    fixture_count = sum(1 for f in failures if f.fixture == row.fixture)
    if fixture_count >= min_count and fixture_count / total > min_share:
        return "fixture"
    lot_count = sum(1 for f in failures if f.lot == row.lot)
    if lot_count >= min_count and lot_count / total > min_share:
        return "lot"
    return "none"
```

Only the note is left, and only reading it needs a model. Section 8 of the guide says plainly how
uneven that source is: most failures carry no note at all, and the ones that do range from a full
diagnosis to a question mark. Nothing the model writes becomes a measurement, a margin or a
verdict. A wrong cause routes a unit to the wrong first step at the bench; it does not change what
the bench already measured.



Last reviewed 2026-09-19.
