Recipe

Sort failing units and operator notes into causes

Failing measurements and free-text operator notes are sorted into the causes the failure analysis guide already lists, then routed. A person confirms before anything is scrapped or reworked. The categories are known in advance, so this is classification into fixed classes and not an agent.

SourcedNeeds level 3

This is a production test page, and production is the setting it is for. Every day’s failing units land in a queue with two things attached to each one: which limit it missed, and whatever the operator happened to type while the unit was on the bench. Orbeck’s own failure analysis guide, evals/bench/corpus/failure-analysis-guide.md, already lists the causes and a routing table by step, so this is not an investigation that starts from nothing. It is sorting failures into causes the guide already names, and sending each one where the guide already says it goes.

The low-volume counterpart of this job is a person reading the notes. An engineer with five prototypes on the bench wrote the notes themselves and can read all of them back faster than anyone can check a classifier’s work. Even the month on this bench is close to that line: 200 units produce eight VOUT failures, and eight notes is an afternoon. This recipe starts paying when there are more notes than there are people to read them.

Two causes are visible before anyone reads a note at all: a fixture with a stale calibration offset, and a lot of output capacitors that measure short at bias, both found by grouping the day’s failures the same way limits-without-a-model groups a whole month of measurements. What is left after that grouping is a genuine one-off, and the operator’s free-text note is the only signal left for it, unevenly written and often blank. This recipe reads that note and sorts it into a cause; it never produces a measurement, a margin or a verdict. Pass or fail was decided by the test executive, at the fixture, against the limits in evals/bench/corpus/srb5030-test-spec.md, before any of this runs.

The walkthrough, on the bench

No recorded run exists for this recipe yet, so what follows is a stepped walkthrough of the code against the real production data, not a played trace: the site has no measured results. Two units from evals/bench/data/production-run-2026-08.csv carry the whole argument.

SRB5030-2608-0011 failed VOUT at 4.9497 V against a 4.9500 V lower limit, on FIX-03, with the note “low again on fix3, thats 3 today.” Grouping the run’s VOUT failures finds FIX-03 behind six of the eight, so _group_signature already returns "fixture" before the note is read. The note agrees, and the route comes out hold_check_fixture: check the calibration record, not the board.

SRB5030-2608-0063 also failed VOUT, on the same fixture, with the note “dead. no vout at all, u1 not switching.” The group signature is identical: FIX-03, flagged. A rule that only asked which fixture a failure landed on would send this unit to the same calibration check, wrongly. This board is not offset by a few tens of millivolts; it reads -0.0300 V, which a calibration offset of -0.030 V cannot produce from a working board, only from one with no output at all. The classifier reads the note and returns dead_board, with the note quoted as evidence, and the route rule sends this one to failure_analysis instead, overriding the group signature.

View code: route
examples/bench_test_failure_triage/run.py · lines 192–212
def _route(group_signature: str, cause: str) -> str:
    """Code, always: the note's label never decides where a unit goes by itself.

    `dead_board` outranks a group signature on purpose: a fixture offset of a few tens of
    millivolts cannot produce zero output, so a dead-board claim is never explained by the
    fixture, whatever the count says. Below that, the group signature outranks the note, because
    a board that merely reads low is exactly what a fixture offset also produces (guide section
    3). An uncorroborated `fixture_signature` -- one note, no group pattern behind it -- is not
    enough on its own to skip failure analysis, for the same reason in reverse: trusting it wrongly
    lets a real defect through on a guess, where the group check would have caught a real fixture
    fault for free. See the recipe page for the cost each direction of that mistake carries.
    """
    if cause == "dead_board":
        return "failure_analysis"
    if group_signature == "fixture":
        return "hold_check_fixture"
    if group_signature == "lot":
        return "hold_check_lot"
    if cause == "board_low_general":
        return "failure_analysis"
    return "retest_other_fixture"

That override is where the asymmetric cost lives. Routing a real fixture problem to failure analysis by mistake costs an engineer an afternoon confirming a good board is good. Routing a real defect to a fixture check by mistake lets that board get retested, possibly pass, and ship. The guide’s own default for an ungrouped VOUT failure, retest once and only escalate on a second failure, already leans toward catching the costlier mistake, and _route keeps that lean: a fixture claim skips failure analysis only when the day’s own data backs it, and a dead-board claim goes through at once because nothing else explains that reading.

The whole thing is a checkpoint, never a disposition on its own:

View code: run
examples/bench_test_failure_triage/run.py · lines 215–245
def run(
    serial: str,
    model: Model,
    tracer: Tracer,
    *,
    measurement: str = DEFAULT_MEASUREMENT,
    production_csv: Path = PRODUCTION_CSV,
) -> Disposition:
    failures = load_failures(measurement, production_csv=production_csv)
    tracer.record(kind="code", decided_by="code", title=f"Load the run's {measurement} failures", detail=f"{len(failures)} rows")

    row = next((f for f in failures if f.serial == serial), None)
    if row is None:
        raise ValueError(f"{serial!r} did not fail {measurement} in {production_csv}")

    signature = _group_signature(failures, row)
    tracer.record(kind="code", decided_by="code", title="Group failures by fixture and by lot", detail=f"signature={signature}")

    cause, evidence = _classify_note(row.note, model, tracer)

    route = _route(signature, cause)
    tracer.record(kind="code", decided_by="code", title="Route from the cause and the group signature", detail=f"route={route}")

    disposition = Disposition(row=row, group_signature=signature, cause=cause, evidence=evidence, route=route)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Hold for a person's confirmation",
        detail="every disposition needs one before anything is scrapped, reworked, or held",
    )
    return disposition

confirm is where a person’s decision is actually recorded, the same shape as the bench’s own OUTP ON: nothing acts on a proposed route until it is confirmed, whatever the route was.

View code: confirm
examples/bench_test_failure_triage/run.py · lines 248–259
def confirm(disposition: Disposition, approved: bool, tracer: Tracer, *, note: str = "") -> str:
    """A person's decision on one proposed disposition. Never automatic, and never skipped: see
    `run`'s last step. Returns the route that was actually acted on."""
    tracer.record(
        kind="code",
        decided_by="code",
        title="Person confirms the disposition",
        detail=f"approved={approved} route={disposition.route}" + (f" note={note!r}" if note else ""),
    )
    if approved:
        return disposition.route
    return f"overridden_by_person: {note or 'no reason given'}"

What it costs

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

244Tokens in, one note
18Tokens out, one note
2 (one retry)Model calls, worst case
6 of 8VOUT failures reaching the model, this run
Compared with classifying every row instead of only the failuresload_failures filters the file’s 1,586 rows to 8 before any note is read, and two of those eight carry no note and never reach the model either. Skipping that filter and asking the model about every row would spend roughly 262 tokens in and out per row across 1,586 rows, about 416,000 tokens a run, to answer the same 6 questions this recipe answers for about 1,570.

The 190 units that passed VOUT cost nothing here; they never reach this code at all. Neither do the two failures with no note, and the group-flagged failures whose notes agree with the group signature pay for a model call that changes no route. Only the genuine one-offs, like the dead board on FIX-03, are where the tokens buy something the grouping alone could not.

How it fails on a real bench, specifically

A note typed on the wrong row

How to notice it
A person reads the disposition for a failing row and the note field is empty or unrelated, while a different row for the same unit turns out to hold the sentence that actually explains it, filed under whichever step the operator happened to be looking at.
How to test for it
Pull every note a serial has, across all eight of its rows, not just the note on the row that failed, and check by hand whether a person reading the full set would have reached a different cause than the code did reading one row alone. This recipe does not stitch notes across rows for that unit; a person confirming the disposition is what is meant to catch it.

A group signature that is right about the fixture and wrong about the unit

How to notice it
A unit on the flagged fixture is not merely offset, it is dead, and a rule that only asked which fixture a failure landed on would send it to a calibration check that explains nothing about it.
How to test for it
Run the recipe on SRB5030-2608-0063: FIX-03 is flagged as a fixture signature, and the note says the board never switches. The route has to come out failure_analysis, not hold_check_fixture, and tests/test_example_bench_test_failure_triage.py checks exactly this case by name.

A cause asserted with nothing behind it

How to notice it
A disposition names a cause that a person, reading the note themselves, cannot find any support for.
How to test for it
Feed the classifier a note like "FAIL" or "?". The schema requires a quoted evidence string for every cause except no_information, so a reply that names dead_board or board_low_general with empty evidence fails validation, gets one retry, and falls back to no_information rather than being accepted on a guess.

What to measure

Build the labeled set from the answer key: every VOUT failure in evals/bench/data/production-run-2026-08.csv, with the cause docs/THE-BENCH.md and the failure analysis guide assign it once someone has checked the fixture record and, where needed, the board. That is eight rows this month, too few to trust a rate from; collect several months before tuning the group thresholds or the prompt, and add every disagreement between the recipe’s route and the confirmed one as it happens.

The confusion that matters is not accuracy on four labels evenly. It is the two ways a fixture call and a dead_board or board_low_general call can be confused, and the two directions cost differently. Calling a real defect a fixture problem sends a bad board toward a calibration check instead of failure analysis, and a board that is only marginally bad can pass a retest and ship. Calling a real fixture problem a defect sends a good board to failure analysis for nothing, which costs an engineer’s time and not a shipped unit. Score the first direction, a false fixture or hold call on a unit the confirmed record says was actually bad, as the number to drive toward zero, even at the cost of more of the second, cheaper kind.

Variations

  • Swap the measurement from VOUT to RIPPLE and the causes to the ones section 4 of the guide names, lot_signature in place of fixture_signature: the group check, the schema-and-retry pattern and the confirmation gate carry over unchanged; only the cause list and the routing table’s targets are this measurement’s own.
  • Retune _group_signature’s min_count and min_share once a few months of confirmed dispositions exist, from what the confusion matrix above actually shows, not from a guess.
  • The same shape sorts a support ticket by product area or an incoming lead by fit: a small fixed set of categories, a lookup from label to handler your code already wrote, and a model reading only the free text a rule cannot.
  • Move to human approval’s own threshold pattern instead of always pausing, once confirmed data shows most dispositions agree with the proposed route.

Design choices

Why this level, and when to use another approach

Three techniques compose this recipe: routing reads the operator’s note and picks one of four causes the failure analysis guide already lists; structured output keeps that answer in a fixed shape, a cause plus the exact words that support it; and human approval holds every result for a person before it becomes a disposition. Level 3 is enough because the categories are fixed in advance and so is what happens once a unit lands in one, the same argument the routing page makes for choosing a rule or a classifier: try a rule first, and reach for a classifier once the wording varies too much for one to catch reliably. An operator’s note is exactly that kind of wording.

Most of the job is level 0 and never reaches a model. Which unit failed, and on which limit, is a column read out of evals/bench/data/production-run-2026-08.csv: 200 units, eight steps each, 1,586 rows, already scored PASS or FAIL by the test executive. The routing table in the guide’s section 1 is a lookup by step number, no judgment involved. And grouping the run’s VOUT failures by fixture, or by lot, is a count and a share: _group_signature finds that FIX-03 carries six of the run’s eight VOUT failures, which is limits-without-a-model’s whole argument playing out inside one recipe rather than across it.

View code: group signature
examples/bench_test_failure_triage/run.py · lines 125–143
def _group_signature(failures: list[FailureRow], row: FailureRow, *, min_count: int = 3, min_share: float = 0.5) -> str:
    """Level 0: does this row's fixture, or its lot, already account for most of the run's
    failures on this measurement? A `GROUP BY` and a count, nothing else -- see
    `evals/bench/corpus/failure-analysis-guide.md` section 7 (fixture) and section 4 (lot), and
    `docs/THE-BENCH.md`'s Story 1 and Story 2 for the numbers this threshold is checked against.
    `min_count` keeps one or two coincidental failures on the same fixture from reading as a
    pattern; `min_share` requires that fixture or lot to be most of the failures, not merely more
    than any other single one.
    """
    total = len(failures)
    if total == 0:
        return "none"
    fixture_count = sum(1 for f in failures if f.fixture == row.fixture)
    if fixture_count >= min_count and fixture_count / total > min_share:
        return "fixture"
    lot_count = sum(1 for f in failures if f.lot == row.lot)
    if lot_count >= min_count and lot_count / total > min_share:
        return "lot"
    return "none"

Only the note is left, and only reading it needs a model. Section 8 of the guide says plainly how uneven that source is: most failures carry no note at all, and the ones that do range from a full diagnosis to a question mark. Nothing the model writes becomes a measurement, a margin or a verdict. A wrong cause routes a unit to the wrong first step at the bench; it does not change what the bench already measured.

Composition

Techniques this recipe uses

The highest level it needs is level 3.

Routing

Sourced

Sorting inputs and sending each one to the right prompt.

Structured output

Sourced

Getting answers in a fixed format such as JSON.

Human approval

Sourced

Pausing for a person to approve or correct.

Same shape, other jobs

Sort incoming items and send each where it belongs

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Tenant, customer or patient messages by urgency
  • Support tickets by product area
  • Failed units by likely cause: fixture, lot, handling or design
  • Bug reports by component and severity
  • Monitoring alerts by who should be paged
  • Incoming leads by fit

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page