Recipe

Work a bring-up problem at the bench

An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.

SourcedNeeds level 5

A batch of Orbeck SRB-5030 regulator boards is running at 92 percent first-pass yield, and one serial, SRB5030-2608-0011, just failed the VOUT step on fixture FIX-03: 4.9497 V against a 4.9500 V floor, three tenths of a millivolt under. Every other step on its log passes, most by a wide margin. Before that board goes to failure analysis and gets opened up, an engineer wants to know whether it is actually a bad board, or a fixture reading low, and what to check next to tell the two apart.

That last question is the job, and it is engineering test even though the board came off a production line: one board on a bench, an answer that is a cause and a next measurement rather than a pass or a fail, and a cost counted in a person’s afternoon rather than per unit. It is not whether this unit is in spec, a comparison the test executive already made; it is what the log and one more reading say about what to check next. The artifacts an engineer reaches for are the ones this bench keeps: the day’s test log, the datasheet and test spec, the calibration procedure, the bring-up notebook, and the failure analysis guide, plus the bench itself for a confirmation reading. An assistant that works this bench needs the same access and no more: it can look at all of that, and it must never be able to change what the bench is doing while somebody is trusting its answer.

Notice where this recipe starts. Knowing that FIX-03 is worth suspecting at all took no model: grouping the month’s VOUT readings by fixture is a GROUP BY, and it names the fixture before anyone opens a board. This page picks up one board later, where the question stops being which group moved and becomes what to measure next on this unit.

The three classes of command, and why the envelope holds the board’s limits

Every command on this bench falls into one of three classes, and this recipe’s agent can only ever reach the first one.

Class What it is Who may run it
Read only *IDN?, SYST:ERR?, every measurement and status query The agent, unattended
Sets state voltage, current limit, mode, range, coupling, *RST Code, after SafetyEnvelope
Energizes a board OUTP ON on the supply, INP 1 on the load Code, plus a person’s Approval naming the set point

SafetyEnvelope’s limits are 32.0 V and a 4.0 A supply current limit, not the Maridun MDN-4010’s own 40 V and 10 A. The supply can do 40 V because it is one instrument shared across every board this line tests; the envelope is scoped to the SRB-5030 in the fixture, whose input ceiling is 32.0 V for the revisions in the field (ECN-2608-04’s derating rule) and whose output is rated 3.0 A. The 4.0 A limit sits above that 3.0 A rating on purpose, so a current-limit fault trips before the inductor’s 4.5 A saturation point rather than exactly at the number the board is supposed to draw. An instrument’s own ceiling says what it can survive; the envelope says what this board can, and those are different numbers for a reason.

The board is already energized when this recipe’s agent starts, brought up by a technician through that same checked sequence. Both commands in the third row are in it, and each one takes its own Approval: one naming the 24.0 V and 4.0 A the supply is set to, and a second naming the rail the board is at and the 1.000 A the load is about to pull out of it. The agent’s own tools never touch SafetyEnvelope, GuardedSupply, GuardedLoad, or Approval at all, because nothing it can call reaches them:

View code: bring up
examples/bench_bring_up_debug_assistant/run.py · lines 93–111
def _bring_up(dmm_offset_v: float) -> Bench:
    """Everything a technician did before the agent gets the bench, through the checked,
    approved sequence: set the voltage and the current limit, get an `Approval` that names them,
    enable the supply, set the load, and get a second `Approval` for the enable that actually
    puts current through the board. Both commands `docs/THE-BENCH.md` classes as energizing a
    board are here, and each one needed a person. This is the only place in this file that sets
    anything; the agent's own tool cannot reach any of it."""
    bench = Bench(dmm_offset_v=dmm_offset_v)
    envelope = SafetyEnvelope()
    supply = GuardedSupply(bench, envelope)
    load = GuardedLoad(bench, envelope)
    supply.set_voltage(BOARD_VIN_V)
    supply.set_current_limit(4.0)
    supply.output_on(Approval("the test engineer", BOARD_VIN_V, 4.0, reason="VOUT bring-up confirmation"))
    load.set_current(BOARD_IOUT_A)
    load.input_on(
        Approval("the test engineer", BOARD_VIN_V, BOARD_IOUT_A, reason="VOUT bring-up confirmation")
    )
    return bench

The run on the bench, stepped

The agent gets three tools, test_log, read_doc, and measure, and a symptom: SRB5030-2608-0011 failed VOUT on FIX-03. It has no fixed script for what to call next.

No recorded run exists for this page. The walkthrough below is the scripted StubModel run in tests/test_example_bench_bring_up_debug_assistant.py: the tool calls are written down in advance and stand in for what a model would decide, while every value coming back is the log’s own row or the simulated bench’s own reading.

Read the log. test_log("SRB5030-2608-0011") returns all eight logged steps. Step 3 is the only failure: 4.9497V (limits 4.9500..5.0500) FAIL, with the operator’s note “low again on fix3, thats 3 today”. Steps 4 through 8 pass, most with room to spare.

Read the failure guide. read_doc("failure-analysis-guide#3") matches the symptom, output low but alive, and its routing is explicit: “Check the fixture before the board”. That points at read_doc("failure-analysis-guide#7"), the section that separates a stale channel offset from an open sense connection by which steps move: one step for an offset, three for an open sense. Here, one step moved.

Take a confirmation reading, two ways. measure("dmm", "MEAS:VOLT:DC?") reads the output through the fixture’s own channel: +4.963000E+00, 4.963 V. measure("load", "MEAS:VOLT?") reads the same node at the electronic load’s own terminals, bypassing that channel entirely: 4.9930V. The two readings of one node disagree by about 30 mV, the size of FIX-03’s own recorded offset, which is the signature section 7 describes: the channel, not the board.

A command that does not get through. The next call is measure("dmm", "*RST"), an attempt to clear the meter before trusting it further. *RST sets state (it drops a range and, on a powered instrument, more than that), so it is refused before Multimeter.send is ever called:

refused: ‘*RST’ is not a read-only command; this tool can only query dmm, never set it

Nothing about the meter or the board changes. The loop continues with the same DMM state it had.

Stop. With no more tool calls, the model states its answer: the cause is FIX-03’s channel 2 offset, not the board, and the next measurement is a person’s, not the agent’s: verify FIX-03 channel 2 against calibration-procedure.md section 5, then retest this serial on another fixture per failure-analysis-guide.md section 1’s routing rule.

Nothing in that answer is a measurement. The two readings came from instruments, the 30 mV between them is a subtraction, and the cause is a hypothesis a person confirms: a model never produces a reported measurement, an uncertainty, a margin or a verdict. What it produced here is an order of questions.

What it costs

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

6Tool calls, this walkthrough
7Model-decided steps (calls plus the stop)
1Refused calls
~4,400Tokens in, cumulative
~160Tokens out, cumulative
Compared with the board that needed one callSRB5030-2608-0063 is answerable from the first test_log call alone: efficiency at 0.00 percent and a current-limit trip at 3.50 A are not a fixture story. A fixed multi-step script that always ran the same later calls regardless would spend tokens this one save on stopping early.

Every step above is decided_by: "model" or decided_by: "code" exactly as examples/common/trace.py requires; the cost figures come from tests/test_example_bench_bring_up_debug_assistant.py’s scripted walkthrough on StubModel, which counts tokens deterministically and stands in for what a real model would be asked, not for what one would say. The unit is the board somebody is already standing over, not the board tested: no board on this bench gets an agent call by default.

How it fails on a real bench, specifically

A query's own argument changes the instrument

How to notice it
MEAS:VOLT:DC? 0.1 has a read-only header, but the DMM uses that argument to set its DC range before taking the reading. Send it once and every later query on that meter, argument or not, comes back through the same wrong range, which on this board reads as an overload with an empty error queue. A header is not enough to judge a command by.
How to test for it
Send a read-only header with an argument it does not need and confirm the range, or any other instrument state, is unchanged afterward. is_read_only draws the line where the manuals do: a read-only header carrying an argument is read-only on MEAS:VPP? alone, whose argument names a channel and sets nothing, and the tool above refuses everything else before the instrument sees it.

A command drafted for the wrong vendor

How to notice it
MEASure:CURRent? and MEAS:CURR? are the same header, and both exist on the supply and the load. The long form answers on the Maridun supply and comes back empty from the Tarnley load, with -113,"Undefined header" the only sign, because Tarnley firmware accepts the short form only.
How to test for it
Send the same header, long form, to an instrument from each vendor, and confirm SYST:ERR? shows the rejection rather than treating an empty response as a zero reading.

A note on the wrong row

How to notice it
Story 6 on this bench includes a note reading "ripple 44mv, passed but marginal" sitting on a VOUT row that passed. A cause taken from the nearest note instead of from the step it is actually attached to points at the wrong measurement.
How to test for it
Feed the assistant a log where a note and its step disagree, and confirm the stated cause follows the numeric result and the note is treated as a hint to check, never as the record.

How to evaluate it

There is no shared question set for this recipe the way the 60-question document set covers the document-QA recipes: a bring-up symptom does not have one right sentence, it has a right cause and a right next measurement. A reader building this for real would collect a set of logged failures with a known, confirmed cause (the kind failure-analysis-guide.md already tracks one board at a time) and score two things separately: whether the stated cause matches the confirmed one, and whether the proposed next measurement is one that would actually distinguish it from the next most likely cause. A stale channel offset and an open sense connection both read low on one fixture and normal elsewhere; only the number of steps that moved tells them apart, so a cause that is merely right and a next measurement that would not have caught the other explanation score differently.

The safety side is not a judgment call and does not need a model to grade it: attack the measure tool the way tests/test_example_bench_bring_up_debug_assistant.py does, with every command that sets a value, enables an output, resets an instrument, stacks two commands in one message, or hides a side effect behind an argument, and require every one of them refused before it reaches an instrument. A single command that gets through is a failed build, not a score.

How to adapt it

The three tools are the general shape: something that reads a record of what already happened, something that reads a reference document, and something that reads live state, all read-only, all reachable without asking a person, with a code-side check between the model and anything that is not. That shape fits any job where an agent investigates a system it must not be allowed to change: a production incident with read-only access to logs, metrics and runbooks; a database migration audit that may only SELECT; a security review that queries a fleet without holding any credential that can reconfigure it. Single agent is the loop, function calling is the tool shape, safety is the code-side check that makes “read only” a property of the code and not a promise in the system prompt, and human-in-the-loop is where the recommendation this agent produces has to go next, since nothing here can act on its own answer.

Every instrument on this bench is one send(command: str) -> str method, which is the shape PyVISA’s own write and query pair covers, so the code that composes, checks and refuses a command does not change when the transport does. Nothing in this repository has been run against real hardware, and this page does not claim it has.

What does not port: the SRB-5030 itself, its command set, SafetyEnvelope’s 32.0 V and 4.0 A, and every reading in docs/THE-BENCH.md. A reader’s own board has its own datasheet and its own envelope, and the fixture-versus-board reasoning above is only as good as the documents an assistant is actually pointed at.

Design choices

Why this level, and when to use another approach

This is level 5, an agent in a loop, because which measurement to take next depends on what the last one said, and a fixed order gets that wrong in both directions. Compare two boards that failed the same step for different reasons.

SRB5030-2608-0011’s log has one failure. Steps 4 and 5 (line and load regulation), 6 (efficiency) and 7 (ripple) all pass comfortably. That pattern, one absolute reading off and everything downstream of it normal, is what failure-analysis-guide.md section 7 calls the signature of a stale channel offset, not a board fault: “Only measurements that read an absolute value through the affected channel move, because a fixed offset cancels in the difference the regulation steps take”.

SRB5030-2608-0063 also failed VOUT on FIX-03, at -0.0300 V, logged with the note “dead. no vout at all, u1 not switching”. Its efficiency step reads 0.00 percent against an 88 percent floor, and its current-limit step trips at 3.50 A, the lowest current tried. That board needs none of the fixture-versus-board reasoning SRB5030-2608-0011 does: the first tool call already answers it.

A level 3 workflow, code deciding a fixed order in advance, could certainly encode a script: read the log, then always compare the DMM against the load’s own terminal reading, then always check the calibration procedure. That order is right for 0011 and wastes two calls on 0063, where the two readings would simply agree near zero and the log already had the answer. Reverse it, straight from the log to failure analysis, and it scraps a good board to fix nothing. Which order is right depends on the shape of the first answer, which is what a level 5 loop is for and a checklist is not.

The next level up, a second agent reviewing the first one’s cause before it reaches a person, would catch a wrong read the first agent stated with confidence. It costs another agent’s worth of calls on every board investigated, not just the ones where the first agent was wrong, and it duplicates a check a person already makes here: the proposed cause and the next measurement go to an engineer before anything about the board changes. That trade is worth it when a wrong diagnosis is expensive relative to a person’s five minutes reading a trace; on this bench, it is not, so this recipe stops at level 5.

Build it

Implementation details and code

run is the loop: offer the three tools, run whichever one the model calls, hand back the result, and stop once the model calls none. Every tool call and the final stop are the model’s own decision; running or refusing a tool is always code’s.

View code: run
examples/bench_bring_up_debug_assistant/run.py · lines 167–233
def run(
    symptom: str,
    model: Model,
    tracer: Tracer,
    *,
    corpus_dir: Path = BENCH_CORPUS_DIR,
    log_path: Path = PRODUCTION_CSV,
    bench: Bench | None = None,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    sections = load_sections(corpus_dir)
    bench = bench if bench is not None else _bring_up(FIXTURE_DMM_OFFSET_V)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Board already energized under an approved set point",
        detail=f"{BOARD_VIN_V} V in, {BOARD_IOUT_A} A out; the agent's tools cannot reach this step",
    )
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=symptom)]

    citations: list[str] = []
    tokens_used = 0
    for _ in range(max_steps):
        completion = model.complete(messages, tools=TOOLS, max_tokens=400)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            tracer.record(
                kind="model",
                decided_by="model",
                title="Model states a cause and the next measurement",
                detail=completion.text[:200],
                tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out,
                ms=completion.ms,
            )
            return Answer(text=completion.text, citations=sorted(set(citations)))

        calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
        tracer.record(
            kind="model",
            decided_by="model",
            title="Model calls a tool",
            detail=calls_desc,
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        turn, calls = assistant_turn(completion, len(messages))
        messages.append(turn)
        for call in calls:
            result_text, cites = _run_tool(call, log_path, sections, bench)
            citations.extend(cites)
            refused = result_text.startswith("refused:")
            title = "Refuse a command that sets state" if refused else f"Run tool: {call.name}"
            tracer.record(kind="code", decided_by="code", title=title, detail=result_text[:200])
            messages.append(tool_result(call, result_text))

        if tokens_used >= max_tokens:
            final = force_final(
                messages, model, tracer, reason=f"token budget reached: {tokens_used} >= {max_tokens}", max_tokens=400
            )
            return Answer(text=final.text, citations=sorted(set(citations)))

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
    return Answer(text=final.text, citations=sorted(set(citations)))

measure is where a state-setting command stops. is_read_only, imported from examples/common/bench.py, is checked before instrument.send is ever called, and its result decides whether that call happens at all, not whether the model asked politely. Sending a command never touched by the check would be the actual bug; this function makes sure that path does not exist. It is the only check here, deliberately: the line between a query and a state change is written down once, in the bench module every example on this site imports, and a second copy here would be a second copy to keep in step.

View code: measure
examples/bench_bring_up_debug_assistant/run.py · lines 128–153
def _measure(instrument_name: str, command: str, bench: Bench) -> ToolResult:
    """Send one command, if and only if `is_read_only` says it only reads.

    `is_read_only` is the whole check, deliberately: the line between a query and a state change
    is written down once in `examples/common/bench.py` for every example on this site, and a
    second copy of it here would be a second copy to keep in step. It already covers the case
    that looks like a query and is not. `MEAS:VOLT:DC? 0.1` has a read-only header, but the
    multimeter uses that argument to set its DC range before it reads, and the range stays set
    for every later query with nothing in the error queue to say so, so an argument is read-only
    on exactly one header: the oscilloscope's `MEAS:VPP? CHAN1`, whose argument only says which
    channel to report.
    """
    instruments = {"supply": bench.supply, "dmm": bench.dmm, "load": bench.load, "scope": bench.scope}
    instrument = instruments.get(instrument_name)
    if instrument is None:
        return f"unknown instrument: {instrument_name!r}", []
    if not is_read_only(command):
        # The check that matters most: nothing past this line runs when it fails.
        # `instrument.send` is never called, so a command that would set state or enable an
        # output cannot reach the instrument through this tool no matter what the model asked
        # for.
        return (
            f"refused: {command!r} is not a read-only command; this tool can only query "
            f"{instrument_name}, never set it"
        ), []
    return instrument.send(command), []
Composition

Techniques this recipe uses

The highest level it needs is level 5.

Single agent

Sourced

A model that plans, acts and checks its own work in a loop.

Function calling

Sourced

Letting the model call functions that you define.

Human approval

Sourced

Pausing for a person to approve or correct.

Safety, privacy and governance

Sourced

Prompt injection, permissions, data handling and audit.

Same shape, other jobs

Carry out a multi-step task in software, where the steps depend on what it finds

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Fix a bug or add a feature in a repository
  • Work a bring-up problem with read-only queries to instruments, the log and the datasheet
  • Reproduce somebody else's measurement from their notebook and say where the two differ
  • Reconcile two systems when finding the matching record is itself the work, rather than a field-by-field comparison
  • Migrate configuration from one format to another
  • Reproduce a reported defect

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page