# Work a bring-up problem at the bench

_Recipe · needs level 5_

An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.


A batch of Orbeck SRB-5030 regulator boards is running at 92 percent first-pass yield, and one
serial, SRB5030-2608-0011, just failed the VOUT step on fixture FIX-03: 4.9497 V against a
4.9500 V floor, three tenths of a millivolt under. Every other step on its log passes, most by a
wide margin. Before that board goes to failure analysis and gets opened up, an engineer wants to
know whether it is actually a bad board, or a fixture reading low, and what to check next to tell
the two apart.

That last question is the job, and it is engineering test even though the board came off a
production line: one board on a bench, an answer that is a cause and a next measurement rather
than a pass or a fail, and a cost counted in a person's afternoon rather than per unit. It is not
whether this unit is in spec, a comparison the test executive already made; it is what the log and
one more reading say about what to check next. The
artifacts an engineer reaches for are the ones this bench keeps: the day's test log,
the datasheet and test spec, the calibration procedure, the bring-up notebook, and the failure
analysis guide, plus the bench itself for a confirmation reading. An assistant that works this
bench needs the same access and no more: it can look at all of that, and it must never be able to
change what the bench is doing while somebody is trusting its answer.

Notice where this recipe starts. Knowing that FIX-03 is worth suspecting at all took no model:
[grouping the month's VOUT readings by fixture](/gradient_ascent/recipes/limits-without-a-model/)
is a `GROUP BY`, and it names the fixture before anyone opens a board. This page picks up one
board later, where the question stops being which group moved and becomes what to measure next on
this unit.

## The three classes of command, and why the envelope holds the board's limits

Every command on this bench falls into one of three classes, and this recipe's agent can only
ever reach the first one.

| Class | What it is | Who may run it |
| --- | --- | --- |
| Read only | `*IDN?`, `SYST:ERR?`, every measurement and status query | The agent, unattended |
| Sets state | voltage, current limit, mode, range, coupling, `*RST` | Code, after `SafetyEnvelope` |
| Energizes a board | `OUTP ON` on the supply, `INP 1` on the load | Code, plus a person's `Approval` naming the set point |

`SafetyEnvelope`'s limits are 32.0 V and a 4.0 A supply current limit, not the Maridun MDN-4010's
own 40 V and 10 A. The supply can do 40 V because it is one instrument shared across every board
this line tests; the envelope is scoped to the SRB-5030 in the fixture, whose input ceiling is
32.0 V for the revisions in the field (ECN-2608-04's derating rule) and whose output is rated
3.0 A. The 4.0 A limit sits above that 3.0 A rating on purpose, so a current-limit fault trips
before the inductor's 4.5 A saturation point rather than exactly at the number the board is
supposed to draw. An instrument's own ceiling says what it can survive; the envelope says what
this board can, and those are different numbers for a reason.

The board is already energized when this recipe's agent starts, brought up by a technician
through that same checked sequence. Both commands in the third row are in it, and each one takes
its own `Approval`: one naming the 24.0 V and 4.0 A the supply is set to, and a second naming the
rail the board is at and the 1.000 A the load is about to pull out of it. The agent's own tools
never touch `SafetyEnvelope`, `GuardedSupply`, `GuardedLoad`, or `Approval` at all, because
nothing it can call reaches them:

`examples/bench_bring_up_debug_assistant/run.py` (lines 93-111)

```python
def _bring_up(dmm_offset_v: float) -> Bench:
    """Everything a technician did before the agent gets the bench, through the checked,
    approved sequence: set the voltage and the current limit, get an `Approval` that names them,
    enable the supply, set the load, and get a second `Approval` for the enable that actually
    puts current through the board. Both commands `docs/THE-BENCH.md` classes as energizing a
    board are here, and each one needed a person. This is the only place in this file that sets
    anything; the agent's own tool cannot reach any of it."""
    bench = Bench(dmm_offset_v=dmm_offset_v)
    envelope = SafetyEnvelope()
    supply = GuardedSupply(bench, envelope)
    load = GuardedLoad(bench, envelope)
    supply.set_voltage(BOARD_VIN_V)
    supply.set_current_limit(4.0)
    supply.output_on(Approval("the test engineer", BOARD_VIN_V, 4.0, reason="VOUT bring-up confirmation"))
    load.set_current(BOARD_IOUT_A)
    load.input_on(
        Approval("the test engineer", BOARD_VIN_V, BOARD_IOUT_A, reason="VOUT bring-up confirmation")
    )
    return bench
```

## The run on the bench, stepped

The agent gets three tools, `test_log`, `read_doc`, and `measure`, and a symptom: SRB5030-2608-0011
failed VOUT on FIX-03. It has no fixed script for what to call next.

No recorded run exists for this page. The walkthrough below is the scripted `StubModel` run in
`tests/test_example_bench_bring_up_debug_assistant.py`: the tool calls are written down in
advance and stand in for what a model would decide, while every value coming back is the log's own
row or the simulated bench's own reading.

**Read the log.** `test_log("SRB5030-2608-0011")` returns all eight logged steps. Step 3 is the
only failure: `4.9497V (limits 4.9500..5.0500) FAIL`, with the operator's note "low again on
fix3, thats 3 today". Steps 4 through 8 pass, most with room to spare.

**Read the failure guide.** `read_doc("failure-analysis-guide#3")` matches the symptom, output low
but alive, and its routing is explicit: "Check the fixture before the board". That points at
`read_doc("failure-analysis-guide#7")`, the section that separates a stale channel offset from an
open sense connection by which steps move: one step for an offset, three for an open sense. Here,
one step moved.

**Take a confirmation reading, two ways.** `measure("dmm", "MEAS:VOLT:DC?")` reads the output
through the fixture's own channel: `+4.963000E+00`, 4.963 V. `measure("load", "MEAS:VOLT?")`
reads the same node at the electronic load's own terminals, bypassing that channel entirely:
`4.9930V`. The two readings of one node disagree by about 30 mV, the size of FIX-03's own
recorded offset, which is the signature section 7 describes: the channel, not the board.

**A command that does not get through.** The next call is `measure("dmm", "*RST")`, an attempt to
clear the meter before trusting it further. `*RST` sets state (it drops a range and, on a powered
instrument, more than that), so it is refused before `Multimeter.send` is ever called:

    refused: '*RST' is not a read-only command; this tool can only query dmm, never set it

Nothing about the meter or the board changes. The loop continues with the same DMM state it had.

**Stop.** With no more tool calls, the model states its answer: the cause is FIX-03's channel 2
offset, not the board, and the next measurement is a person's, not the agent's: verify FIX-03
channel 2 against `calibration-procedure.md` section 5, then retest this serial on another
fixture per `failure-analysis-guide.md` section 1's routing rule.

Nothing in that answer is a measurement. The two readings came from instruments, the 30 mV between
them is a subtraction, and the cause is a hypothesis a person confirms: a model never produces a
reported measurement, an uncertainty, a margin or a verdict. What it produced here is an order of
questions.

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Tool calls, this walkthrough:** 6
- **Model-decided steps (calls plus the stop):** 7
- **Refused calls:** 1
- **Tokens in, cumulative:** ~4,400
- **Tokens out, cumulative:** ~160

**Compared with the board that needed one call.** SRB5030-2608-0063 is answerable from the first test_log call alone: efficiency at 0.00 percent and a current-limit trip at 3.50 A are not a fixture story. A fixed multi-step script that always ran the same later calls regardless would spend tokens this one save on stopping early.

Every step above is `decided_by: "model"` or `decided_by: "code"` exactly as
`examples/common/trace.py` requires; the cost figures come from
`tests/test_example_bench_bring_up_debug_assistant.py`'s scripted walkthrough on `StubModel`,
which counts tokens deterministically and stands in for what a real model would be asked, not for
what one would say. The unit is the board somebody is already standing over, not the board tested:
no board on this bench gets an agent call by default.

## How it fails on a real bench, specifically

### A query's own argument changes the instrument

- **How to notice it:** MEAS:VOLT:DC? 0.1 has a read-only header, but the DMM uses that argument to set its DC range before taking the reading. Send it once and every later query on that meter, argument or not, comes back through the same wrong range, which on this board reads as an overload with an empty error queue. A header is not enough to judge a command by.
- **How to test for it:** Send a read-only header with an argument it does not need and confirm the range, or any other instrument state, is unchanged afterward. is_read_only draws the line where the manuals do: a read-only header carrying an argument is read-only on MEAS:VPP? alone, whose argument names a channel and sets nothing, and the tool above refuses everything else before the instrument sees it.

### A command drafted for the wrong vendor

- **How to notice it:** MEASure:CURRent? and MEAS:CURR? are the same header, and both exist on the supply and the load. The long form answers on the Maridun supply and comes back empty from the Tarnley load, with -113,"Undefined header" the only sign, because Tarnley firmware accepts the short form only.
- **How to test for it:** Send the same header, long form, to an instrument from each vendor, and confirm SYST:ERR? shows the rejection rather than treating an empty response as a zero reading.

### A note on the wrong row

- **How to notice it:** Story 6 on this bench includes a note reading "ripple 44mv, passed but marginal" sitting on a VOUT row that passed. A cause taken from the nearest note instead of from the step it is actually attached to points at the wrong measurement.
- **How to test for it:** Feed the assistant a log where a note and its step disagree, and confirm the stated cause follows the numeric result and the note is treated as a hint to check, never as the record.

## How to evaluate it

There is no shared question set for this recipe the way the 60-question document set covers the
document-QA recipes: a bring-up symptom does not have one right sentence, it has a right cause and
a right next measurement. A reader building this for real would collect a set of logged
failures with a known, confirmed cause (the kind `failure-analysis-guide.md` already tracks one
board at a time) and score two things separately: whether the stated cause matches the confirmed
one, and whether the proposed next measurement is one that would actually distinguish it from the
next most likely cause. A stale channel offset and an open sense connection both read low on one
fixture and normal elsewhere; only the number of steps that moved tells them apart, so a cause
that is merely right and a next measurement that would not have caught the other explanation
score differently.

The safety side is not a judgment call and does not need a model to grade it: attack the `measure`
tool the way `tests/test_example_bench_bring_up_debug_assistant.py` does, with every command that
sets a value, enables an output, resets an instrument, stacks two commands in one message, or
hides a side effect behind an argument, and require every one of them refused before it reaches an
instrument. A single command that gets through is a failed build, not a score.

## How to adapt it

The three tools are the general shape: something that reads a record of what already happened,
something that reads a reference document, and something that reads live state, all read-only,
all reachable without asking a person, with a code-side check between the model and anything that
is not. That shape fits any job where an agent investigates a system it must not be allowed to
change: a production incident with read-only access to logs, metrics and runbooks; a database
migration audit that may only `SELECT`; a security review that queries a fleet without holding any
credential that can reconfigure it. [Single agent](/gradient_ascent/techniques/single-agent/) is
the loop, [function calling](/gradient_ascent/techniques/function-calling/) is the tool shape,
[safety](/gradient_ascent/techniques/safety/) is the code-side check that makes "read only" a
property of the code and not a promise in the system prompt, and
[human-in-the-loop](/gradient_ascent/techniques/human-in-the-loop/) is where the recommendation
this agent produces has to go next, since nothing here can act on its own answer.

Every instrument on this bench is one `send(command: str) -> str` method, which is the shape
PyVISA's own `write` and `query` pair covers, so the code that composes, checks and refuses a
command does not change when the transport does. Nothing in this repository has been run against
real hardware, and this page does not claim it has.

What does not port: the SRB-5030 itself, its command set, `SafetyEnvelope`'s 32.0 V and 4.0 A, and
every reading in `docs/THE-BENCH.md`. A reader's own board has its own datasheet and its own
envelope, and the fixture-versus-board reasoning above is only as good as the documents an
assistant is actually pointed at.
## Design choices

### Why this level, and when to use another approach

This is level 5, an agent in a loop, because which measurement to take next depends on what the
last one said, and a fixed order gets that wrong in both directions. Compare two boards that
failed the same step for different reasons.

SRB5030-2608-0011's log has one failure. Steps 4 and 5 (line and load regulation), 6 (efficiency)
and 7 (ripple) all pass comfortably. That pattern, one absolute reading off and everything
downstream of it normal, is what `failure-analysis-guide.md` section 7 calls the signature of a
stale channel offset, not a board fault: "Only measurements that read an absolute value through
the affected channel move, because a fixed offset cancels in the difference the regulation steps
take".

SRB5030-2608-0063 also failed VOUT on FIX-03, at -0.0300 V, logged with the note "dead. no vout at
all, u1 not switching". Its efficiency step reads 0.00 percent against an 88 percent floor, and
its current-limit step trips at 3.50 A, the lowest current tried. That board needs none of the
fixture-versus-board reasoning SRB5030-2608-0011 does: the first tool call already answers it.

A level 3 workflow, code deciding a fixed order in advance, could certainly encode a script: read
the log, then always compare the DMM against the load's own terminal reading, then always check
the calibration procedure. That order is right for 0011 and wastes two calls on 0063, where the
two readings would simply agree near zero and the log already had the answer. Reverse it, straight
from the log to failure analysis, and it scraps a good board to fix nothing. Which order is right
depends on the shape of the first answer, which is what a level 5 loop is for and a checklist is
not.

The next level up, [a second agent](/gradient_ascent/techniques/orchestrator-workers/) reviewing
the first one's cause before it reaches a person, would catch a wrong read the first agent stated
with confidence. It costs another agent's worth of calls on every board investigated, not just the
ones where the first agent was wrong, and it duplicates a check a person already makes here: the
proposed cause and the next measurement go to an engineer before anything about the board changes.
That trade is worth it when a wrong diagnosis is expensive relative to a person's five minutes
reading a trace; on this bench, it is not, so this recipe stops at level 5.

## Build it

### Implementation details and code

`run` is the loop: offer the three tools, run whichever one the model calls, hand back the
result, and stop once the model calls none. Every tool call and the final stop are the model's
own decision; running or refusing a tool is always code's.

`examples/bench_bring_up_debug_assistant/run.py` (lines 167-233)

```python
def run(
    symptom: str,
    model: Model,
    tracer: Tracer,
    *,
    corpus_dir: Path = BENCH_CORPUS_DIR,
    log_path: Path = PRODUCTION_CSV,
    bench: Bench | None = None,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    sections = load_sections(corpus_dir)
    bench = bench if bench is not None else _bring_up(FIXTURE_DMM_OFFSET_V)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Board already energized under an approved set point",
        detail=f"{BOARD_VIN_V} V in, {BOARD_IOUT_A} A out; the agent's tools cannot reach this step",
    )
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=symptom)]

    citations: list[str] = []
    tokens_used = 0
    for _ in range(max_steps):
        completion = model.complete(messages, tools=TOOLS, max_tokens=400)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            tracer.record(
                kind="model",
                decided_by="model",
                title="Model states a cause and the next measurement",
                detail=completion.text[:200],
                tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out,
                ms=completion.ms,
            )
            return Answer(text=completion.text, citations=sorted(set(citations)))

        calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
        tracer.record(
            kind="model",
            decided_by="model",
            title="Model calls a tool",
            detail=calls_desc,
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        turn, calls = assistant_turn(completion, len(messages))
        messages.append(turn)
        for call in calls:
            result_text, cites = _run_tool(call, log_path, sections, bench)
            citations.extend(cites)
            refused = result_text.startswith("refused:")
            title = "Refuse a command that sets state" if refused else f"Run tool: {call.name}"
            tracer.record(kind="code", decided_by="code", title=title, detail=result_text[:200])
            messages.append(tool_result(call, result_text))

        if tokens_used >= max_tokens:
            final = force_final(
                messages, model, tracer, reason=f"token budget reached: {tokens_used} >= {max_tokens}", max_tokens=400
            )
            return Answer(text=final.text, citations=sorted(set(citations)))

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
    return Answer(text=final.text, citations=sorted(set(citations)))
```

`measure` is where a state-setting command stops. `is_read_only`, imported from
`examples/common/bench.py`, is checked before `instrument.send` is ever called, and its result
decides whether that call happens at all, not whether the model asked politely. Sending a command
never touched by the check would be the actual bug; this function makes sure that path does not
exist. It is the only check here, deliberately: the line between a query and a state change is
written down once, in the bench module every example on this site imports, and a second copy here
would be a second copy to keep in step.

`examples/bench_bring_up_debug_assistant/run.py` (lines 128-153)

```python
def _measure(instrument_name: str, command: str, bench: Bench) -> ToolResult:
    """Send one command, if and only if `is_read_only` says it only reads.

    `is_read_only` is the whole check, deliberately: the line between a query and a state change
    is written down once in `examples/common/bench.py` for every example on this site, and a
    second copy of it here would be a second copy to keep in step. It already covers the case
    that looks like a query and is not. `MEAS:VOLT:DC? 0.1` has a read-only header, but the
    multimeter uses that argument to set its DC range before it reads, and the range stays set
    for every later query with nothing in the error queue to say so, so an argument is read-only
    on exactly one header: the oscilloscope's `MEAS:VPP? CHAN1`, whose argument only says which
    channel to report.
    """
    instruments = {"supply": bench.supply, "dmm": bench.dmm, "load": bench.load, "scope": bench.scope}
    instrument = instruments.get(instrument_name)
    if instrument is None:
        return f"unknown instrument: {instrument_name!r}", []
    if not is_read_only(command):
        # The check that matters most: nothing past this line runs when it fails.
        # `instrument.send` is never called, so a command that would set state or enable an
        # output cannot reach the instrument through this tool no matter what the model asked
        # for.
        return (
            f"refused: {command!r} is not a read-only command; this tool can only query "
            f"{instrument_name}, never set it"
        ), []
    return instrument.send(command), []
```



Last reviewed 2026-09-19.
