Work a bring-up problem at the bench
An agent with read-only tools, instrument queries, the test log and the datasheet, works a low output down to a cause and proposes the next measurement. Queries run unattended; anything that sets a voltage, a current limit or an output goes through the envelope and a person. Level 5 because each measurement depends on the last.
SourcedNeeds level 5
A batch of Orbeck SRB-5030 regulator boards is running at 92 percent first-pass yield, and one serial, SRB5030-2608-0011, just failed the VOUT step on fixture FIX-03: 4.9497 V against a 4.9500 V floor, three tenths of a millivolt under. Every other step on its log passes, most by a wide margin. Before that board goes to failure analysis and gets opened up, an engineer wants to know whether it is actually a bad board, or a fixture reading low, and what to check next to tell the two apart.
That last question is the job, and it is engineering test even though the board came off a production line: one board on a bench, an answer that is a cause and a next measurement rather than a pass or a fail, and a cost counted in a person’s afternoon rather than per unit. It is not whether this unit is in spec, a comparison the test executive already made; it is what the log and one more reading say about what to check next. The artifacts an engineer reaches for are the ones this bench keeps: the day’s test log, the datasheet and test spec, the calibration procedure, the bring-up notebook, and the failure analysis guide, plus the bench itself for a confirmation reading. An assistant that works this bench needs the same access and no more: it can look at all of that, and it must never be able to change what the bench is doing while somebody is trusting its answer.
Notice where this recipe starts. Knowing that FIX-03 is worth suspecting at all took no model:
grouping the month’s VOUT readings by fixture
is a GROUP BY, and it names the fixture before anyone opens a board. This page picks up one
board later, where the question stops being which group moved and becomes what to measure next on
this unit.
The three classes of command, and why the envelope holds the board’s limits
Every command on this bench falls into one of three classes, and this recipe’s agent can only ever reach the first one.
| Class | What it is | Who may run it |
|---|---|---|
| Read only | *IDN?, SYST:ERR?, every measurement and status query |
The agent, unattended |
| Sets state | voltage, current limit, mode, range, coupling, *RST |
Code, after SafetyEnvelope |
| Energizes a board | OUTP ON on the supply, INP 1 on the load |
Code, plus a person’s Approval naming the set point |
SafetyEnvelope’s limits are 32.0 V and a 4.0 A supply current limit, not the Maridun MDN-4010’s
own 40 V and 10 A. The supply can do 40 V because it is one instrument shared across every board
this line tests; the envelope is scoped to the SRB-5030 in the fixture, whose input ceiling is
32.0 V for the revisions in the field (ECN-2608-04’s derating rule) and whose output is rated
3.0 A. The 4.0 A limit sits above that 3.0 A rating on purpose, so a current-limit fault trips
before the inductor’s 4.5 A saturation point rather than exactly at the number the board is
supposed to draw. An instrument’s own ceiling says what it can survive; the envelope says what
this board can, and those are different numbers for a reason.
The board is already energized when this recipe’s agent starts, brought up by a technician
through that same checked sequence. Both commands in the third row are in it, and each one takes
its own Approval: one naming the 24.0 V and 4.0 A the supply is set to, and a second naming the
rail the board is at and the 1.000 A the load is about to pull out of it. The agent’s own tools
never touch SafetyEnvelope, GuardedSupply, GuardedLoad, or Approval at all, because
nothing it can call reaches them:
View code: bring up
def _bring_up(dmm_offset_v: float) -> Bench:
"""Everything a technician did before the agent gets the bench, through the checked,
approved sequence: set the voltage and the current limit, get an `Approval` that names them,
enable the supply, set the load, and get a second `Approval` for the enable that actually
puts current through the board. Both commands `docs/THE-BENCH.md` classes as energizing a
board are here, and each one needed a person. This is the only place in this file that sets
anything; the agent's own tool cannot reach any of it."""
bench = Bench(dmm_offset_v=dmm_offset_v)
envelope = SafetyEnvelope()
supply = GuardedSupply(bench, envelope)
load = GuardedLoad(bench, envelope)
supply.set_voltage(BOARD_VIN_V)
supply.set_current_limit(4.0)
supply.output_on(Approval("the test engineer", BOARD_VIN_V, 4.0, reason="VOUT bring-up confirmation"))
load.set_current(BOARD_IOUT_A)
load.input_on(
Approval("the test engineer", BOARD_VIN_V, BOARD_IOUT_A, reason="VOUT bring-up confirmation")
)
return benchThe run on the bench, stepped
The agent gets three tools, test_log, read_doc, and measure, and a symptom: SRB5030-2608-0011
failed VOUT on FIX-03. It has no fixed script for what to call next.
No recorded run exists for this page. The walkthrough below is the scripted StubModel run in
tests/test_example_bench_bring_up_debug_assistant.py: the tool calls are written down in
advance and stand in for what a model would decide, while every value coming back is the log’s own
row or the simulated bench’s own reading.
Read the log. test_log("SRB5030-2608-0011") returns all eight logged steps. Step 3 is the
only failure: 4.9497V (limits 4.9500..5.0500) FAIL, with the operator’s note “low again on
fix3, thats 3 today”. Steps 4 through 8 pass, most with room to spare.
Read the failure guide. read_doc("failure-analysis-guide#3") matches the symptom, output low
but alive, and its routing is explicit: “Check the fixture before the board”. That points at
read_doc("failure-analysis-guide#7"), the section that separates a stale channel offset from an
open sense connection by which steps move: one step for an offset, three for an open sense. Here,
one step moved.
Take a confirmation reading, two ways. measure("dmm", "MEAS:VOLT:DC?") reads the output
through the fixture’s own channel: +4.963000E+00, 4.963 V. measure("load", "MEAS:VOLT?")
reads the same node at the electronic load’s own terminals, bypassing that channel entirely:
4.9930V. The two readings of one node disagree by about 30 mV, the size of FIX-03’s own
recorded offset, which is the signature section 7 describes: the channel, not the board.
A command that does not get through. The next call is measure("dmm", "*RST"), an attempt to
clear the meter before trusting it further. *RST sets state (it drops a range and, on a powered
instrument, more than that), so it is refused before Multimeter.send is ever called:
refused: ‘*RST’ is not a read-only command; this tool can only query dmm, never set it
Nothing about the meter or the board changes. The loop continues with the same DMM state it had.
Stop. With no more tool calls, the model states its answer: the cause is FIX-03’s channel 2
offset, not the board, and the next measurement is a person’s, not the agent’s: verify FIX-03
channel 2 against calibration-procedure.md section 5, then retest this serial on another
fixture per failure-analysis-guide.md section 1’s routing rule.
Nothing in that answer is a measurement. The two readings came from instruments, the 30 mV between them is a subtraction, and the cause is a hypothesis a person confirms: a model never produces a reported measurement, an uncertainty, a margin or a verdict. What it produced here is an order of questions.
What it costs
Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.
Every step above is decided_by: "model" or decided_by: "code" exactly as
examples/common/trace.py requires; the cost figures come from
tests/test_example_bench_bring_up_debug_assistant.py’s scripted walkthrough on StubModel,
which counts tokens deterministically and stands in for what a real model would be asked, not for
what one would say. The unit is the board somebody is already standing over, not the board tested:
no board on this bench gets an agent call by default.
How it fails on a real bench, specifically
A query's own argument changes the instrument
- How to notice it
- MEAS:VOLT:DC? 0.1 has a read-only header, but the DMM uses that argument to set its DC range before taking the reading. Send it once and every later query on that meter, argument or not, comes back through the same wrong range, which on this board reads as an overload with an empty error queue. A header is not enough to judge a command by.
- How to test for it
- Send a read-only header with an argument it does not need and confirm the range, or any other instrument state, is unchanged afterward. is_read_only draws the line where the manuals do: a read-only header carrying an argument is read-only on MEAS:VPP? alone, whose argument names a channel and sets nothing, and the tool above refuses everything else before the instrument sees it.
A command drafted for the wrong vendor
- How to notice it
- MEASure:CURRent? and MEAS:CURR? are the same header, and both exist on the supply and the load. The long form answers on the Maridun supply and comes back empty from the Tarnley load, with -113,"Undefined header" the only sign, because Tarnley firmware accepts the short form only.
- How to test for it
- Send the same header, long form, to an instrument from each vendor, and confirm SYST:ERR? shows the rejection rather than treating an empty response as a zero reading.
A note on the wrong row
- How to notice it
- Story 6 on this bench includes a note reading "ripple 44mv, passed but marginal" sitting on a VOUT row that passed. A cause taken from the nearest note instead of from the step it is actually attached to points at the wrong measurement.
- How to test for it
- Feed the assistant a log where a note and its step disagree, and confirm the stated cause follows the numeric result and the note is treated as a hint to check, never as the record.
How to evaluate it
There is no shared question set for this recipe the way the 60-question document set covers the
document-QA recipes: a bring-up symptom does not have one right sentence, it has a right cause and
a right next measurement. A reader building this for real would collect a set of logged
failures with a known, confirmed cause (the kind failure-analysis-guide.md already tracks one
board at a time) and score two things separately: whether the stated cause matches the confirmed
one, and whether the proposed next measurement is one that would actually distinguish it from the
next most likely cause. A stale channel offset and an open sense connection both read low on one
fixture and normal elsewhere; only the number of steps that moved tells them apart, so a cause
that is merely right and a next measurement that would not have caught the other explanation
score differently.
The safety side is not a judgment call and does not need a model to grade it: attack the measure
tool the way tests/test_example_bench_bring_up_debug_assistant.py does, with every command that
sets a value, enables an output, resets an instrument, stacks two commands in one message, or
hides a side effect behind an argument, and require every one of them refused before it reaches an
instrument. A single command that gets through is a failed build, not a score.
How to adapt it
The three tools are the general shape: something that reads a record of what already happened,
something that reads a reference document, and something that reads live state, all read-only,
all reachable without asking a person, with a code-side check between the model and anything that
is not. That shape fits any job where an agent investigates a system it must not be allowed to
change: a production incident with read-only access to logs, metrics and runbooks; a database
migration audit that may only SELECT; a security review that queries a fleet without holding any
credential that can reconfigure it. Single agent is
the loop, function calling is the tool shape,
safety is the code-side check that makes “read only” a
property of the code and not a promise in the system prompt, and
human-in-the-loop is where the recommendation
this agent produces has to go next, since nothing here can act on its own answer.
Every instrument on this bench is one send(command: str) -> str method, which is the shape
PyVISA’s own write and query pair covers, so the code that composes, checks and refuses a
command does not change when the transport does. Nothing in this repository has been run against
real hardware, and this page does not claim it has.
What does not port: the SRB-5030 itself, its command set, SafetyEnvelope’s 32.0 V and 4.0 A, and
every reading in docs/THE-BENCH.md. A reader’s own board has its own datasheet and its own
envelope, and the fixture-versus-board reasoning above is only as good as the documents an
assistant is actually pointed at.
Design choices
Why this level, and when to use another approach
This is level 5, an agent in a loop, because which measurement to take next depends on what the last one said, and a fixed order gets that wrong in both directions. Compare two boards that failed the same step for different reasons.
SRB5030-2608-0011’s log has one failure. Steps 4 and 5 (line and load regulation), 6 (efficiency)
and 7 (ripple) all pass comfortably. That pattern, one absolute reading off and everything
downstream of it normal, is what failure-analysis-guide.md section 7 calls the signature of a
stale channel offset, not a board fault: “Only measurements that read an absolute value through
the affected channel move, because a fixed offset cancels in the difference the regulation steps
take”.
SRB5030-2608-0063 also failed VOUT on FIX-03, at -0.0300 V, logged with the note “dead. no vout at all, u1 not switching”. Its efficiency step reads 0.00 percent against an 88 percent floor, and its current-limit step trips at 3.50 A, the lowest current tried. That board needs none of the fixture-versus-board reasoning SRB5030-2608-0011 does: the first tool call already answers it.
A level 3 workflow, code deciding a fixed order in advance, could certainly encode a script: read the log, then always compare the DMM against the load’s own terminal reading, then always check the calibration procedure. That order is right for 0011 and wastes two calls on 0063, where the two readings would simply agree near zero and the log already had the answer. Reverse it, straight from the log to failure analysis, and it scraps a good board to fix nothing. Which order is right depends on the shape of the first answer, which is what a level 5 loop is for and a checklist is not.
The next level up, a second agent reviewing the first one’s cause before it reaches a person, would catch a wrong read the first agent stated with confidence. It costs another agent’s worth of calls on every board investigated, not just the ones where the first agent was wrong, and it duplicates a check a person already makes here: the proposed cause and the next measurement go to an engineer before anything about the board changes. That trade is worth it when a wrong diagnosis is expensive relative to a person’s five minutes reading a trace; on this bench, it is not, so this recipe stops at level 5.
Build it
Implementation details and code
run is the loop: offer the three tools, run whichever one the model calls, hand back the
result, and stop once the model calls none. Every tool call and the final stop are the model’s
own decision; running or refusing a tool is always code’s.
View code: run
def run(
symptom: str,
model: Model,
tracer: Tracer,
*,
corpus_dir: Path = BENCH_CORPUS_DIR,
log_path: Path = PRODUCTION_CSV,
bench: Bench | None = None,
max_steps: int = MAX_STEPS,
max_tokens: int = MAX_TOKENS,
) -> Answer:
sections = load_sections(corpus_dir)
bench = bench if bench is not None else _bring_up(FIXTURE_DMM_OFFSET_V)
tracer.record(
kind="code",
decided_by="code",
title="Board already energized under an approved set point",
detail=f"{BOARD_VIN_V} V in, {BOARD_IOUT_A} A out; the agent's tools cannot reach this step",
)
messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=symptom)]
citations: list[str] = []
tokens_used = 0
for _ in range(max_steps):
completion = model.complete(messages, tools=TOOLS, max_tokens=400)
tokens_used += completion.tokens_in + completion.tokens_out
if not completion.tool_calls:
tracer.record(
kind="model",
decided_by="model",
title="Model states a cause and the next measurement",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
return Answer(text=completion.text, citations=sorted(set(citations)))
calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
tracer.record(
kind="model",
decided_by="model",
title="Model calls a tool",
detail=calls_desc,
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
turn, calls = assistant_turn(completion, len(messages))
messages.append(turn)
for call in calls:
result_text, cites = _run_tool(call, log_path, sections, bench)
citations.extend(cites)
refused = result_text.startswith("refused:")
title = "Refuse a command that sets state" if refused else f"Run tool: {call.name}"
tracer.record(kind="code", decided_by="code", title=title, detail=result_text[:200])
messages.append(tool_result(call, result_text))
if tokens_used >= max_tokens:
final = force_final(
messages, model, tracer, reason=f"token budget reached: {tokens_used} >= {max_tokens}", max_tokens=400
)
return Answer(text=final.text, citations=sorted(set(citations)))
final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
return Answer(text=final.text, citations=sorted(set(citations)))measure is where a state-setting command stops. is_read_only, imported from
examples/common/bench.py, is checked before instrument.send is ever called, and its result
decides whether that call happens at all, not whether the model asked politely. Sending a command
never touched by the check would be the actual bug; this function makes sure that path does not
exist. It is the only check here, deliberately: the line between a query and a state change is
written down once, in the bench module every example on this site imports, and a second copy here
would be a second copy to keep in step.
View code: measure
def _measure(instrument_name: str, command: str, bench: Bench) -> ToolResult:
"""Send one command, if and only if `is_read_only` says it only reads.
`is_read_only` is the whole check, deliberately: the line between a query and a state change
is written down once in `examples/common/bench.py` for every example on this site, and a
second copy of it here would be a second copy to keep in step. It already covers the case
that looks like a query and is not. `MEAS:VOLT:DC? 0.1` has a read-only header, but the
multimeter uses that argument to set its DC range before it reads, and the range stays set
for every later query with nothing in the error queue to say so, so an argument is read-only
on exactly one header: the oscilloscope's `MEAS:VPP? CHAN1`, whose argument only says which
channel to report.
"""
instruments = {"supply": bench.supply, "dmm": bench.dmm, "load": bench.load, "scope": bench.scope}
instrument = instruments.get(instrument_name)
if instrument is None:
return f"unknown instrument: {instrument_name!r}", []
if not is_read_only(command):
# The check that matters most: nothing past this line runs when it fails.
# `instrument.send` is never called, so a command that would set state or enable an
# output cannot reach the instrument through this tool no matter what the model asked
# for.
return (
f"refused: {command!r} is not a read-only command; this tool can only query "
f"{instrument_name}, never set it"
), []
return instrument.send(command), []Techniques this recipe uses
The highest level it needs is level 5.
Carry out a multi-step task in software, where the steps depend on what it finds
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- Fix a bug or add a feature in a repository
- Work a bring-up problem with read-only queries to instruments, the log and the datasheet
- Reproduce somebody else's measurement from their notebook and say where the two differ
- Reconcile two systems when finding the matching record is itself the work, rather than a field-by-field comparison
- Migrate configuration from one format to another
- Reproduce a reported defect
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page