Recipe

Draft an instrument control script from its programming manual

The model drafts commands from the manual for that instrument; code checks every one against the documented command set, runs the script on the simulated instrument, and feeds the errors back for another pass. A person bench-checks before it drives real hardware, and every set point goes through a code-side envelope.

SourcedNeeds level 4

A test engineer bringing up step 3 of a production test spec needs a script for the electronic load: set it to constant current, enable it, read the board back, disable it. The commands are in the load’s own programming manual, a PDF of tables and a worked example, and writing them by hand means paging through it for every keyword, every unit, every argument order. That’s the job: turn a manual into a script, for one instrument, without hand-typing every line.

The walkthrough below is a production test step, because that is the script this bench’s manual prints a worked example for. The same loop is worth more at low volume: a one-off characterization script runs five times, has no golden run to be checked against, and costs a person their afternoon if it is wrong. Volume changes the economics, not the level and not one of the four steps below.

The trap is that this bench has two instrument vendors whose SCPI dialects disagree. Maridun Instruments’ supply and meter accept VOLTage or VOLT and ON or 1, and answer a query with a bare number. Tarnley Test Systems’ load and scope take the short keyword and 1/0 only, and answer with a unit stuck to the number: 1.0000A. An engineer who has just written the supply’s script types CURRent and INP ON out of habit, and the load takes neither.

Walking a draft through the bench

No recorded trace exists for this page; docs/EVALS.md explains why. The numbers below come from running the example on StubModel, scripted to make exactly the mistake a Maridun-trained habit makes, against run.TASK: set the load to 1.000 A, enable it, read it back, disable it.

Draft 1, sent to a scratch TRN-2400 with no board wired to it at all:

MODE CC CURRent 1.000 INP ON MEAS:VOLT? MEAS:CURR? INP OFF

CURRent is the Maridun long form; Tarnley firmware takes the short form only and answers with -113,"Undefined header". INP ON and INP OFF are the Maridun boolean words; Tarnley wants 1 and 0 and answers both with -224,"Illegal parameter value". Both come back only when the script asks SYST:ERR?, not as anything that looks like a crash.

Draft 2, after the two exact SYST:ERR? lines go back to the model and nothing else:

MODE CC CURR 1.000 INP 1 MEAS:VOLT? MEAS:CURR? INP 0

Clean. Now, and only now, code runs it for real: the supply brings the board to 24.000 V with a 4.000 A current limit approved by the test engineer, CURR 1.000 is checked against SafetyEnvelope before it is sent, INP 1 needs its own matching approval, and the load reads back 4.9930V and 1.0000A. Those are not invented numbers: they are what the DUT model in examples/common/bench.py computes at 24 V in, 1.000 A out, and they match the load manual’s own worked example exactly, because both describe the same board. Neither went through the model. It drafted the commands that took them, and that is where its output stops: a model never produces a reported measurement, an uncertainty, a margin or a verdict.

What it costs

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls, clean draft
2Model calls, one dialect mistake
~760Tokens, clean draft
~1,575Tokens, one revision

Those are the illustrative run’s own numbers, counted by examples/common/model.py’s count_tokens estimate rather than a provider’s real tokenizer. It is a per-script cost in both settings, and the reader’s own unit decides which line matters.

Per unit, in production: zero. Once a draft runs clean, running it against the next board and all 50 a day costs no model call at all, so the model is paid for once and the comparison is against the seconds the test step itself takes.

Per session, at low volume: those same 1,575 tokens are the whole model cost of the afternoon, because a characterization script runs five times and there is nothing to amortize over. The comparison is against the hour of manual-paging the draft replaced, and against the afternoon lost if the draft was wrong. That second cost is larger here, not smaller: in production a wrong keyword is caught by the golden run, and at five runs there is no golden run, so step 2 above is the only thing between a Maridun habit and a wasted afternoon.

How this fails on a real bench

A command drafted from the wrong vendor's dialect

How to notice it
The step the manual's worked example shows a response for produces nothing, and the next SYST:ERR? holds -113 or -224 instead of 0, not an exception and not a crash.
How to test for it
Send every drafted command to the target instrument's own simulated class and read SYST:ERR? after each one, the way _check_against_manual does, before a person ever sees the script.

A unit suffix parsed with a bare float()

How to notice it
A Maridun reply is a bare number and float() works; the identical call on a Tarnley reply like "1.0000A" raises, which is the good case, or a fixed-width slice returns a wrong number silently, which is worse.
How to test for it
Call the parser on both a suffixed and a bare reply and check the number, not just that it runs; parse_reading's own tests do this for "4.9930V", "1000.0000OHM" and "24.0000".

A reading taken before the board has settled

How to notice it
Step 4 of the manual's own load sequence is "Wait 100 ms for the board and the load to settle". A drafted script that goes straight from INP 1 to MEAS:VOLT? still runs clean, because a simulated instrument settles instantly and the error queue has nothing to say about timing. The reading is a number, not a measurement.
How to test for it
Nothing in the draft-check-revise loop catches this; the simulator is the wrong instrument to ask. It is what the fourth step is for: a person comparing the first reading after the enable with one taken a second later on the real load.

An enable command with no approval, or the wrong one

How to notice it
Nothing on a simulator catches fire, which is exactly the danger: a script that reaches INP 1 with no Approval, or one naming a different set point than the board is actually at, has to be refused before it is sent, not after.
How to test for it
Call the enable path with no Approval, and with one naming the wrong voltage or current, and require SafetyRefusal both times.

How to evaluate it

Whether a drafted script is acceptable is a pass or fail a simulator already computes, never a model’s opinion of its own work: did it run clean within the revision cap, and once run for real, did the reading land where the DUT model says it should. In production there is a set to collect: one drafting task per test step per instrument, eight steps across four manuals, each with a known-good script as the answer key, scored on how many converge within the cap and how many revisions each took. A script that never converges is not a partial credit case; it is the “Stop: revision cap reached” step in the trace, and it means a person looks at the manual next, not the model again.

At low volume that set does not exist: one script, no answer key. The check to do first costs nothing and is mechanical. Every command in the final script appears in the instrument’s own command table, and every reading lands where the datasheet says. A script that runs clean and measures the wrong node is the failure no simulator catches, in either setting.

Adapting it to your own instrument

Every instrument on this bench is one send(command: str) -> str method, the shape PyVISA’s write and query pair covers, so the command-checking and revision-loop code above does not change; only the transport underneath send does. Nothing here has been run against real hardware, and this page does not claim it has.

What does not port: the TRN-2400’s own command table, its two dialect quirks, and the DUT model that makes 4.9930V the right answer at 24 V and 1.000 A. Your instrument has its own manual and your board has its own datasheet, and checking a drafted command against those, not against this one, is the entire lesson.

The shape is draft, check, revise against a fixed rule, and it shows up anywhere a draft has to meet a standard that code, not a person’s read of the draft, can check: code checked against its own tests, a SQL query checked against the schema it queries, a report checked against a required template. What is unusual about an instrument script is only that a rejected command can reach a mains-powered board, which is why the fourth step, here, is a person and not another check.

Design choices

Why this level, and when to use another approach

Level 0 already runs the checked script: once it is known good, replaying it costs nothing and needs no model at any level, whether against 50 boards a day or across the corners of one prototype, and so does everything done with the readings it takes. The job a model helps with is the one-time draft, when the script does not exist yet or the manual has changed, and a first draft off two disagreeing manuals is where a wrong keyword or a wrong boolean word actually gets typed.

A single, ungraded draft (level 1) is not enough, because “does this run clean” is not a matter of taste, it’s a fact the simulated instrument already knows: send the command, read SYST:ERR?. Throwing that answer away wastes it. Level 3 puts that check in a loop: draft, check every command against the instrument’s own documented set, and if anything was rejected, hand back exactly what SYST:ERR? said and draft again, capped at a fixed number of tries. Climbing past level 3, to a model that also decides when to stop or reads back its own success, buys nothing here: whether a script is clean is binary and already computed by the simulator, and a model deciding when “clean enough” has been reached would be grading the one thing code grades for free.

Four things stay true whether the load on the other end is this simulator or a real TRN-2400 on a bench, because the same code runs either way:

  1. The model drafts. It proposes a list of SCPI commands. It never sends one.
  2. Code checks every command against the documented set, by sending each one to a simulated instrument that implements exactly what the manual documents and nothing else, and reading SYST:ERR? after every line.
  3. The script runs on the simulated instrument first, and whatever SYST:ERR? found goes back to the model for another draft, until it runs clean or a fixed number of tries runs out.
  4. A person bench-checks it before it ever points at a real load, with the board’s own current limit set low and somebody watching.

Nothing here has run against real hardware. Even the “for real” step below still means the simulator: it is the step the same code takes when the target is not simulated.

Build it

Implementation details and code

Checking a command means sending it to an instrument that only knows what its manual documents:

View code: check against manual
examples/bench_instrument_script_from_the_manual/run.py · lines 147–166
def _check_against_manual(commands: list[str], tracer: Tracer) -> list[tuple[str, str]]:
    """Try a draft against a scratch TRN-2400: unwired, no board, safe to send anything.

    This is the documented command set for the instrument, not a copy of it: `ElectronicLoad`
    accepts exactly what `trn2400-programming-manual.md` documents and rejects everything else,
    so a command the manual does not support fails here the same way it would on the real load.
    """
    scratch = ElectronicLoad()
    errors: list[tuple[str, str]] = []
    for command in commands:
        scratch.send(command)
        error = scratch.send("SYST:ERR?")
        if error != NO_ERROR:
            errors.append((command, error))
    detail = "clean" if not errors else "; ".join(f"{c!r} -> {e}" for c, e in errors)
    tracer.record(
        kind="code", decided_by="code",
        title="Check commands against the documented command set", detail=detail,
    )
    return errors

Enabling the load is the one line in the whole script that actually puts current through the board, so it is the one line the model never gets to send outright:

View code: enable load
examples/bench_instrument_script_from_the_manual/run.py · lines 169–180
def _enable_load(load: GuardedLoad, approval: Approval) -> None:
    """Enable the TRN-2400's input: the one command in this script that puts current through the
    board, and so the one command a person has to have approved.

    The gate itself is `GuardedLoad.input_on` in `examples/common/bench.py`, next to the identical
    one `GuardedSupply.output_on` puts in front of `OUTP ON`. It refuses an enable with no
    `Approval`, one that names a rail or a current the bench is not actually at, and one that has
    already been spent, and it re-checks the load's own set point against `SafetyEnvelope` on the
    way through. This recipe adds nothing of its own to that; it names the step, because a reader
    following the drafted script needs to see where the model's line stops being the model's.
    """
    load.input_on(approval)

That is a call and not a check, on purpose. The gate lives in examples/common/bench.py as GuardedLoad.input_on, beside the one GuardedSupply.output_on puts in front of OUTP ON, so every recipe that enables a load gets it from one place. It is gated because the TRN-2400 will sink 30 A into a board rated for 3.0 A, and because code cannot tell a deliberate limit hunt from a set point nobody meant. A person names the rail and the current; code checks the bench is actually at them.

And reading a Tarnley reply back:

View code: parse reading
examples/bench_instrument_script_from_the_manual/run.py · lines 67–77
def parse_reading(reply: str) -> float:
    """Parse one numeric reply, Maridun's bare or Tarnley's suffixed (`"1.0000A"` -> `1.0000`).

    `float()` is tried first, which is the correct parse for a Maridun reply and the good failure
    for a Tarnley one: it raises rather than silently returning a wrong number, which is what
    slicing a fixed number of characters off the reply would do instead.
    """
    try:
        return float(reply)
    except ValueError:
        return float(_UNIT_SUFFIX_RE.sub("", reply))
Composition

Techniques this recipe uses

The highest level it needs is level 4.

Retrieval-augmented generation (RAG)

Measured

Searching your documents and giving the results to the model.

Write and check

Sourced

One prompt writes, another checks, and the loop repeats until the check passes.

Code execution

Sourced

Letting the model write code and run it in a sandbox.

Same shape, other jobs

Produce something that has to meet a standard, and check it before anyone sees it

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • Marketing copy against brand and legal rules
  • An instrument control script checked against the documented command set and run on a simulator
  • A measurement report checked figure by figure against the numbers code computed
  • Code against its tests
  • A SQL query against the schema
  • A report against a required template
  • A test procedure against the requirement it verifies

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page