# Draft an instrument control script from its programming manual

_Recipe · needs level 4_

The model drafts commands from the manual for that instrument; code checks every one against the documented command set, runs the script on the simulated instrument, and feeds the errors back for another pass. A person bench-checks before it drives real hardware, and every set point goes through a code-side envelope.


A test engineer bringing up step 3 of a production test spec needs a script for the electronic
load: set it to constant current, enable it, read the board back, disable it. The commands are in
the load's own programming manual, a PDF of tables and a worked example, and writing them by hand
means paging through it for every keyword, every unit, every argument order. That's the job: turn
a manual into a script, for one instrument, without hand-typing every line.

The walkthrough below is a production test step, because that is the script this bench's manual
prints a worked example for. The same loop is worth more at low volume: a one-off characterization
script runs five times, has no golden run to be checked against, and costs a person their
afternoon if it is wrong. Volume changes the economics, not the level and not one of the four
steps below.

The trap is that this bench has two instrument vendors whose SCPI dialects disagree. Maridun
Instruments' supply and meter accept `VOLTage` or `VOLT` and `ON` or `1`, and answer a query with a
bare number. Tarnley Test Systems' load and scope take the short keyword and `1`/`0` only, and
answer with a unit stuck to the number: `1.0000A`. An engineer who has just written the supply's
script types `CURRent` and `INP ON` out of habit, and the load takes neither.

## Walking a draft through the bench

No recorded trace exists for this page; `docs/EVALS.md` explains why. The numbers below come from
running the example on `StubModel`, scripted to make exactly the mistake a Maridun-trained habit
makes, against `run.TASK`: set the load to 1.000 A, enable it, read it back, disable it.

**Draft 1**, sent to a scratch TRN-2400 with no board wired to it at all:

    MODE CC
    CURRent 1.000
    INP ON
    MEAS:VOLT?
    MEAS:CURR?
    INP OFF

`CURRent` is the Maridun long form; Tarnley firmware takes the short form only and answers with
`-113,"Undefined header"`. `INP ON` and `INP OFF` are the Maridun boolean words; Tarnley wants `1`
and `0` and answers both with `-224,"Illegal parameter value"`. Both come back only when the
script asks `SYST:ERR?`, not as anything that looks like a crash.

**Draft 2**, after the two exact `SYST:ERR?` lines go back to the model and nothing else:

    MODE CC
    CURR 1.000
    INP 1
    MEAS:VOLT?
    MEAS:CURR?
    INP 0

Clean. Now, and only now, code runs it for real: the supply brings the board to 24.000 V with a
4.000 A current limit approved by the test engineer, `CURR 1.000` is checked against `SafetyEnvelope`
before it is sent, `INP 1` needs its own matching approval, and the load reads back `4.9930V` and
`1.0000A`. Those are not invented numbers: they are what the DUT model in
`examples/common/bench.py` computes at 24 V in, 1.000 A out, and they match the load manual's own
worked example exactly, because both describe the same board. Neither went through the model. It
drafted the commands that took them, and that is where its output stops: a model never produces a
reported measurement, an uncertainty, a margin or a verdict.

## What it costs

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, clean draft:** 1
- **Model calls, one dialect mistake:** 2
- **Tokens, clean draft:** ~760
- **Tokens, one revision:** ~1,575

Those are the illustrative run's own numbers, counted by `examples/common/model.py`'s
`count_tokens` estimate rather than a provider's real tokenizer. It is a per-script cost in both
settings, and the reader's own unit decides which line matters.

**Per unit, in production:** zero. Once a draft runs clean, running it against the next board and
all 50 a day costs no model call at all, so the model is paid for once and the comparison is
against the seconds the test step itself takes.

**Per session, at low volume:** those same 1,575 tokens are the whole model cost of the afternoon,
because a characterization script runs five times and there is nothing to amortize over. The
comparison is against the hour of manual-paging the draft replaced, and against the afternoon lost
if the draft was wrong. That second cost is larger here, not smaller: in production a wrong keyword
is caught by the golden run, and at five runs there is no golden run, so step 2 above is the only
thing between a Maridun habit and a wasted afternoon.

## How this fails on a real bench

### A command drafted from the wrong vendor's dialect

- **How to notice it:** The step the manual's worked example shows a response for produces nothing, and the next SYST:ERR? holds -113 or -224 instead of 0, not an exception and not a crash.
- **How to test for it:** Send every drafted command to the target instrument's own simulated class and read SYST:ERR? after each one, the way _check_against_manual does, before a person ever sees the script.

### A unit suffix parsed with a bare float()

- **How to notice it:** A Maridun reply is a bare number and float() works; the identical call on a Tarnley reply like "1.0000A" raises, which is the good case, or a fixed-width slice returns a wrong number silently, which is worse.
- **How to test for it:** Call the parser on both a suffixed and a bare reply and check the number, not just that it runs; parse_reading's own tests do this for "4.9930V", "1000.0000OHM" and "24.0000".

### A reading taken before the board has settled

- **How to notice it:** Step 4 of the manual's own load sequence is "Wait 100 ms for the board and the load to settle". A drafted script that goes straight from INP 1 to MEAS:VOLT? still runs clean, because a simulated instrument settles instantly and the error queue has nothing to say about timing. The reading is a number, not a measurement.
- **How to test for it:** Nothing in the draft-check-revise loop catches this; the simulator is the wrong instrument to ask. It is what the fourth step is for: a person comparing the first reading after the enable with one taken a second later on the real load.

### An enable command with no approval, or the wrong one

- **How to notice it:** Nothing on a simulator catches fire, which is exactly the danger: a script that reaches INP 1 with no Approval, or one naming a different set point than the board is actually at, has to be refused before it is sent, not after.
- **How to test for it:** Call the enable path with no Approval, and with one naming the wrong voltage or current, and require SafetyRefusal both times.

## How to evaluate it

Whether a drafted script is acceptable is a pass or fail a simulator already computes, never a
model's opinion of its own work: did it run clean within the revision cap, and once run for real,
did the reading land where the DUT model says it should. In production there is a set to collect:
one drafting task per test step per instrument, eight steps across four manuals, each with a
known-good script as the answer key, scored on how many converge within the cap and how many
revisions each took. A script that never converges is not a partial credit case; it is the "Stop:
revision cap reached" step in the trace, and it means a person looks at the manual next, not the
model again.

At low volume that set does not exist: one script, no answer key. The check to do first costs
nothing and is mechanical. Every command in the final script appears in the instrument's own
command table, and every reading lands where the datasheet says. A script that runs clean and
measures the wrong node is the failure no simulator catches, in either setting.

## Adapting it to your own instrument

Every instrument on this bench is one `send(command: str) -> str` method, the shape PyVISA's
`write` and `query` pair covers, so the command-checking and revision-loop code above does not
change; only the transport underneath `send` does. Nothing here has been run against real
hardware, and this page does not claim it has.

What does not port: the TRN-2400's own command table, its two dialect quirks, and the DUT model
that makes `4.9930V` the right answer at 24 V and 1.000 A. Your instrument has its own manual and
your board has its own datasheet, and checking a drafted command against those, not against this
one, is the entire lesson.

The shape is
[draft, check, revise against a fixed rule](/gradient_ascent/techniques/evaluator-optimizer/),
and it shows up anywhere a draft has to meet a standard that code, not a person's read of the
draft, can check: code checked against its own tests, a SQL query checked against the schema it
queries, a report checked against a required template. What is unusual about an instrument script
is only that a rejected command can reach a mains-powered board, which is why the fourth step,
here, is a person and not another check.
## Design choices

### Why this level, and when to use another approach

Level 0 already runs the checked script: once it is known good, replaying it costs nothing and
needs no model at any level, whether against 50 boards a day or across the corners of one
prototype, and so does [everything done with the
readings it takes](/gradient_ascent/recipes/limits-without-a-model/). The job a model helps with is the one-time draft, when the script does
not exist yet or the manual has changed, and a first draft off two disagreeing manuals is where a
wrong keyword or a wrong boolean word actually gets typed.

A single, ungraded draft (level 1) is not enough, because "does this run clean" is not a matter of
taste, it's a fact the simulated instrument already knows: send the command, read `SYST:ERR?`.
Throwing that answer away wastes it. Level 3 puts that check in a loop: draft, check every command
against the instrument's own documented set, and if anything was rejected, hand back exactly what
`SYST:ERR?` said and draft again, capped at a fixed number of tries. Climbing past level 3, to a
model that also decides *when* to stop or reads back its own success, buys nothing here: whether a
script is clean is binary and already computed by the simulator, and a model deciding when "clean
enough" has been reached would be grading the one thing code grades for free.

Four things stay true whether the load on the other end is this simulator or a real TRN-2400 on a
bench, because the same code runs either way:

1. **The model drafts.** It proposes a list of SCPI commands. It never sends one.
2. **Code checks every command against the documented set**, by sending each one to a simulated
   instrument that implements exactly what the manual documents and nothing else, and reading
   `SYST:ERR?` after every line.
3. **The script runs on the simulated instrument first**, and whatever `SYST:ERR?` found goes back
   to the model for another draft, until it runs clean or a fixed number of tries runs out.
4. **A person bench-checks it** before it ever points at a real load, with the board's own current
   limit set low and somebody watching.

Nothing here has run against real hardware. Even the "for real" step below still means the
simulator: it is the step the same code takes when the target is not simulated.

## Build it

### Implementation details and code

Checking a command means sending it to an instrument that only knows what its manual documents:

`examples/bench_instrument_script_from_the_manual/run.py` (lines 147-166)

```python
def _check_against_manual(commands: list[str], tracer: Tracer) -> list[tuple[str, str]]:
    """Try a draft against a scratch TRN-2400: unwired, no board, safe to send anything.

    This is the documented command set for the instrument, not a copy of it: `ElectronicLoad`
    accepts exactly what `trn2400-programming-manual.md` documents and rejects everything else,
    so a command the manual does not support fails here the same way it would on the real load.
    """
    scratch = ElectronicLoad()
    errors: list[tuple[str, str]] = []
    for command in commands:
        scratch.send(command)
        error = scratch.send("SYST:ERR?")
        if error != NO_ERROR:
            errors.append((command, error))
    detail = "clean" if not errors else "; ".join(f"{c!r} -> {e}" for c, e in errors)
    tracer.record(
        kind="code", decided_by="code",
        title="Check commands against the documented command set", detail=detail,
    )
    return errors
```

Enabling the load is the one line in the whole script that actually puts current through the
board, so it is the one line the model never gets to send outright:

`examples/bench_instrument_script_from_the_manual/run.py` (lines 169-180)

```python
def _enable_load(load: GuardedLoad, approval: Approval) -> None:
    """Enable the TRN-2400's input: the one command in this script that puts current through the
    board, and so the one command a person has to have approved.

    The gate itself is `GuardedLoad.input_on` in `examples/common/bench.py`, next to the identical
    one `GuardedSupply.output_on` puts in front of `OUTP ON`. It refuses an enable with no
    `Approval`, one that names a rail or a current the bench is not actually at, and one that has
    already been spent, and it re-checks the load's own set point against `SafetyEnvelope` on the
    way through. This recipe adds nothing of its own to that; it names the step, because a reader
    following the drafted script needs to see where the model's line stops being the model's.
    """
    load.input_on(approval)
```

That is a call and not a check, on purpose. The gate lives in `examples/common/bench.py` as
`GuardedLoad.input_on`, beside the one `GuardedSupply.output_on` puts in front of `OUTP ON`, so
every recipe that enables a load gets it from one place. It is gated because the TRN-2400 will
sink 30 A into a board rated for 3.0 A, and because code cannot tell a deliberate limit hunt from
a set point nobody meant. A person names the rail and the current; code checks the bench is
actually at them.

And reading a Tarnley reply back:

`examples/bench_instrument_script_from_the_manual/run.py` (lines 67-77)

```python
def parse_reading(reply: str) -> float:
    """Parse one numeric reply, Maridun's bare or Tarnley's suffixed (`"1.0000A"` -> `1.0000`).

    `float()` is tried first, which is the correct parse for a Maridun reply and the good failure
    for a Tarnley one: it raises rather than silently returning a wrong number, which is what
    slicing a fixed number of characters off the reply would do instead.
    """
    try:
        return float(reply)
    except ValueError:
        return float(_UNIT_SUFFIX_RE.sub("", reply))
```



Last reviewed 2026-09-19.
