# Ask questions of a production test log

_Recipe · needs level 4_

Starts where the dashboard stopped: limits, yield and Cpk are already charted and did not answer the question. The model writes analysis code that runs in a sandbox over the CSV, and a person reads the code as well as the answer. Includes the trap of a column in millivolts under a header that says volts.


This recipe starts from the other end. By the time anyone asks it a question, the dashboard is
already built: [limits, first-pass yield and Cpk](/gradient_ascent/recipes/limits-without-a-model/)
are charted from the production log, grouped by lot, by fixture, by day and by shift. The question
that shows up next is never one of those charts. An engineer reads the August 31 retest export and
the numbers look wrong. A soak log from a 90-minute burn-in has three units in it and a person wants
to know, in words, whether any of them drifted. A September sweep of five prototypes has 900
readings in it and nobody has asked whether any of them is impossible. Two of those three are
production test and the third is engineering test, which is the point: what makes a question this
recipe's is not the setting it came from, it is that nobody built a chart for it.

## The run on the bench, stepped

**"Do the retested boards actually pass?"** Eighteen boards were pulled back for a retest on August
31 into `retest-2026-08-31.csv`, whose `value_v` column is meant to be volts. Before any snippet
touches it, code checks every value against 40.0 V, the widest node an SRB-5030 has anywhere
(`VIN_ABS_MAX_V`, the datasheet's absolute maximum input). Sixteen of the eighteen, numbers like
4973.5 and 5007.7, clear that ceiling by two orders of magnitude, and dividing by 1000 brings all
sixteen back under it, so code applies that correction and says so every time the table loads,
whether or not the model calls the tool:

`examples/bench_test_data_by_conversation/run.py` (lines 79-108)

```python
def _as_plausible_volts(
    rows: list[dict], *, field: str = "value_v", ceiling_v: float = VIN_ABS_MAX_V
) -> tuple[list[dict], str]:
    """Range-check `field` against the widest node this board has anywhere, before any analysis
    sees it. Every reading in `field` is a volts measurement on an SRB-5030, and this board never
    carries more than `ceiling_v` volts on any pin (`docs/THE-BENCH.md`); a value past that by
    three orders of magnitude is not a surprising board, it is the wrong unit. If dividing by 1000
    brings every value back inside the ceiling, the column was millivolts and code corrects it and
    says so; if it still does not fit, this refuses rather than guess further.
    """

    def implausible(values: list[float]) -> list[float]:
        return [v for v in values if abs(v) > ceiling_v]

    raw = [row[field] for row in rows]
    bad = implausible(raw)
    if not bad:
        return rows, ""
    scaled = [{**row, field: row[field] / 1000.0} for row in rows]
    still_bad = implausible([row[field] for row in scaled])
    if still_bad:
        raise ImplausibleUnits(
            f"{len(bad)} of {len(rows)} {field!r} readings exceed {ceiling_v} V even after "
            f"dividing by 1000; refusing to guess the unit."
        )
    return scaled, (
        f"{len(bad)} of {len(rows)} {field!r} readings (as high as {max(bad):.1f}) exceeded "
        f"{ceiling_v} V, the widest node this board has anywhere. Code divided {field} by 1000 "
        f"before any analysis ran: the export is millivolts under a header that says volts."
    )
```

Read literally as volts, those sixteen readings clear a 4.9500 V lower limit by three orders of
magnitude, which is the trap: a check that only looks at the lower limit calls all of them passes.
Once the model asks for the snippet and code runs it on the corrected table, the real numbers come
back: fifteen of the eighteen sit between 4.9678 V and 5.0077 V, mean 4.9840 V; three do not
(SRB5030-2608-0052, SRB5030-2608-0063, SRB5030-2608-0178). The file's own `result` column already
had this right: whatever tool wrote the export judged pass and fail correctly and only printed the
wrong header.

**"Did any soak unit fail to settle?"** Three boards ran 90 minutes at full load on August 27, five
minutes between samples, in `soak-2026-08-27.csv`. Asked to look at the whole run rather than the
last row, the model's snippet takes the first and last sample per serial and reports the output
drop and the case-temperature rise. Two boards settle: 2.5 mV and 1.2 mV of drop against a case
that climbs to about 60 degC and stays there. The third, SRB5030-2608-0121, drops 75.5 mV while its
case keeps climbing to 91.7 degC and never levels off, which is what a resistive joint does:
current squared times resistance heats it, and a hotter joint is more resistive. `docs/THE-BENCH.md`
records that this same serial is the only genuine efficiency failure in the whole production run,
at 79.6%. Nothing in this snippet reads the production log or draws that connection; a person who
has read both tables is the one who gets to make it.

**"Is any block of the sweep impossible?"** Five prototypes were swept over line, load and
temperature in September into `characterization-2026-09.csv`: 900 readings, no limits column and
no verdict, because what an engineer computes from it is a margin. The same range check runs on
its `vout_v` column, finds every value plausible as volts, and says so: a guard that only speaks
when it fires cannot be told apart from one nobody wired up. Then the question, which no dashboard
has a tile for. A board cannot put out more power than it takes in, so the snippet averages each
of the 180 blocks and compares `vin_v * iin_a` against `iout_a * vout_v`. One block comes back:
SRB5030-2609-0001 at a labeled 12.0 V and 3.000 A draws 0.6729 A where that same point on the
other four boards draws 1.3123 A, which is 8.08 W in against 14.95 W out. The supply was still at
24.0 V from the block before it. The output voltage gives nothing away, since holding it steady
while the input moves is the whole job of the part, and no uncertainty budget would have caught it
either: every reading in that block is a good reading of a condition nobody asked for.

The same file's other trap is not this recipe's. A block of readings scatters seven times as wide
as the rest because the meter was left on the 100 V range, and finding that is a `GROUP BY` on a
column nobody thinks to group by, which is
[the engineering-test level-0 page](/gradient_ascent/recipes/characterize-a-design/)'s work. Ask
this recipe only what a grouping has already failed to answer.

## What it costs

The unit here is the question, in all three settings, because a question is what a person asks
once. Two model calls each: one to write the snippet, one to turn the result into a sentence.
Nothing on this site has called a live model, so this is an estimate, worked in the open with the
same deterministic token counter the stub runs use: the retest question above runs 617 input
tokens and 155 output tokens across its two calls, counting the system prompt describing all three
tables, the snippet, the sandbox's JSON result and the final sentence. A real model's tokenizer
will not match that exactly, but the shape holds whichever model runs it. This recipe is not
competing with the dashboard, which answers most questions
for nothing per question; it is competing with the ten minutes an ad hoc question costs a person
with a spreadsheet and a filter box.

## How it fails on a real bench, specifically

A retest export merged on a column named for volts and filled with millivolts is the whole first
story above, and what catches it is the range check, in code, before any snippet runs, not a model
noticing the numbers look large. That check is blunt on purpose and it can be wrong in one
direction: a column that really is volts with a single mistyped row in it would be scaled by a
thousand along with the bad row, and every value would still be inside the ceiling afterward. That
is why the correction goes into the trace on every run instead of being applied quietly. A snippet
that runs clean and computes the wrong thing is not caught by the sandbox at all; it is caught by a
person reading the code next to the number, which is why the code is on the page and not just the
answer. No number this recipe reports is the model's:
the sandbox computes every figure and the model writes the sentence around it, because a model
never produces a reported measurement, an uncertainty, a margin or a verdict. The pass or fail on
the retested boards already happened, in the file's own `result` column and the limits every
earlier step checked. The sweep's margins are not on this page at all, for the same reason: a
margin is a subtraction, and a subtraction belongs where it costs nothing per question.

## How to evaluate it

This example runs on the bench in `evals/bench/`, which has no question file of its own, so the
site's 60-question document set (`docs/EVALS.md`) has nothing to grade it against; `scripts/
eval_run.py` refuses to score it and says what to measure instead. Two things, both checkable by a
program rather than a person's judgment: whether the figures in the final answer equal the figures
the executed snippet actually printed, which the tests in this recipe's own example package check
against numbers recomputed independently from the CSVs, and whether the range check stops the
mislabeled `value_v` column before a snippet ever sees it. Beyond that, a real eval is a small set
of ad hoc questions over these three tables with answers worked out by hand, scored on whether the
snippet ran clean or the sandbox refused it, and whether the final number matches the hand-worked
one exactly. Track a sandbox refusal separately from a wrong-but-clean answer; they are different
failures with different fixes, one the sandbox doing its job and the other a prompt that needs to
say more about what the tables hold.

## How to adapt it

The shape ports past this bench whenever a table and a question already exist and a fixed report
cannot anticipate the question: a regional sales dip, survey responses, server logs after an
incident. What ports specifically: one tool, a grammar checked node by node before anything runs,
a range check on any column whose unit is a fact you already know, a physical identity the data
has to satisfy the way a power balance does, and a person reading the snippet next to the number.

What does not port: the 40 V ceiling, which is this board's absolute maximum input and nobody
else's; the column names; and the three tables themselves, which are `evals/bench/data/` CSVs
invented for this site. A reader's own data has its own plausible ranges, its own known units and
its own identities that have to hold, and the lesson of the first and third stories above is that
those are worth checking in code before a snippet runs, on every table, not just these three.
## Design choices

### Why this level, and when to use another approach

The model writes one short Python snippet, a sandbox runs it against the loaded tables, and the
model turns whatever came back into a sentence. That is level 4: one real decision (what the
snippet says), then code that always runs it and always asks for a final answer. It composes two
techniques: writing the snippet is [function calling](/gradient_ascent/techniques/function-calling/),
one tool, offered once; running it is [code execution](/gradient_ascent/techniques/code-execution/),
a sandbox that never trusts the model's text on its own.

It is not level 5. [Data analysis by conversation](/gradient_ascent/recipes/data-analysis/) is the
same shape with a loop around it, for a question whose second computation depends on what the first
one showed: "how does this quarter compare to last" cannot even name its second query until the
first one has answered. None of the three questions below works that way. Each takes one snippet
and one look at the result, and a loop would buy nothing here but a second, unnecessary model
call. If a reader's own questions turn out to chain, that recipe is where to go.

It is also not level 0, and that is worth saying plainly: a fixed report cannot answer a question
nobody wrote a query for yet. What is level 0 is the arithmetic inside the snippet, and the check
that runs before any snippet does. Both are ordinary code, and neither is the model's to get right
or wrong.

## Build it

### Implementation details and code

The model's one decision, and what code does with it, in `run`:

`examples/bench_test_data_by_conversation/run.py` (lines 368-388)

```python
    call = first.tool_calls[0]
    dropped = "" if len(first.tool_calls) == 1 else f" (dropped {len(first.tool_calls) - 1} further call(s))"
    code = str(call.arguments.get("code", ""))
    tracer.record(
        kind="model",
        decided_by="model",
        title="Model writes analysis code",
        detail=code[:300] + dropped,
        tokens_in=first.tokens_in,
        tokens_out=first.tokens_out,
        ms=first.ms,
    )

    try:
        result = run_snippet(code, tables)
    except UnsafeCode as exc:
        tracer.record(kind="code", decided_by="code", title="Sandbox refused the snippet", detail=str(exc))
        return Answer(text=f"Could not safely run that analysis: {exc}", citations=[])

    result_text = _render_result(result)
    tracer.record(kind="code", decided_by="code", title="Sandbox runs the snippet", detail=result_text[:400])
```

The sandbox itself never calls Python's own `eval`. It does call `exec`, unlike
[code execution](/gradient_ascent/techniques/code-execution/)'s own single-expression evaluator,
but only after every node in the snippet has been walked and refused if it is not on a short
allow-list: no import, no attribute access at all (so no `x.y`, and nothing dunder-chained off a
literal), no `lambda`, no `while`, no function or class definitions, calls only to a fixed list of
names, and loops nested no more than two deep:

`examples/bench_test_data_by_conversation/run.py` (lines 256-280)

```python
def run_snippet(code: str, tables: dict[str, list[dict]]) -> object:
    """Run one analysis snippet against `tables` and return whatever it assigned to `result`.

    Never calls Python's own `eval`. It does call `exec`, but only on a tree `_check_grammar` has
    already walked node by node, in a namespace with `__builtins__` emptied out, holding nothing
    but `TABLES` and the functions in `_SANDBOX_FUNCTIONS` -- so there is no name anywhere in
    scope that reaches a file, a socket, or the interpreter itself.
    """
    if len(code) > MAX_SOURCE_CHARS:
        raise UnsafeCode(f"snippet is {len(code)} characters; the limit is {MAX_SOURCE_CHARS}")
    try:
        tree = ast.parse(code, mode="exec")
    except SyntaxError as exc:
        raise UnsafeCode(f"not valid analysis code: {exc}") from exc
    _check_grammar(tree)
    namespace: dict[str, object] = {"__builtins__": {}, "TABLES": tables, **_SANDBOX_FUNCTIONS}
    try:
        exec(compile(tree, "<analysis-snippet>", "exec"), namespace)  # noqa: S102 -- see docstring
    except UnsafeCode:
        raise
    except Exception as exc:  # the snippet's own runtime error: a KeyError, a ZeroDivisionError...
        raise UnsafeCode(f"the snippet raised {type(exc).__name__}: {exc}") from exc
    if "result" not in namespace:
        raise UnsafeCode("the snippet must assign its answer to a variable named result")
    return namespace["result"]
```

What that grammar protects against: every classic escape this site's own tests try against it
(`(1).__class__`, `__import__('os')`, `open(...)`, an f-string, a walrus, a `while True`) is
refused the same way, because the node type it needs is not in the allow-list, not because the
code recognized an attack. What it does not protect against: a snippet that runs clean and answers
a question other than the one asked, which is why the sandbox's own output goes next to the prose
and not instead of it, and why the loop and node caps are a coarse defense against a runaway
snippet rather than a wall-clock timeout. A production deployment wants an OS-level sandbox around
this too, the same point [code execution](/gradient_ascent/techniques/code-execution/) makes about
a real container.



Last reviewed 2026-09-19.
