Recipe

Ask questions of a production test log

Starts where the dashboard stopped: limits, yield and Cpk are already charted and did not answer the question. The model writes analysis code that runs in a sandbox over the CSV, and a person reads the code as well as the answer. Includes the trap of a column in millivolts under a header that says volts.

SourcedNeeds level 4

This recipe starts from the other end. By the time anyone asks it a question, the dashboard is already built: limits, first-pass yield and Cpk are charted from the production log, grouped by lot, by fixture, by day and by shift. The question that shows up next is never one of those charts. An engineer reads the August 31 retest export and the numbers look wrong. A soak log from a 90-minute burn-in has three units in it and a person wants to know, in words, whether any of them drifted. A September sweep of five prototypes has 900 readings in it and nobody has asked whether any of them is impossible. Two of those three are production test and the third is engineering test, which is the point: what makes a question this recipe’s is not the setting it came from, it is that nobody built a chart for it.

The run on the bench, stepped

“Do the retested boards actually pass?” Eighteen boards were pulled back for a retest on August 31 into retest-2026-08-31.csv, whose value_v column is meant to be volts. Before any snippet touches it, code checks every value against 40.0 V, the widest node an SRB-5030 has anywhere (VIN_ABS_MAX_V, the datasheet’s absolute maximum input). Sixteen of the eighteen, numbers like 4973.5 and 5007.7, clear that ceiling by two orders of magnitude, and dividing by 1000 brings all sixteen back under it, so code applies that correction and says so every time the table loads, whether or not the model calls the tool:

View code: as plausible volts
examples/bench_test_data_by_conversation/run.py · lines 79–108
def _as_plausible_volts(
    rows: list[dict], *, field: str = "value_v", ceiling_v: float = VIN_ABS_MAX_V
) -> tuple[list[dict], str]:
    """Range-check `field` against the widest node this board has anywhere, before any analysis
    sees it. Every reading in `field` is a volts measurement on an SRB-5030, and this board never
    carries more than `ceiling_v` volts on any pin (`docs/THE-BENCH.md`); a value past that by
    three orders of magnitude is not a surprising board, it is the wrong unit. If dividing by 1000
    brings every value back inside the ceiling, the column was millivolts and code corrects it and
    says so; if it still does not fit, this refuses rather than guess further.
    """

    def implausible(values: list[float]) -> list[float]:
        return [v for v in values if abs(v) > ceiling_v]

    raw = [row[field] for row in rows]
    bad = implausible(raw)
    if not bad:
        return rows, ""
    scaled = [{**row, field: row[field] / 1000.0} for row in rows]
    still_bad = implausible([row[field] for row in scaled])
    if still_bad:
        raise ImplausibleUnits(
            f"{len(bad)} of {len(rows)} {field!r} readings exceed {ceiling_v} V even after "
            f"dividing by 1000; refusing to guess the unit."
        )
    return scaled, (
        f"{len(bad)} of {len(rows)} {field!r} readings (as high as {max(bad):.1f}) exceeded "
        f"{ceiling_v} V, the widest node this board has anywhere. Code divided {field} by 1000 "
        f"before any analysis ran: the export is millivolts under a header that says volts."
    )

Read literally as volts, those sixteen readings clear a 4.9500 V lower limit by three orders of magnitude, which is the trap: a check that only looks at the lower limit calls all of them passes. Once the model asks for the snippet and code runs it on the corrected table, the real numbers come back: fifteen of the eighteen sit between 4.9678 V and 5.0077 V, mean 4.9840 V; three do not (SRB5030-2608-0052, SRB5030-2608-0063, SRB5030-2608-0178). The file’s own result column already had this right: whatever tool wrote the export judged pass and fail correctly and only printed the wrong header.

“Did any soak unit fail to settle?” Three boards ran 90 minutes at full load on August 27, five minutes between samples, in soak-2026-08-27.csv. Asked to look at the whole run rather than the last row, the model’s snippet takes the first and last sample per serial and reports the output drop and the case-temperature rise. Two boards settle: 2.5 mV and 1.2 mV of drop against a case that climbs to about 60 degC and stays there. The third, SRB5030-2608-0121, drops 75.5 mV while its case keeps climbing to 91.7 degC and never levels off, which is what a resistive joint does: current squared times resistance heats it, and a hotter joint is more resistive. docs/THE-BENCH.md records that this same serial is the only genuine efficiency failure in the whole production run, at 79.6%. Nothing in this snippet reads the production log or draws that connection; a person who has read both tables is the one who gets to make it.

“Is any block of the sweep impossible?” Five prototypes were swept over line, load and temperature in September into characterization-2026-09.csv: 900 readings, no limits column and no verdict, because what an engineer computes from it is a margin. The same range check runs on its vout_v column, finds every value plausible as volts, and says so: a guard that only speaks when it fires cannot be told apart from one nobody wired up. Then the question, which no dashboard has a tile for. A board cannot put out more power than it takes in, so the snippet averages each of the 180 blocks and compares vin_v * iin_a against iout_a * vout_v. One block comes back: SRB5030-2609-0001 at a labeled 12.0 V and 3.000 A draws 0.6729 A where that same point on the other four boards draws 1.3123 A, which is 8.08 W in against 14.95 W out. The supply was still at 24.0 V from the block before it. The output voltage gives nothing away, since holding it steady while the input moves is the whole job of the part, and no uncertainty budget would have caught it either: every reading in that block is a good reading of a condition nobody asked for.

The same file’s other trap is not this recipe’s. A block of readings scatters seven times as wide as the rest because the meter was left on the 100 V range, and finding that is a GROUP BY on a column nobody thinks to group by, which is the engineering-test level-0 page’s work. Ask this recipe only what a grouping has already failed to answer.

What it costs

The unit here is the question, in all three settings, because a question is what a person asks once. Two model calls each: one to write the snippet, one to turn the result into a sentence. Nothing on this site has called a live model, so this is an estimate, worked in the open with the same deterministic token counter the stub runs use: the retest question above runs 617 input tokens and 155 output tokens across its two calls, counting the system prompt describing all three tables, the snippet, the sandbox’s JSON result and the final sentence. A real model’s tokenizer will not match that exactly, but the shape holds whichever model runs it. This recipe is not competing with the dashboard, which answers most questions for nothing per question; it is competing with the ten minutes an ad hoc question costs a person with a spreadsheet and a filter box.

How it fails on a real bench, specifically

A retest export merged on a column named for volts and filled with millivolts is the whole first story above, and what catches it is the range check, in code, before any snippet runs, not a model noticing the numbers look large. That check is blunt on purpose and it can be wrong in one direction: a column that really is volts with a single mistyped row in it would be scaled by a thousand along with the bad row, and every value would still be inside the ceiling afterward. That is why the correction goes into the trace on every run instead of being applied quietly. A snippet that runs clean and computes the wrong thing is not caught by the sandbox at all; it is caught by a person reading the code next to the number, which is why the code is on the page and not just the answer. No number this recipe reports is the model’s: the sandbox computes every figure and the model writes the sentence around it, because a model never produces a reported measurement, an uncertainty, a margin or a verdict. The pass or fail on the retested boards already happened, in the file’s own result column and the limits every earlier step checked. The sweep’s margins are not on this page at all, for the same reason: a margin is a subtraction, and a subtraction belongs where it costs nothing per question.

How to evaluate it

This example runs on the bench in evals/bench/, which has no question file of its own, so the site’s 60-question document set (docs/EVALS.md) has nothing to grade it against; scripts/ eval_run.py refuses to score it and says what to measure instead. Two things, both checkable by a program rather than a person’s judgment: whether the figures in the final answer equal the figures the executed snippet actually printed, which the tests in this recipe’s own example package check against numbers recomputed independently from the CSVs, and whether the range check stops the mislabeled value_v column before a snippet ever sees it. Beyond that, a real eval is a small set of ad hoc questions over these three tables with answers worked out by hand, scored on whether the snippet ran clean or the sandbox refused it, and whether the final number matches the hand-worked one exactly. Track a sandbox refusal separately from a wrong-but-clean answer; they are different failures with different fixes, one the sandbox doing its job and the other a prompt that needs to say more about what the tables hold.

How to adapt it

The shape ports past this bench whenever a table and a question already exist and a fixed report cannot anticipate the question: a regional sales dip, survey responses, server logs after an incident. What ports specifically: one tool, a grammar checked node by node before anything runs, a range check on any column whose unit is a fact you already know, a physical identity the data has to satisfy the way a power balance does, and a person reading the snippet next to the number.

What does not port: the 40 V ceiling, which is this board’s absolute maximum input and nobody else’s; the column names; and the three tables themselves, which are evals/bench/data/ CSVs invented for this site. A reader’s own data has its own plausible ranges, its own known units and its own identities that have to hold, and the lesson of the first and third stories above is that those are worth checking in code before a snippet runs, on every table, not just these three.

Design choices

Why this level, and when to use another approach

The model writes one short Python snippet, a sandbox runs it against the loaded tables, and the model turns whatever came back into a sentence. That is level 4: one real decision (what the snippet says), then code that always runs it and always asks for a final answer. It composes two techniques: writing the snippet is function calling, one tool, offered once; running it is code execution, a sandbox that never trusts the model’s text on its own.

It is not level 5. Data analysis by conversation is the same shape with a loop around it, for a question whose second computation depends on what the first one showed: “how does this quarter compare to last” cannot even name its second query until the first one has answered. None of the three questions below works that way. Each takes one snippet and one look at the result, and a loop would buy nothing here but a second, unnecessary model call. If a reader’s own questions turn out to chain, that recipe is where to go.

It is also not level 0, and that is worth saying plainly: a fixed report cannot answer a question nobody wrote a query for yet. What is level 0 is the arithmetic inside the snippet, and the check that runs before any snippet does. Both are ordinary code, and neither is the model’s to get right or wrong.

Build it

Implementation details and code

The model’s one decision, and what code does with it, in run:

View code: run.py
examples/bench_test_data_by_conversation/run.py · lines 368–388
    call = first.tool_calls[0]
    dropped = "" if len(first.tool_calls) == 1 else f" (dropped {len(first.tool_calls) - 1} further call(s))"
    code = str(call.arguments.get("code", ""))
    tracer.record(
        kind="model",
        decided_by="model",
        title="Model writes analysis code",
        detail=code[:300] + dropped,
        tokens_in=first.tokens_in,
        tokens_out=first.tokens_out,
        ms=first.ms,
    )

    try:
        result = run_snippet(code, tables)
    except UnsafeCode as exc:
        tracer.record(kind="code", decided_by="code", title="Sandbox refused the snippet", detail=str(exc))
        return Answer(text=f"Could not safely run that analysis: {exc}", citations=[])

    result_text = _render_result(result)
    tracer.record(kind="code", decided_by="code", title="Sandbox runs the snippet", detail=result_text[:400])

The sandbox itself never calls Python’s own eval. It does call exec, unlike code execution’s own single-expression evaluator, but only after every node in the snippet has been walked and refused if it is not on a short allow-list: no import, no attribute access at all (so no x.y, and nothing dunder-chained off a literal), no lambda, no while, no function or class definitions, calls only to a fixed list of names, and loops nested no more than two deep:

View code: run snippet
examples/bench_test_data_by_conversation/run.py · lines 256–280
def run_snippet(code: str, tables: dict[str, list[dict]]) -> object:
    """Run one analysis snippet against `tables` and return whatever it assigned to `result`.

    Never calls Python's own `eval`. It does call `exec`, but only on a tree `_check_grammar` has
    already walked node by node, in a namespace with `__builtins__` emptied out, holding nothing
    but `TABLES` and the functions in `_SANDBOX_FUNCTIONS` -- so there is no name anywhere in
    scope that reaches a file, a socket, or the interpreter itself.
    """
    if len(code) > MAX_SOURCE_CHARS:
        raise UnsafeCode(f"snippet is {len(code)} characters; the limit is {MAX_SOURCE_CHARS}")
    try:
        tree = ast.parse(code, mode="exec")
    except SyntaxError as exc:
        raise UnsafeCode(f"not valid analysis code: {exc}") from exc
    _check_grammar(tree)
    namespace: dict[str, object] = {"__builtins__": {}, "TABLES": tables, **_SANDBOX_FUNCTIONS}
    try:
        exec(compile(tree, "<analysis-snippet>", "exec"), namespace)  # noqa: S102 -- see docstring
    except UnsafeCode:
        raise
    except Exception as exc:  # the snippet's own runtime error: a KeyError, a ZeroDivisionError...
        raise UnsafeCode(f"the snippet raised {type(exc).__name__}: {exc}") from exc
    if "result" not in namespace:
        raise UnsafeCode("the snippet must assign its answer to a variable named result")
    return namespace["result"]

What that grammar protects against: every classic escape this site’s own tests try against it ((1).__class__, __import__('os'), open(...), an f-string, a walrus, a while True) is refused the same way, because the node type it needs is not in the allow-list, not because the code recognized an attack. What it does not protect against: a snippet that runs clean and answers a question other than the one asked, which is why the sandbox’s own output goes next to the prose and not instead of it, and why the loop and node caps are a coarse defense against a runaway snippet rather than a wall-clock timeout. A production deployment wants an OS-level sandbox around this too, the same point code execution makes about a real container.

Composition

Techniques this recipe uses

The highest level it needs is level 4.

Code execution

Sourced

Letting the model write code and run it in a sandbox.

Function calling

Sourced

Letting the model call functions that you define.

Same shape, other jobs

Ask questions of data you do not fully understand yet

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • A production yield drop: bad lot, drifting fixture or real design margin
  • A fall in sales in one region
  • Survey results
  • Server logs after an incident
  • Results of an experiment with many factors
  • Characterization data across temperature and voltage
  • Why one block of readings in a session scatters wider than the rest

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page