Recipe

Data analysis by conversation

A single agent writes and runs code against a dataset, one question at a time, to answer questions a fixed query could not anticipate.

SourcedNeeds level 5

Sales and returns for a small retailer live in a spreadsheet, and nobody built a dashboard for every question somebody might ask of it. A manager wants something specific (which category has the worst return rate this quarter, and how that compares with last quarter) answered from the data rather than guessed fluently.

Data analysis by conversation is that job: a single agent writes and runs code against the dataset itself, one question at a time, seeing each result before deciding whether it needs another pass or already has enough to answer.

Example run

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Data analysis by conversation, assembled

The model writes a plan, then writes and runs code against the dataset, deciding after each result whether it needs another pass.

Level 5 · Agent loops
Question about the datasetQuestion aboutthe datasetMODELwrites a short planwrites a short planMODELpicks the next steppicks the next stepTOOLrun_python(snippet)run_python(snippet)AnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 07Your code chose

The question arrives

"Which product category has the highest return
rate this quarter, and how does that compare
to last quarter?"
0 tokens · 0 ms

Walkthrough

The loop composes the runnable code already on the single agent and code execution pages, adapted in one specific way: single agent’s own example offers two tools, search and lookup_part, over the site’s document corpus; this job offers one, run_python(snippet), over a dataset instead. The plan-then-loop shape carries over exactly: a first call writes a short plan your code moves on from regardless of what it says, then a loop lets the model pick an action, see the result, and decide whether to act again or stop, matching examples/single_agent/run.py’s own two hard caps on steps and tokens.

Code execution’s own example is honest about a real limit worth carrying into this job: its sandbox is an arithmetic whitelist (numbers, + - * /, nothing else) built to make one safe numeric answer, not to run a real analysis over a table. A sandbox built for this recipe needs a wider allow-list (filtering, grouping, aggregating a table) while keeping the same shape that example’s whitelist does: refuse anything that isn’t on it, rather than trying to recognize an attack. The production containers that page quotes are bounded from the outside the same way: it cites one maker’s own settings, internet access disabled, fixed memory, a maximum execution time, so whatever the model writes can only compute against the data it was given.

The run above shows two computations run in sequence. The model computes this quarter’s return rate by category, sees that outdoor gear stands out, and, without being told to, runs the comparable query for last quarter before answering, because the question asked for a comparison, not one number.

Two things the composition doesn’t settle. The dataset has to be inside the sandbox before the first snippet runs, put there by your code, since a sandbox with no network cannot go and fetch it. And every result the sandbox returns goes into the model’s context, so rows it computes over are rows it reads; log each snippet, the result it produced, and whether the allow-list refused it.

What to measure

Single agent is scored by the site’s shared 60-question set (see docs/EVALS.md); its model_decided_steps count and its cap-hit rate both carry over directly to this job, on different questions. Code execution is not (it answers only the 12 numeric questions there), and a spreadsheet task needs its own set regardless: a small, fixed dataset with known answers computed by hand, and a batch of questions ranging from one computation to several. Score final answer accuracy against those known answers, the average and worst-case number of actions a question takes, and the refusal rate on a snippet that tries to reach outside the sandbox’s allow-list (the number code execution’s own page tracks separately from a correct answer). No result file exists for either technique on this task yet, so this recipe cannot claim a score for any of it.

Variations

  • Widen the sandbox’s allow-list toward real tabular operations while keeping it off the network and the filesystem; code execution’s own failure modes cover the tension between a whitelist too strict to be useful and one too permissive to be safe.
  • Tighten the step cap for a dataset with a small, known set of useful queries, so a loop that isn’t making progress fails fast instead of burning the whole cap.
  • Add human approval before a computed number turns into a report or a decision with a real consequence, the same gate support desk uses before a costly reply ships.
  • If the same few questions get asked every week, a fixed query or a dashboard answers them for less than an agent re-deriving them each time; save the agent for the question nobody wrote a query for yet.

Design choices

Why this level, and when to use another approach

Two techniques compose this recipe. Code execution is the sandbox: the model writes a short piece of code, and a restricted interpreter (never the model itself) actually runs it and returns whatever it computed. A single agent is what lets that happen more than once: the model sees each result and decides, itself, whether another computation is needed or the question is already answered.

Code execution alone (one write-and-run cycle, bounded from outside) would be enough if every question here needed exactly one computation. Some do. Many ad hoc questions don’t: answering “how does this quarter compare to last” means computing one quarter, looking at what came back, and only then knowing that a second, comparable computation is needed before there’s anything to compare. The single agent page states the condition directly (move up once the number and order of actions cannot be known before the model sees the question), and that is exactly what a “how does X compare to Y” or “what changed since last time” question does: the second step’s shape depends on the first step’s result, not on anything decided in advance.

It is not worth climbing past a single agent for this job. One agent with one broad sandbox tool already reruns, refines and compares on its own; a second model would help only if the questions needed independent verification of a specific number, or spanned more data than one context window holds, and an ordinary ad hoc question over one spreadsheet is neither. Multiple agents would add cost and coordination for a comparison one agent already reaches by looking at its own last result.

Composition

Techniques this recipe uses

The highest level it needs is level 5.

Code execution

Sourced

Letting the model write code and run it in a sandbox.

Single agent

Sourced

A model that plans, acts and checks its own work in a loop.

Same shape, other jobs

Ask questions of data you do not fully understand yet

This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.

  • A production yield drop: bad lot, drifting fixture or real design margin
  • A fall in sales in one region
  • Survey results
  • Server logs after an incident
  • Results of an experiment with many factors
  • Characterization data across temperature and voltage
  • Why one block of readings in a session scatters wider than the rest

Last reviewed 09/18/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page