The example is plan-and-execute: one call with no tools asks the model to write a short plan, then
a loop offers two tools, search and lookup_part, and lets the model act on the plan, revise it,
and decide when to stop. This is deliberately not the same shape as
agentic RAG’s example, which is a single continuous
loop with no separate planning call: compare the two traces and the two papers behind them.
ReAct interleaves reasoning traces with task-specific actions[1]. Plan-and-Solve draws
its plan up before any action: Wang et al. list three pitfalls in zero-shot chain-of-thought
(calculation errors, missing-step errors and semantic misunderstanding errors), propose
Plan-and-Solve against the missing steps, and extend it to PS+ for the calculation
errors[2].
The planning call is the one line in this example that looks like a model decision but is not
one. It is kind: "model" because a model ran, but decided_by: "code": your code always makes
this call, and always moves on to the execution loop next, whatever the plan actually says: the
same rule the single call in RAG’s example follows. Only once
the loop starts offering tools does the model’s own output pick what happens next: which tool,
with what arguments, or to stop. Every one of those steps is decided_by: "model", matching
examples/agentic_rag/.
max_steps (default 5) and max_tokens (default 3000) are the hard caps, checked after every
tool result. Hitting either forces one last no-tools call for a final answer (decided_by: "code", since the model never chose to stop), and the trace records which cap did it, so a
partial answer never looks like a normal one. tests/test_example_single_agent.py scripts a model
that never stops calling tools on its own and checks both caps actually cut the run short. Those
caps stop one run; they carry no state into a new one. Once the work outlives a single context
window or a single sitting, see long-running tasks
for what picks up across that gap.
The same loop, aimed at a bench instead of a document set, is
bringing up a failed board. That is
engineering test, not production test, even though the board came off a production line: one
board, an afternoon, and an answer that is a cause and a next measurement rather than a pass or a
fail. Three read-only tools query the test log, the bench documents and one instrument, and each
measurement narrows in on a cause the way each tool call here narrows in on an answer. Nothing in
that recipe’s tools can set a voltage or enable an
output; the board is already energized under a technician’s own approved sequence before the
agent’s loop ever starts.
examples/single_agent/run.py · lines 51–99
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
max_steps: int = MAX_STEPS,
max_tokens: int = MAX_TOKENS,
) -> Answer:
del embedder # a single agent retrieves through its tools, not a vector index
sections = load_sections(corpus_dir)
plan_call = model.complete(
[Message(role="system", content=PLAN_SYSTEM), Message(role="user", content=question)], max_tokens=200
)
record_completion(tracer, decided_by="code", title="Model writes a plan", completion=plan_call)
tokens_used = plan_call.tokens_in + plan_call.tokens_out
messages = [
Message(role="system", content=ACT_SYSTEM.format(plan=plan_call.text)),
Message(role="user", content=question),
]
citations: list[str] = []
for _ in range(max_steps):
completion = model.complete(messages, tools=TOOLS, max_tokens=400)
tokens_used += completion.tokens_in + completion.tokens_out
if not completion.tool_calls:
record_completion(tracer, decided_by="model", title="Model stops and answers", completion=completion)
return Answer.from_text(completion.text, retrieved_sources=citations)
calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
record_completion(tracer, decided_by="model", title="Model acts on the plan", completion=completion, detail=calls_desc)
turn, calls = assistant_turn(completion, len(messages))
messages.append(turn)
for call in calls:
result_text, cites = _run_tool(call, sections)
citations.extend(cites)
tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
messages.append(tool_result(call, result_text))
if tokens_used >= max_tokens:
reason = f"token budget reached: {tokens_used} >= {max_tokens}"
final = force_final(messages, model, tracer, reason=reason, max_tokens=400)
return Answer.from_text(final.text, retrieved_sources=citations)
final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
return Answer.from_text(final.text, retrieved_sources=citations)
Run it yourself:
examples/single_agent/README.md · lines 17–17
python -m examples.single_agent --model stub:scripted