# Single agent

_Level 05 · Agent loops · sourced_

A model that plans, acts and checks its own work in a loop.

## Conceptual architecture: A decision loop, with a way out.

The model chooses a proposed next step. Software decides whether it can run.

- **Goal + context:** Task, instructions, selected history
- **Model decision:** Request a tool or return an answer
- **Execution gate:** Arguments, permissions, budgets
- **Run allowed tool:** Bounded operation in the environment
- **Observe the result:** Return output or a useful error
- **Finish or hand back:** Return work, evidence, and gaps
- **Pause or refuse:** Approval needed, denied, or capped

Connections:
- Goal + context → context → Model decision
- Model decision → tool request → Execution gate
- Execution gate → allowed → Run allowed tool
- Run allowed tool → observation → Observe the result
- Observe the result → next decision → Model decision
- Model decision → final answer → Finish or hand back
- Execution gate → cannot proceed → Pause or refuse

Reasoning helps the model choose useful actions. The loop supplies feedback; the harness supplies execution, state, and enforced limits. None of those makes the answer automatically correct.
- **Control:** Tool output is evidence, not permission to take another action.
- **Stopping:** Finish, ask for help, or stop at a step, time, or cost limit.
- **Verification:** Inspect the environment and the final artifact, not just the model’s account of its work.

## Try this in a recipe
- [Investigate an incident with bounded tools](/gradient_ascent/recipes/incident-runbook.md): Let a model choose read-only diagnostic tools, then require an evidence-backed handoff within six calls.

## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow one agent choosing its next action from the results it receives. Watch the loop gather information, revise its approach, and decide whether it has enough to finish.

**Assumptions:** The request needs a recognizable completion condition and available tools. Autonomy does not imply unrestricted access or unlimited attempts.

**Design choices:** Let the model choose between useful next steps when the route is uncertain. Use a fixed workflow when the steps are already well known.

**Request:** Find why our booking failed and explain the next step.

**Starting evidence:** Tools: check_room, read_policy. Room is free. Policy: under two hours' notice requires staff review.

**Action and control:** The model chooses policy lookup after availability fails to explain rejection; observations guide its next choice.

**Stage records (authored, not executed):**

### Investigation state · opened

Request: explain a failed booking.
Available tools: check_room, read_policy.
Known initially: booking failed.
Unknown initially: availability and applicable booking rules.
No booking-write tool is exposed in this fixture.

What changed: The agent has an investigation task, not authority to confirm a booking.

### First action · selected

Proposed next call: check_room.
Purpose: test whether room availability explains the failure.
Alternative: inspect policy first.
This authored trace chooses availability; another valid investigation could start elsewhere.

What changed: There is a concrete evidence question behind the action, not a claim about hidden model reasoning.

### Observation and next choice

check_room result: room is free.
Updated state: unavailability does not explain the failure.
Next proposed call: read_policy.
Returned policy: under two hours' notice requires staff review.
Still missing: proof that this rule was the actual rejection reason.

What changed: A tool observation changes the useful next step. The policy supplies a possible explanation, not the booking system's rejection log.

### Finding · qualified

Likely cause: short-notice review rule, if this request fell within that window.
Next step: confirm request timing or obtain the rejection reason; ask staff to review if applicable.
Booking status: not confirmed.
No bypass attempted.

What changed: The finding is deliberately conditional because the fixture lacks request timing and a rejection log.

### Bounded unsuccessful run

Alternate run stops before policy lookup.
Completed: availability check.
Unresolved: why the booking failed.
Handoff: report partial findings and the next useful source to inspect.
Do not replace missing observations with a guessed cause.

What changed: A step budget bounds work; it does not establish completion.

### Transfer the loop

Replace: booking tools with your task's read or action tools.
Define: what counts as completion and when another lookup is useful.
Record: observation → proposed next action → result.
Use a fixed workflow instead when the route is already predictable.

What changed: The reusable pattern is choosing a next step from evidence, with a clear way to stop or hand back uncertainty.

**Sample result:** Possible cause: the short-notice rule, if this request was made under two hours before the booking. Confirm request timing or the rejection reason, then seek staff review if applicable. No booking is confirmed.

**Change something — Exhaust the step budget before policy lookup:** Partial result: availability checked, cause unresolved. The cap does not turn uncertainty into a conclusion.

**Decision:** Is a capped run a completed investigation?

**Answer:** No; report the unresolved question.

**Why:** Distinguish model-chosen next actions from a fixed chain; stop on insufficient evidence or a step cap.

**Review criteria:** An observation record separating availability, policy, unknown request timing, and the actual rejection reason; a partial handoff when the run is capped.

**Recovery:** When progress stalls, inspect the evidence gap and choose a different approach, ask a question, or return a partial result. Repeated identical calls are not progress.

**Adapt it:** Apply the loop to investigation, planning, or bounded project work. Choose tools, budgets, and stopping conditions proportionate to the task rather than copying another agent's limits.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow one agent choosing its next action from the results it receives. Watch the loop gather information, revise its approach, and decide whether it has enough to finish.

**Assumptions:** The request needs a recognizable completion condition and available tools. Autonomy does not imply unrestricted access or unlimited attempts.

**Design choices:** Let the model choose between useful next steps when the route is uncertain. Use a fixed workflow when the steps are already well known.

**Request:** Find a workable library visit time using the supplied tools.

**Starting evidence:** Tools: read_hours, check_bus_schedule. Library open Saturday; one bus route is suspended.

**Action and control:** The model chooses a follow-up transport lookup after seeing opening hours; each result informs the next choice.

**Stage records (authored, not executed):**

### Input record

Tools: read_hours, check_bus_schedule. Library open Saturday; one bus route is suspended.

What changed: Establish the facts supplied for this version of the task.

### Design note

Let the model choose between useful next steps when the route is uncertain. Use a fixed workflow when the steps are already well known.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

The model chooses a follow-up transport lookup after seeing opening hours; each result informs the next choice.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Propose a visit using the remaining route if its schedule supports arrival. Do not claim a reservation or buy a ticket.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Trace tool results, the resulting choices, and unresolved constraints.

If the result falls short:
When progress stalls, inspect the evidence gap and choose a different approach, ask a question, or return a partial result. Repeated identical calls are not progress.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply the loop to investigation, planning, or bounded project work. Choose tools, budgets, and stopping conditions proportionate to the task rather than copying another agent's limits.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Propose a visit using the remaining route if its schedule supports arrival. Do not claim a reservation or buy a ticket.

**Change something — No route can arrive during opening hours:** Stop with the constraint conflict and ask about another date or transport option.

**Decision:** Should an agent force a plan when tool evidence shows none works?

**Answer:** No; report the conflict and ask.

**Why:** Autonomous next-step choice still needs evidence and an honest stop condition.

**Review criteria:** Trace tool results, the resulting choices, and unresolved constraints.

**Recovery:** When progress stalls, inspect the evidence gap and choose a different approach, ask a question, or return a partial result. Repeated identical calls are not progress.

**Adapt it:** Apply the loop to investigation, planning, or bounded project work. Choose tools, budgets, and stopping conditions proportionate to the task rather than copying another agent's limits.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow one agent choosing its next action from the results it receives. Watch the loop gather information, revise its approach, and decide whether it has enough to finish.

**Assumptions:** The request needs a recognizable completion condition and available tools. Autonomy does not imply unrestricted access or unlimited attempts.

**Design choices:** Let the model choose between useful next steps when the route is uncertain. Use a fixed workflow when the steps are already well known.

**Request:** Investigate a failed archived test run without operating hardware.

**Starting evidence:** Tools: read_test_log, inspect_config, read_driver_docs. Log shows timeout after a mode change.

**Action and control:** The model decides which records to inspect next, comparing configured settling time with documentation.

**Stage records (authored, not executed):**

### Input record

Tools: read_test_log, inspect_config, read_driver_docs. Log shows timeout after a mode change.

What changed: Establish the facts supplied for this version of the task.

### Design note

Let the model choose between useful next steps when the route is uncertain. Use a fixed workflow when the steps are already well known.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

The model decides which records to inspect next, comparing configured settling time with documentation.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Candidate cause: settling time mismatch. Propose a reviewed change and validation plan; do not claim root cause proven.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Inspect observation sequence, competing explanations, missing evidence, and proposed validation.

If the result falls short:
When progress stalls, inspect the evidence gap and choose a different approach, ask a question, or return a partial result. Repeated identical calls are not progress.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply the loop to investigation, planning, or bounded project work. Choose tools, budgets, and stopping conditions proportionate to the task rather than copying another agent's limits.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Candidate cause: settling time mismatch. Propose a reviewed change and validation plan; do not claim root cause proven.

**Change something — Logs lack timestamps needed to test the hypothesis:** Stop with a hypothesis and request evidence. Do not rewrite the framework to make the symptom disappear.

**Decision:** Is a plausible diagnosis the same as a verified cause?

**Answer:** No; identify what evidence is still needed.

**Why:** An agent can investigate adaptively without gaining authority to execute or overstate conclusions.

**Review criteria:** Inspect observation sequence, competing explanations, missing evidence, and proposed validation.

**Recovery:** When progress stalls, inspect the evidence gap and choose a different approach, ask a question, or return a partial result. Repeated identical calls are not progress.

**Adapt it:** Apply the loop to investigation, planning, or bounded project work. Choose tools, budgets, and stopping conditions proportionate to the task rather than copying another agent's limits.

A single agent puts the model in charge of a loop, not just one choice. At
[function calling](/gradient_ascent/techniques/function-calling/), the model picks one tool once
and your code takes it from there. A single agent feeds its own output back in and decides again
each time, including deciding when it is done.

The research names two shapes for that loop. ReAct generates "both reasoning traces and
task-specific actions in an interleaved manner", the traces there to help the model "induce,
track, and update action plans as well as handle exceptions"[1]. Plan-and-execute writes
a plan first and works through it a step at a time; Wang et al. call this Plan-and-Solve and
propose it against one of the three pitfalls they list in zero-shot chain-of-thought prompting,
missing-step errors[2].

Level 5 is the first level where two things are both the model's decision: which action to take,
and when to stop. Your code still runs every tool and returns every result, and enforces caps the
model cannot override: on steps, on tokens, and on which tools it may attempt at all. Hit a cap
first and your code forces the final answer; that stop is the program's decision, not the
model's.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 5 · Single agent. The same steps are described in the sections below._

## Practical guidance

"Single agent" is not a feature you turn on: it is the loop running underneath a coding agent
deciding which file to open next, a deep-research mode deciding whether to search again, a voice
agent deciding whether to keep talking, or a skill-picking agent deciding which instructions to
load. See [coding agents](/gradient_ascent/techniques/coding-agents/), [agentic RAG](/gradient_ascent/techniques/agentic-rag/), [voice
agents](/gradient_ascent/techniques/voice-agents/) and [skills](/gradient_ascent/techniques/skills/) for what to actually do with each
of those. This page is about the one thing they share: what it means when one of them just stops.

Every product built this way has a limit on it you will never see a setting for, only its effect.
Anthropic's Agent SDK caps the number of tool-use round trips and how much a session may spend, and
ends the run with a specific reason attached rather than continuing forever[4]. OpenAI's
Agents SDK does the same by default, raising a distinct error once a run's turn count passes its
limit[5]. When a coding agent, a research assistant or any other tool built this way stops
partway through a long task with no obvious error, that is very likely what happened: it hit a cap
built into the product, not a crash. Asking it to continue, or splitting the task into a smaller
piece and starting fresh, is the right response, not repeating the same request and expecting a
different stopping point.

Setbacks along the way do not usually end a run on their own. Anthropic's documentation says that
when a step does not go as planned, the agent "typically attempts a different approach or reports
that it couldn't proceed"[4], which is usually good, but it also means a wrong turn early
on can carry through several more steps before anything looks visibly wrong. Anthropic's broader
advice on this is worth taking at face value: autonomy brings "higher costs, and the potential for
compounding errors,"[3] so read the final answer against what you actually asked for, not
just against whether it sounds finished.

None of the settings themselves (how many turns, how much it may spend, which tools it may use) are
something you set: whoever built the product you are using chose those. Your job is reading the
result and knowing that a run stopping short is usually a limit, not a failure.

## Implementation details

The example is plan-and-execute: one call with no tools asks the model to write a short plan, then
a loop offers two tools, `search` and `lookup_part`, and lets the model act on the plan, revise it,
and decide when to stop. This is deliberately not the same shape as
[agentic RAG](/gradient_ascent/techniques/agentic-rag/)'s example, which is a single continuous
loop with no separate planning call: compare the two traces and the two papers behind them.
ReAct interleaves reasoning traces with task-specific actions[1]. Plan-and-Solve draws
its plan up before any action: Wang et al. list three pitfalls in zero-shot chain-of-thought
(calculation errors, missing-step errors and semantic misunderstanding errors), propose
Plan-and-Solve against the missing steps, and extend it to PS+ for the calculation
errors[2].

The planning call is the one line in this example that looks like a model decision but is not
one. It is `kind: "model"` because a model ran, but `decided_by: "code"`: your code always makes
this call, and always moves on to the execution loop next, whatever the plan actually says: the
same rule the single call in [RAG](/gradient_ascent/techniques/rag/)'s example follows. Only once
the loop starts offering tools does the model's own output pick what happens next: which tool,
with what arguments, or to stop. Every one of those steps is `decided_by: "model"`, matching
`examples/agentic_rag/`.

`max_steps` (default 5) and `max_tokens` (default 3000) are the hard caps, checked after every
tool result. Hitting either forces one last no-tools call for a final answer (`decided_by:
"code"`, since the model never chose to stop), and the trace records which cap did it, so a
partial answer never looks like a normal one. `tests/test_example_single_agent.py` scripts a model
that never stops calling tools on its own and checks both caps actually cut the run short. Those
caps stop one run; they carry no state into a new one. Once the work outlives a single context
window or a single sitting, see [long-running tasks](/gradient_ascent/techniques/long-horizon/)
for what picks up across that gap.

The same loop, aimed at a bench instead of a document set, is
[bringing up a failed board](/gradient_ascent/recipes/bring-up-debug-assistant/). That is
engineering test, not production test, even though the board came off a production line: one
board, an afternoon, and an answer that is a cause and a next measurement rather than a pass or a
fail. Three read-only tools query the test log, the bench documents and one instrument, and each
measurement narrows in on a cause the way each tool call here narrows in on an answer. Nothing in
that recipe's tools can set a voltage or enable an
output; the board is already energized under a technician's own approved sequence before the
agent's loop ever starts.

`examples/single_agent/run.py` (lines 51-99)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    del embedder  # a single agent retrieves through its tools, not a vector index
    sections = load_sections(corpus_dir)

    plan_call = model.complete(
        [Message(role="system", content=PLAN_SYSTEM), Message(role="user", content=question)], max_tokens=200
    )
    record_completion(tracer, decided_by="code", title="Model writes a plan", completion=plan_call)
    tokens_used = plan_call.tokens_in + plan_call.tokens_out
    messages = [
        Message(role="system", content=ACT_SYSTEM.format(plan=plan_call.text)),
        Message(role="user", content=question),
    ]

    citations: list[str] = []
    for _ in range(max_steps):
        completion = model.complete(messages, tools=TOOLS, max_tokens=400)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            record_completion(tracer, decided_by="model", title="Model stops and answers", completion=completion)
            return Answer.from_text(completion.text, retrieved_sources=citations)

        calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
        record_completion(tracer, decided_by="model", title="Model acts on the plan", completion=completion, detail=calls_desc)
        turn, calls = assistant_turn(completion, len(messages))
        messages.append(turn)
        for call in calls:
            result_text, cites = _run_tool(call, sections)
            citations.extend(cites)
            tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
            messages.append(tool_result(call, result_text))

        if tokens_used >= max_tokens:
            reason = f"token budget reached: {tokens_used} >= {max_tokens}"
            final = force_final(messages, model, tracer, reason=reason, max_tokens=400)
            return Answer.from_text(final.text, retrieved_sources=citations)

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
    return Answer.from_text(final.text, retrieved_sources=citations)
```

Run it yourself:

`examples/single_agent/README.md` (lines 17-17)

```text
python -m examples.single_agent --model stub:scripted
```

## When you do not need this

Try [function calling](/gradient_ascent/techniques/function-calling/) first if one tool call,
chosen once, is enough for the task: most single lookups are.

Try a fixed sequence of steps instead ([workflow
graphs](/gradient_ascent/techniques/workflow-graphs/) or a simpler chain) if you can write down in advance which actions the task needs
and in what order. A workflow like that is wrong the same way every time it is wrong, and it costs
the same every time it runs, which a single agent does not.

Move up to a single agent once the number and order of actions cannot be known before the model
sees the question: a search that might take one lookup or five, depending on what the first one
turns up.

## Failure modes

### Looping without progress

- **How to notice it:** The model calls the same tool with the same or a barely different argument several times in a row, learning nothing new from the result, until the step cap forces a stop.
- **How to test for it:** Run a question the tools cannot actually answer and read the tool-call arguments in order. Real progress looks like each call narrowing in on something; a loop looks like the same call repeated with cosmetic changes.

### Drift from the question

- **How to notice it:** The agent's later actions chase a detail it noticed mid-loop rather than the question it was actually asked. Anthropic describes this as part of the cost of autonomy: agents left to direct themselves carry "the potential for compounding errors".
- **How to test for it:** Read every tool call in order and ask whether each one still serves the original question, not just whether it returned something plausible.

### Tool misuse

- **How to notice it:** The model calls a tool with an argument it invented rather than one it actually retrieved earlier in the trace: a part number it guessed, not one a search or a prior lookup returned.
- **How to test for it:** Trace every tool argument back to where it came from: a prior tool result, or nowhere. An argument that traces to nowhere is a guess, whether or not the tool call itself succeeds.

### Overconfidence at the stop

- **How to notice it:** The model stops and states an answer with no hedge, even though an earlier tool result only partly supported it or the two results it gathered actually conflicted.
- **How to test for it:** Compare the final answer's claims against the tool results actually returned in the trace, not against whether tools were called at all.

### The cap ships a known-partial answer

- **How to notice it:** A step or token cap is reached before the model stopped on its own, and the forced final answer goes out anyway, silently unless the "Force a final answer" step is surfaced somewhere a person or a downstream system can see it.
- **How to test for it:** Script a model that never stops calling tools (this page's own test suite does exactly this) and confirm the run still returns an answer, and that the answer's origin says which cap forced it.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, best case (plan, one action, stop):** 3
- **Model calls, worst case (step cap reached):** 7
- **Tokens in, one action round:** ~340–610
- **Wall time, one round trip:** ~0.7s

**Compared with a fixed workflow with the same steps (level 3).** Cost here tracks how many actions the question actually needs, not a count fixed in advance: a question one action can answer costs close to what a two-call workflow costs, and one that reaches the cap costs several times that, for the same question.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

`single_agent` answers a question about the documents and cites what it used, the same task `rag`
and `agentic_rag` are scored on, so it fits the site's own 60-question set the same way: exact or
rubric match, citation hit rate, and the count of trace steps the model itself decided against the
ones the code decided, which the trace already carries.

`scripts/eval_run.py` counts `single_agent` among the examples the question set can score. No
result file exists for it yet, so this page cannot say a number for any of it. Run
`python scripts/eval_run.py --example single_agent --model <spec> --dry` to project the cost of a
real run before spending anything on one.

## Run it

**What to monitor.** The share of runs that end with a forced final answer instead of a real stop (the cap-hit rate), the average number of actions per run, and how often a tool call's argument cannot be traced back to an earlier result in the same run.

**Cost at volume.** Cost per question is not fixed the way a workflow's is: it depends on how many actions the model decides it needs. Budget for a run that uses the full step cap on every question, not the average, since a bad batch of questions becomes a bad batch of spend.

**How it fails in production.** The loop calls the same or a near-identical tool repeatedly without making progress, silently consuming the whole step cap on a question the tools were never going to answer.

**What to log.** The plan text, every tool call with its arguments and result, in order, and which cap (if any) forced the final answer, so a bad answer traces back to a specific decision instead of an unexplained partial result.

## Try it

1. **Use it.** Ask a coding agent or a deep-research tool to do something that takes several steps, and watch what happens if it does not finish: does it say it hit a turn or time limit, or does it just hand back a partial answer with no explanation?
2. **Build it.** Run python -m examples.single_agent --model stub:scripted from the repo root. The model writes a plan, acts on it twice (search, then lookup_part), and stops on its own with a priced, warranty-scoped answer, every cited section one a tool really returned. Run it again with --model stub: the echo is never a tool call, so the run stops after the plan step with no citations. For the caps, run python -m unittest tests.test_example_single_agent -v.
3. **Either lane.** Pick one of the failure modes above and try to script a StubModel response that causes it on purpose, using the pattern in tests/test_example_single_agent.py.


## Sources

1. [ReAct: Synergizing Reasoning and Acting in Language Models](https://arxiv.org/abs/2210.03629) — arXiv (Princeton University, Google Research) (accessed 2026-09-19)
2. [Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models](https://arxiv.org/abs/2305.04091) — arXiv (accessed 2026-09-19)
3. [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) — Anthropic, 2024-12-19 (accessed 2026-09-19)
4. [How the agent loop works](https://code.claude.com/docs/en/agent-sdk/agent-loop) — Anthropic (Claude Agent SDK documentation) (accessed 2026-09-19)
5. [Running agents](https://openai.github.io/openai-agents-python/running_agents/) — OpenAI (Agents SDK documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
