# Agentic RAG and deep research

_Level 05 · Agent loops · measured_

An agent that runs its own searches until it has an answer.

## Conceptual architecture: Search again only when the evidence calls for it.

An agent chooses searches and document reads; the application enforces access and a search budget.

- **Research question:** Scope, permitted sources, constraints
- **Model decision:** Request a tool or return an answer
- **Execution gate:** Arguments, permissions, budgets
- **Search or read:** Retrieve passages or open a source
- **Inspect evidence:** Coverage, conflicts, source provenance
- **Answer or report gaps:** Citations, uncertainty, unresolved facts
- **Pause or refuse:** Approval needed, denied, or capped

Connections:
- Research question → context → Model decision
- Model decision → tool request → Execution gate
- Execution gate → allowed → Search or read
- Search or read → observation → Inspect evidence
- Inspect evidence → next decision → Model decision
- Model decision → final answer → Answer or report gaps
- Execution gate → cannot proceed → Pause or refuse

More searches create opportunities to fill gaps and to introduce errors. Judge evidence coverage and claim support separately from how many steps the agent took.
- **Control:** Tool output is evidence, not permission to take another action.
- **Stopping:** Finish, ask for help, or stop at a step, time, or cost limit.
- **Verification:** Inspect the environment and the final artifact, not just the model’s account of its work.

## Try this in a recipe
- [Answer a warranty question with evidence](/gradient_ascent/recipes/document-qa.md): Retrieve the relevant policy, answer each part of the question, and distinguish an unknown fact from a retrieval miss.

## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow an investigation in which the model chooses follow-up searches as evidence arrives. Inspect how each new source changes the question and whether further searching is still useful.

**Assumptions:** Relevant evidence may span sources or contain contradictions. Search autonomy cannot compensate for missing access or a collection that lacks the answer.

**Design choices:** Use a fixed retrieval pass for simple questions; allow iterative retrieval when the first result exposes a real gap. Record support for final claims rather than search volume.

**Request:** Investigate whether this DW-480 water-damage repair is covered.

**Starting evidence:** Fictional manual v3 gives two-year coverage. Service notes link addendum A3, which excludes water damage for the DW-480. The first search returns only duration; a later lookup can retrieve A3.

**Action and control:** Agent chooses a follow-up search for exclusions and combines evidence before answering.

**Stage records (authored, not executed):**

### Input record

Fictional manual v3 gives two-year coverage. Service notes link addendum A3, which excludes water damage for the DW-480. The first search returns only duration; a later lookup can retrieve A3.

What changed: Establish the facts supplied for this version of the task.

### Design note

Use a fixed retrieval pass for simple questions; allow iterative retrieval when the first result exposes a real gap. Record support for final claims rather than search volume.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Agent chooses a follow-up search for exclusions and combines evidence before answering.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Not covered under the supplied exclusion. Cite both the warranty and addendum with their versions.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Search trajectory, evidence accumulated per step, citations, conflicts, and a stop/abstain case.

If the result falls short:
When sources disagree, identify the conflict and seek discriminating evidence. Stop with a qualified finding when the remaining uncertainty cannot be resolved within the task.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for technical investigation, policy research, or project analysis. Define acceptable sources, evidence freshness, and what counts as enough support for your decision.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Not covered under the supplied exclusion. Cite both the warranty and addendum with their versions.

**Change something — The addendum cannot be retrieved:** Duration alone does not settle coverage. Stop and identify missing evidence rather than search indefinitely or guess.

**Decision:** Should repeated searching force a definitive answer?

**Answer:** No; stop or abstain when evidence is insufficient.

**Why:** Compare against one-pass RAG on the same evidence; repeated search can still miss facts or exceed its budget.

**Review criteria:** Search trajectory, evidence accumulated per step, citations, conflicts, and a stop/abstain case.

**Recovery:** When sources disagree, identify the conflict and seek discriminating evidence. Stop with a qualified finding when the remaining uncertainty cannot be resolved within the task.

**Adapt it:** Use this for technical investigation, policy research, or project analysis. Define acceptable sources, evidence freshness, and what counts as enough support for your decision.

Agentic RAG puts the search loop itself under the model's control. [RAG](/gradient_ascent/techniques/rag/) always searches once and asks the model once. Agentic RAG
instead offers the model a search tool it can call as many times as it decides it needs, lets it
read what comes back, and lets it decide whether to search again, read more closely, or stop and
answer: the [single-agent](/gradient_ascent/techniques/single-agent/) loop aimed at retrieval.

Google describes its own version this way: at each step, "the model has to ground itself on all
information gathered so far, then identify missing information and discrepancies it wants to
explore", continuing until "the model determines enough information has been gathered"[1]. OpenAI's deep research models work the same way: "agentic" systems that "conduct
multi-step research" and return a listing of every search made along the way[3]. In both
cases the model chooses the next query and when to stop; your code still runs every search and can
cut the loop off with a hard cap regardless of what the model would have done next.

This page is measured: the cost and the score under How to Evaluate It come from a recorded run of this example on a real model, beside the RAG page's run on the same model and questions, and hold for that model's class. The step-through just below is still a scripted illustration, and source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 5 · Agentic RAG. The same steps are described in the sections below._

## Practical guidance

This is what "Deep Research" or "DeepSearch" does in a chat app: ChatGPT, Claude, Gemini,
Perplexity and Grok DeepSearch all ship a mode like it, usually a toggle or a separate button next
to the ordinary send button. Reach for it when a question has more than one part living in
different places, or when two sources might disagree and you want that checked rather than guessed
past: "Compare what our returns policy says about damaged items against what the shipping
carrier's own terms say, and tell me where they conflict." A question one search can already answer
does not need it, and costs more here for nothing extra.

Once it is running, Google's own description of the mechanism is specific: the model "oversees the
execution of" a research plan, and at each step "the model reasons over information available to
decide its next move"[1]. That means the searches it runs are worth reading, not just the
report at the end. Open the sources or search-steps panel most of these products show: OpenAI's
deep research output "will contain a listing of web search calls, code interpreter calls, and
remote MCP calls made to get to the answer,"[3] so the individual searches are visible,
not hidden inside the final prose.

Check a claim in the report against a source it actually shows, the way you would check a citation
in plain RAG search: does the linked source really say what the report claims, and does the report
ever cite a source it does not appear to have opened at all.

There is a cap you do not see. OpenAI documents a setting a developer can turn to control the total
number of tool calls a deep-research run may make before returning a result[3], and every
product like it has some version of the same limit. A report that reads thinner than the question
deserved, especially one with several parts, may be a run that hit its cap rather than one that ran
out of things to find; asking it to keep going, or narrowing the question, is worth trying before
trusting a thin answer.

If a single document already has the answer, upload it and ask directly: deep research is for
questions that need several sources found and weighed against each other, not for reading a file
you already have.

## Implementation details

`examples/agentic_rag/` is the site's running example for this page; nothing new was written for
it here. It offers the model two tools, `search(query)`, which returns titles and citations but no
text, and `read(cite)`, which returns one section's full text, and loops until the model stops
calling tools or a cap is hit. Withholding the text from `search` is what makes the loop genuinely
iterative rather than a slower RAG: the model has to decide, itself, which of the titles it saw are
worth opening before it can cite anything with confidence.

Every tool call and the decision to stop are `decided_by: "model"`: the model's own output picks
the query, picks which section to read, and picks when it has enough. Running a tool and returning
its result to the model are always `decided_by: "code"`, the same rule
[single agent](/gradient_ascent/techniques/single-agent/)'s example follows. `MAX_STEPS` (6) and
`MAX_TOKENS` (4000) are the hard caps; when either is hit before the model stops on its own, the
code forces one last no-tools call for a final answer, and that forced stop is `decided_by:
"code"`: the model never chose to stop, so it is not credited with a decision it did not make.

`examples/agentic_rag/run.py` (lines 43-99)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    max_steps: int = MAX_STEPS,
    max_tokens: int = MAX_TOKENS,
) -> Answer:
    del embedder  # level 5 retrieves through its tools, not a vector index
    sections = load_sections(corpus_dir)
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
    tracer.record(kind="code", decided_by="code", title="Build prompt with tool definitions", detail="search, read")

    citations: list[str] = []
    tokens_used = 0
    for _ in range(max_steps):
        completion = model.complete(messages, tools=TOOLS, max_tokens=400)
        tokens_used += completion.tokens_in + completion.tokens_out

        if not completion.tool_calls:
            tracer.record(
                kind="model",
                decided_by="model",
                title="Model stops and answers",
                detail=completion.text[:200],
                tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out,
                ms=completion.ms,
            )
            return Answer.from_text(completion.text, retrieved_sources=citations)

        calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
        tracer.record(
            kind="model",
            decided_by="model",
            title="Model calls tool(s)",
            detail=calls_desc,
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        turn, calls = assistant_turn(completion, len(messages))
        messages.append(turn)
        for call in calls:
            result_text, cites = _run_tool(call, sections)
            citations.extend(cites)
            tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
            messages.append(tool_result(call, result_text))

        if tokens_used >= max_tokens:
            final = force_final(messages, model, tracer, reason=f"token budget reached: {tokens_used} >= {max_tokens}", max_tokens=400)
            return Answer.from_text(final.text, retrieved_sources=citations)

    final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
    return Answer.from_text(final.text, retrieved_sources=citations)
```

Run it yourself:

`examples/agentic_rag/README.md` (lines 16-16)

```text
python -m examples.agentic_rag --model stub:scripted
```

Anthropic's account of building a production research agent puts a number on what a loop like this
costs, and the published sentence carries two figures, not one: "In our data, agents typically use
about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens
than chats"[2]. The 4× is the half that belongs on this page: one agent running its own
searches. The 15× is for the system of several agents Anthropic was describing, which is level 6,
not this one. Both are Anthropic's numbers, not anything measured here. Anthropic also lists what
went wrong in early versions of that system: agents "continuing when they already had sufficient
results, using overly verbose search queries, or selecting incorrect tools"[2], the same
failures a step cap and a careful stop condition exist to catch here, at a much smaller scale.

That escalation only pays off once the searches stop depending on each other: when a question
splits into independent lines of research that together need more context than one agent can
hold, see [lead agent and workers](/gradient_ascent/techniques/orchestrator-workers/) for
spreading them across several agents instead of running one longer loop.

## When you do not need this

Try [RAG](/gradient_ascent/techniques/rag/) first if one search, over one fixed set of documents,
can actually answer the question: most lookups can, and RAG costs one model call every time
instead of a number that varies with how hard the question turns out to be.

Try a fixed multi-step [workflow](/gradient_ascent/techniques/workflow-graphs/) instead of an
agent if you already know how many searches a question needs and in what order: a two-step
chain that always searches, then always searches again with a refined query, is cheaper and more
predictable than a loop when the shape of the task never actually varies.

Move up to agentic RAG once the next query genuinely depends on what the last one found, so the
number of searches cannot be fixed in advance.

## Failure modes

### The loop stops on a thin answer

- **How to notice it:** The model decides it has enough after one or two searches when the question actually needed a third, and answers confidently from an incomplete set of sources.
- **How to test for it:** Ask a question you know needs sources from more than one document and check the trace: did the model search again after the first result, or answer from what the first search alone returned?

### The loop never stops on its own

- **How to notice it:** Anthropic's own account of building a research agent describes early versions "continuing when they already had sufficient results, using overly verbose search queries, or selecting incorrect tools": cost without any added accuracy.
- **How to test for it:** Compare the number of searches a question actually needed against the number the trace shows. Extra searches that return the same information as an earlier one are this failure, not thoroughness.

### The cap cuts off a real search partway through

- **How to notice it:** The step or token cap is reached before the model was actually done, and the forced final answer reads as complete even though a source it was about to check never got opened.
- **How to test for it:** Force a low cap (examples/agentic_rag/run.py's max_steps argument) on a question that needs more searches than the cap allows, and confirm the trace records which cap stopped it rather than presenting the answer as a normal stop.

### A confident source beats a correct one

- **How to notice it:** The model settles on the first source that looks authoritative rather than the one that actually answers the question, especially when two sources disagree.
- **How to test for it:** Use a conflicting-sources question from evals/corpus/ and check whether the answer notices the conflict or just reports whichever source its search happened to rank first.

## Cost and latency

_Measured: averages over the 60-question run on Muse Glimmer 30B, a model in the Large local (about 30B) class, on one local GPU. Tokens out include the model's hidden reasoning, which it spends before answering. Holds for this model class only._

- **Tokens in, per question:** 4,037
- **Tokens out, per question:** 1,247
- **Wall time, per question:** 9.6s
- **Questions in the run:** 60

**Compared with RAG (level 2), same model.** Per question, RAG (level 2) took 512 tokens in, 687 out and 6.6s on Muse Glimmer 30B; this page took 4,037 in, 1,247 out and 9.6s, on the same 60 questions.

Most of the extra input is the loop itself: every step sends the whole conversation again,
searches and read sections included, so the prompt grows with each round. Anthropic separately
reports that in its data agents typically use about 4x more tokens than chat interactions, and
multi-agent systems about 15x more than chats; those are Anthropic's figures, not ones measured
here.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

`agentic_rag` is registered and scored on the same 60-question set as every other technique here:
exact or rubric match, citation hit rate, and the count of model-decided steps the trace carries.
Multi-hop and conflicting-source questions are where the extra cost is supposed to earn its
keep: a multi-hop question needs two sections found and used together, which single-pass RAG
often cannot do in one search, and a conflicting-source question needs the loop to notice two
retrieved sections disagree rather than stopping after the first one that looks like an answer. A
lookup question that RAG already answers in one call is the wrong place to look for agentic RAG's
advantage; if the extra cost does not show up as a better score on multi-hop and conflicting
questions specifically, it is not paying for itself.

### Measured result: Muse Glimmer 30B

**55 of 60 correct** on the site's 60-question set, run 09/23/2026 with Muse Glimmer 30B by Meta, a model in the Large local (about 30B) class. Open weights at 4-bit (Q4_K_M), run on one local GPU through Ollama. The tag is a local build of muse-glimmer:30b.

RAG (level 2) scored 46 of 60 on the same questions with the same model.

| Question kind | This page | RAG (level 2), same model |
| --- | --- | --- |
| Lookup | 12 of 12 | 12 of 12 |
| Numeric | 12 of 12 | 11 of 12 |
| Conflicting sources | 11 of 12 | 10 of 12 |
| Not in the documents | 12 of 12 | 11 of 12 |
| Multi-hop | 8 of 12 | 2 of 12 |

- **Retrieval coverage:** 90% of the sections the questions need reached the prompt.
- **Citation coverage:** 91% of the sections the questions need were cited in the answer.
- **Model-decided steps:** 225. Steps where the model chose what happened next.
- **Empty replies:** 0. **Ungraded answers:** 0.
- **Ended by a cap:** 38 of 60 questions, where the step or token budget in the code stopped the loop and forced an answer.
- **Grader:** the same model, on 28 rubric questions, the rest by exact match. Checked by a person on 09/23/2026: All 5 answers scored wrong were read, and each is wrong by its rubric or pattern. Three say outright that the section they needed was found but not yet read when the token budget forced an answer; one never reached the warranty terms; and M04 never says the DW-300 has no leak sensor, which its rubric requires.

This holds for the Large local (about 30B) class only. Not yet run: Small local (about 8B); Frontier API.

Result file: https://github.com/reedos/gradient_ascent/blob/main/evals/results/agentic_rag/ollama_muse-glimmer_30b-q4_K_M-dflash.json · recorded trace: https://github.com/reedos/gradient_ascent/blob/main/examples/agentic_rag/trace.json

On this model the extra cost paid for itself where this page said it should: most of the
multi-hop questions single-pass RAG missed, the loop answered, because after reading one section
it searched again for the next. The other kinds moved less, since RAG already did well on them.

Most loops did not stop on their own. The token budget in `run` ended most questions and forced
an answer from whatever had been read so far, and three of the answers this run got wrong say so
outright: the section they needed had turned up in a search but had not been read yet. Whether a
larger budget would answer more is not tested here. The budget is the setting that trades cost
for completeness, and a real deployment has to choose it.

To run it yourself, `python scripts/eval_run.py --example agentic_rag --model <spec> --dry`
projects the cost first; `docs/FIRST-LIVE-RUN.md` is the full sequence and the checks to read
before the score.

## Run it

**What to monitor.** Searches per question and the cap-hit rate (the share of runs that end in a forced final answer). A rising average search count with no change to the questions arriving is worth investigating before it shows up as a cost spike.

**Cost at volume.** Cost per question varies with how many searches it actually takes, unlike RAG's fixed one call. Budget from the cap, not the average, and watch the tail: a handful of hard questions that each use the full cap can cost as much as the rest of a batch combined.

**How it fails in production.** The model keeps searching past the point of diminishing returns on an easy question, or stops one search short on a hard one, and both look identical from outside unless the trace is actually read.

**What to log.** Every query the model chose, every citation returned, every section it read in full, and which cap (if any) ended the run, so a thin or wrong answer traces back to a specific search decision instead of an unexplained gap.

## Try it

1. **Use it.** Give a deep-research mode a question with two parts that live in different sources, and check its shown searches or sources: did it actually run more than one search, and does each part of the answer trace to one it ran?
2. **Build it.** Run python -m examples.agentic_rag --model stub:scripted from the repo root. The model searches, reads the one section its own search turned up, and stops: three turns, one citation. Run it again with --model stub, where the echo is never a tool call, and the loop ends on the first turn having retrieved nothing.
3. **Either lane.** Compare this page's run to RAG's: RAG's five steps are all solid (code-decided); count how many of this run's eight are dashed instead. What does the difference buy, and what does it cost?


## Sources

1. [Deep Research](https://gemini.google/overview/deep-research/) — Google (accessed 2026-09-19)
2. [How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system) — Anthropic (accessed 2026-09-19)
3. [Deep research](https://developers.openai.com/api/docs/guides/deep-research) — OpenAI (API documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
