Primary sources
- Deep Research · Google (accessed 09/19/2026)
- How we built our multi-agent research system · Anthropic (accessed 09/19/2026)
- Deep research · OpenAI (API documentation) (accessed 09/19/2026)
An agent that runs its own searches until it has an answer.
Measured
How it works · conceptual architecture
An agent chooses searches and document reads; the application enforces access and a search budget.
Scope, permitted sources, constraints
Request a tool or return an answer
Arguments, permissions, budgets
Retrieve passages or open a source
Coverage, conflicts, source provenance
Citations, uncertainty, unresolved facts
Approval needed, denied, or capped
Tool output is evidence, not permission to take another action.
Finish, ask for help, or stop at a step, time, or cost limit.
Inspect the environment and the final artifact, not just the model’s account of its work.
A focused everyday life example. Additional perspectives appear where they provide a useful contrast.
Retrieval where an agent chooses follow-up searches and reads until it can answer or must stop.
Follow an investigation in which the model chooses follow-up searches as evidence arrives. Inspect how each new source changes the question and whether further searching is still useful.
Investigate whether this DW-480 water-damage repair is covered.
Search trajectory, evidence accumulated per step, citations, conflicts, and a stop/abstain case.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Investigate whether this DW-480 water-damage repair is covered.
Authored case. Select any record below; nothing is sent to a model.What changed: Establish the facts supplied for this version of the task.
Relevant evidence may span sources or contain contradictions. Search autonomy cannot compensate for missing access or a collection that lacks the answer.
Describe your task to your own model and use Agentic RAG and deep research as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Agentic RAG puts the search loop itself under the model’s control. RAG always searches once and asks the model once. Agentic RAG instead offers the model a search tool it can call as many times as it decides it needs, lets it read what comes back, and lets it decide whether to search again, read more closely, or stop and answer: the single-agent loop aimed at retrieval.
Google describes its own version this way: at each step, “the model has to ground itself on all information gathered so far, then identify missing information and discrepancies it wants to explore”, continuing until “the model determines enough information has been gathered”[1]. OpenAI’s deep research models work the same way: “agentic” systems that “conduct multi-step research” and return a listing of every search made along the way[3]. In both cases the model chooses the next query and when to stop; your code still runs every search and can cut the loop off with a hard cap regardless of what the model would have done next.
This page is measured: the cost and the score under How to Evaluate It come from a recorded run of this example on a real model, beside the RAG page’s run on the same model and questions, and hold for that model’s class. The step-through just below is still a scripted illustration, and source references do not establish the correctness of every implementation or outcome.
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
The model runs its own searches and decides when it has enough.
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
"How long is the warranty on the DW-480, and what voids it?"
This is what “Deep Research” or “DeepSearch” does in a chat app: ChatGPT, Claude, Gemini, Perplexity and Grok DeepSearch all ship a mode like it, usually a toggle or a separate button next to the ordinary send button. Reach for it when a question has more than one part living in different places, or when two sources might disagree and you want that checked rather than guessed past: “Compare what our returns policy says about damaged items against what the shipping carrier’s own terms say, and tell me where they conflict.” A question one search can already answer does not need it, and costs more here for nothing extra.
Once it is running, Google’s own description of the mechanism is specific: the model “oversees the execution of” a research plan, and at each step “the model reasons over information available to decide its next move”[1]. That means the searches it runs are worth reading, not just the report at the end. Open the sources or search-steps panel most of these products show: OpenAI’s deep research output “will contain a listing of web search calls, code interpreter calls, and remote MCP calls made to get to the answer,”[3] so the individual searches are visible, not hidden inside the final prose.
Check a claim in the report against a source it actually shows, the way you would check a citation in plain RAG search: does the linked source really say what the report claims, and does the report ever cite a source it does not appear to have opened at all.
There is a cap you do not see. OpenAI documents a setting a developer can turn to control the total number of tool calls a deep-research run may make before returning a result[3], and every product like it has some version of the same limit. A report that reads thinner than the question deserved, especially one with several parts, may be a run that hit its cap rather than one that ran out of things to find; asking it to keep going, or narrowing the question, is worth trying before trusting a thin answer.
If a single document already has the answer, upload it and ask directly: deep research is for questions that need several sources found and weighed against each other, not for reading a file you already have.
examples/agentic_rag/ is the site’s running example for this page; nothing new was written for
it here. It offers the model two tools, search(query), which returns titles and citations but no
text, and read(cite), which returns one section’s full text, and loops until the model stops
calling tools or a cap is hit. Withholding the text from search is what makes the loop genuinely
iterative rather than a slower RAG: the model has to decide, itself, which of the titles it saw are
worth opening before it can cite anything with confidence.
Every tool call and the decision to stop are decided_by: "model": the model’s own output picks
the query, picks which section to read, and picks when it has enough. Running a tool and returning
its result to the model are always decided_by: "code", the same rule
single agent’s example follows. MAX_STEPS (6) and
MAX_TOKENS (4000) are the hard caps; when either is hit before the model stops on its own, the
code forces one last no-tools call for a final answer, and that forced stop is decided_by: "code": the model never chose to stop, so it is not credited with a decision it did not make.
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
max_steps: int = MAX_STEPS,
max_tokens: int = MAX_TOKENS,
) -> Answer:
del embedder # level 5 retrieves through its tools, not a vector index
sections = load_sections(corpus_dir)
messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
tracer.record(kind="code", decided_by="code", title="Build prompt with tool definitions", detail="search, read")
citations: list[str] = []
tokens_used = 0
for _ in range(max_steps):
completion = model.complete(messages, tools=TOOLS, max_tokens=400)
tokens_used += completion.tokens_in + completion.tokens_out
if not completion.tool_calls:
tracer.record(
kind="model",
decided_by="model",
title="Model stops and answers",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
return Answer.from_text(completion.text, retrieved_sources=citations)
calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
tracer.record(
kind="model",
decided_by="model",
title="Model calls tool(s)",
detail=calls_desc,
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
turn, calls = assistant_turn(completion, len(messages))
messages.append(turn)
for call in calls:
result_text, cites = _run_tool(call, sections)
citations.extend(cites)
tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
messages.append(tool_result(call, result_text))
if tokens_used >= max_tokens:
final = force_final(messages, model, tracer, reason=f"token budget reached: {tokens_used} >= {max_tokens}", max_tokens=400)
return Answer.from_text(final.text, retrieved_sources=citations)
final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
return Answer.from_text(final.text, retrieved_sources=citations)Run it yourself:
python -m examples.agentic_rag --model stub:scriptedAnthropic’s account of building a production research agent puts a number on what a loop like this costs, and the published sentence carries two figures, not one: “In our data, agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats”[2]. The 4× is the half that belongs on this page: one agent running its own searches. The 15× is for the system of several agents Anthropic was describing, which is level 6, not this one. Both are Anthropic’s numbers, not anything measured here. Anthropic also lists what went wrong in early versions of that system: agents “continuing when they already had sufficient results, using overly verbose search queries, or selecting incorrect tools”[2], the same failures a step cap and a careful stop condition exist to catch here, at a much smaller scale.
That escalation only pays off once the searches stop depending on each other: when a question splits into independent lines of research that together need more context than one agent can hold, see lead agent and workers for spreading them across several agents instead of running one longer loop.
Try RAG first if one search, over one fixed set of documents, can actually answer the question: most lookups can, and RAG costs one model call every time instead of a number that varies with how hard the question turns out to be.
Try a fixed multi-step workflow instead of an agent if you already know how many searches a question needs and in what order: a two-step chain that always searches, then always searches again with a refined query, is cheaper and more predictable than a loop when the shape of the task never actually varies.
Move up to agentic RAG once the next query genuinely depends on what the last one found, so the number of searches cannot be fixed in advance.
Measured: averages over the 60-question run on Muse Glimmer 30B, a model in the Large local (about 30B) class, on one local GPU. Tokens out include the model's hidden reasoning, which it spends before answering. Holds for this model class only.
Most of the extra input is the loop itself: every step sends the whole conversation again, searches and read sections included, so the prompt grows with each round. Anthropic separately reports that in its data agents typically use about 4x more tokens than chat interactions, and multi-agent systems about 15x more than chats; those are Anthropic’s figures, not ones measured here.
agentic_rag is registered and scored on the same 60-question set as every other technique here:
exact or rubric match, citation hit rate, and the count of model-decided steps the trace carries.
Multi-hop and conflicting-source questions are where the extra cost is supposed to earn its
keep: a multi-hop question needs two sections found and used together, which single-pass RAG
often cannot do in one search, and a conflicting-source question needs the loop to notice two
retrieved sections disagree rather than stopping after the first one that looks like an answer. A
lookup question that RAG already answers in one call is the wrong place to look for agentic RAG’s
advantage; if the extra cost does not show up as a better score on multi-hop and conflicting
questions specifically, it is not paying for itself.
RAG (level 2): 46 of 60 on the same questions, same model.
On the site's 60-question set, run 09/23/2026 with Muse Glimmer 30B by Meta, a model in the Large local (about 30B) class. Open weights at 4-bit (Q4_K_M), run on one local GPU through Ollama. The tag is a local build of muse-glimmer:30b.
| Question kind | This page | RAG (level 2) | Share |
|---|---|---|---|
| Lookup | 12 of 12 | 12 of 12 | |
| Numeric | 12 of 12 | 11 of 12 | |
| Conflicting sources | 11 of 12 | 10 of 12 | |
| Not in the documents | 12 of 12 | 11 of 12 | |
| Multi-hop | 8 of 12 | 2 of 12 |
of the sections the questions need reached the prompt
of the sections the questions need were cited in the answer
steps where the model chose what happened next
answers the score could not read
questions where the code's step or token budget stopped the loop
Graded by the same model on 28 rubric questions, the rest by exact match. Checked by a person on 09/23/2026: All 5 answers scored wrong were read, and each is wrong by its rubric or pattern. Three say outright that the section they needed was found but not yet read when the token budget forced an answer; one never reached the warranty terms; and M04 never says the DW-300 has no leak sensor, which its rubric requires. This holds for the Large local (about 30B) class only. Not yet run: Small local (about 8B); Frontier API.
On this model the extra cost paid for itself where this page said it should: most of the multi-hop questions single-pass RAG missed, the loop answered, because after reading one section it searched again for the next. The other kinds moved less, since RAG already did well on them.
Most loops did not stop on their own. The token budget in run ended most questions and forced
an answer from whatever had been read so far, and three of the answers this run got wrong say so
outright: the section they needed had turned up in a search but had not been read yet. Whether a
larger budget would answer more is not tested here. The budget is the setting that trades cost
for completeness, and a real deployment has to choose it.
To run it yourself, python scripts/eval_run.py --example agentic_rag --model <spec> --dry
projects the cost first; docs/FIRST-LIVE-RUN.md is the full sequence and the checks to read
before the score.
Searches per question and the cap-hit rate (the share of runs that end in a forced final answer). A rising average search count with no change to the questions arriving is worth investigating before it shows up as a cost spike.
Cost per question varies with how many searches it actually takes, unlike RAG's fixed one call. Budget from the cap, not the average, and watch the tail: a handful of hard questions that each use the full cap can cost as much as the rest of a batch combined.
The model keeps searching past the point of diminishing returns on an easy question, or stops one search short on a hard one, and both look identical from outside unless the trace is actually read.
Every query the model chose, every citation returned, every section it read in full, and which cap (if any) ended the run, so a thin or wrong answer traces back to a specific search decision instead of an unexplained gap.
Give a deep-research mode a question with two parts that live in different sources, and check its shown searches or sources: did it actually run more than one search, and does each part of the answer trace to one it ran?
Run python -m examples.agentic_rag --model stub:scripted from the repo root. The model searches, reads the one section its own search turned up, and stops: three turns, one citation. Run it again with --model stub, where the echo is never a tool call, and the loop ends on the first turn having retrieved nothing.
Compare this page's run to RAG's: RAG's five steps are all solid (code-decided); count how many of this run's eight are dashed instead. What does the difference buy, and what does it cost?
6 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.
Research agent
Maker’s documentation Checked 09/18/2026Research agent
Maker’s documentation Checked 09/18/2026Research agent
Maker’s documentation Checked 09/18/2026Research agent
Maker’s documentation Checked 09/18/2026Research agent
Maker’s documentation Checked 09/18/2026Retrieval framework
Maker’s documentation Checked 09/18/2026An agent finds a manual, notices a later bulletin, searches again, and reconciles the evidence in its answer.
Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page