Level 05 · Agent loops

Agentic RAG and deep research

An agent that runs its own searches until it has an answer.

Measured

How it works · conceptual architecture

Search again only when the evidence calls for it.

An agent chooses searches and document reads; the application enforces access and a search budget.

Step / conditionInformation / relationshipReturn / repeatHighlighted box: model
Search again only when the evidence calls for it.Research question → context → Model decision. Model decision → tool request → Execution gate. Execution gate → allowed → Search or read. Search or read → observation → Inspect evidence. Inspect evidence → next decision → Model decision. Model decision → final answer → Answer or report gaps. Execution gate → cannot proceed → Pause or refuse.contexttool requestallowedobservationnext decisionfinal answercannot proceedAResearch questionScope, permitted sources,constraintsBModel decisionRequest a tool or return ananswerCExecution gateArguments, permissions,budgetsDSearch or readRetrieve passages or open asourceEInspect evidenceCoverage, conflicts, sourceprovenanceFAnswer or report gapsCitations, uncertainty,unresolved factsGPause or refuseApproval needed, denied, orcapped
A
Research question

Scope, permitted sources, constraints

  • context B · Model decision
B
Model decision

Request a tool or return an answer

  • tool request C · Execution gate
  • final answer F · Answer or report gaps
C
Execution gate

Arguments, permissions, budgets

  • allowed D · Search or read
  • cannot proceed G · Pause or refuse
D
Search or read

Retrieve passages or open a source

  • observation E · Inspect evidence
E
Inspect evidence

Coverage, conflicts, source provenance

  • next decision B · Model decision
F
Answer or report gaps

Citations, uncertainty, unresolved facts

    G
    Pause or refuse

    Approval needed, denied, or capped

      More searches create opportunities to fill gaps and to introduce errors. Judge evidence coverage and claim support separately from how many steps the agent took.
      The details that change the design

      Control

      Tool output is evidence, not permission to take another action.

      Stopping

      Finish, ask for help, or stop at a step, time, or cost limit.

      Verification

      Inspect the environment and the final artifact, not just the model’s account of its work.

      A focused everyday life example. Additional perspectives appear where they provide a useful contrast.

      GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

      Agentic RAG and deep research: see it in practice.

      Retrieval where an agent chooses follow-up searches and reads until it can answer or must stop.

      What you’ll walk through

      Follow an investigation in which the model chooses follow-up searches as evidence arrives. Inspect how each new source changes the question and whether further searching is still useful.

      The task in this version

      Investigate whether this DW-480 water-damage repair is covered.

      What you’ll learn to check

      Search trajectory, evidence accumulated per step, citations, conflicts, and a stop/abstain case.

      The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

      Everyday lifeAn authored case with its own evidence, changed condition, and decision.
      The task in this example

      Investigate whether this DW-480 water-damage repair is covered.

      Authored case. Select any record below; nothing is sent to a model.
      FOLLOW THE EXAMPLE1 / 6
      Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
      THE VISIBLE WORKStarting evidence
      Input record
      AUTHORED TEACHING RECORD · NOT A LIVE RUN
      Fictional manual v3 gives two-year coverage. Service notes link addendum A3, which excludes water damage for the DW-480. The first search returns only duration; a later lookup can retrieve A3.

      What changed: Establish the facts supplied for this version of the task.

      WHY THIS MATTERS

      What this case assumes

      Relevant evidence may span sources or contain contradictions. Search autonomy cannot compensate for missing access or a collection that lacks the answer.

      1 / 6

      Apply this to your project

      Describe your task to your own model and use Agentic RAG and deep research as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

      Go deeper: practical guidance, failure modes, and implementation

      Agentic RAG puts the search loop itself under the model’s control. RAG always searches once and asks the model once. Agentic RAG instead offers the model a search tool it can call as many times as it decides it needs, lets it read what comes back, and lets it decide whether to search again, read more closely, or stop and answer: the single-agent loop aimed at retrieval.

      Google describes its own version this way: at each step, “the model has to ground itself on all information gathered so far, then identify missing information and discrepancies it wants to explore”, continuing until “the model determines enough information has been gathered”[1]. OpenAI’s deep research models work the same way: “agentic” systems that “conduct multi-step research” and return a listing of every search made along the way[3]. In both cases the model chooses the next query and when to stop; your code still runs every search and can cut the loop off with a hard cap regardless of what the model would have done next.

      This page is measured: the cost and the score under How to Evaluate It come from a recorded run of this example on a real model, beside the RAG page’s run on the same model and questions, and hold for that model’s class. The step-through just below is still a scripted illustration, and source references do not establish the correctness of every implementation or outcome.

      Optional: inspect the implementation trace

      This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

      Agentic RAG

      The model runs its own searches and decides when it has enough.

      Level 5 · Agent loops
      QuestionQuestionMODELpicks the next steppicks the next stepTOOLsearch(query)search(query)TOOLread(doc)read(doc)AnswerAnswer
      0of 1 step so far chosen by the model
      your code chose this stepthe model chose this step

      The run, step by step

      This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

      STEP 01 / 08Your code chose

      The question arrives

      "How long is the warranty on the DW-480,
      and what voids it?"
      0 tokens · 0 ms

      Practical guidance

      This is what “Deep Research” or “DeepSearch” does in a chat app: ChatGPT, Claude, Gemini, Perplexity and Grok DeepSearch all ship a mode like it, usually a toggle or a separate button next to the ordinary send button. Reach for it when a question has more than one part living in different places, or when two sources might disagree and you want that checked rather than guessed past: “Compare what our returns policy says about damaged items against what the shipping carrier’s own terms say, and tell me where they conflict.” A question one search can already answer does not need it, and costs more here for nothing extra.

      Once it is running, Google’s own description of the mechanism is specific: the model “oversees the execution of” a research plan, and at each step “the model reasons over information available to decide its next move”[1]. That means the searches it runs are worth reading, not just the report at the end. Open the sources or search-steps panel most of these products show: OpenAI’s deep research output “will contain a listing of web search calls, code interpreter calls, and remote MCP calls made to get to the answer,”[3] so the individual searches are visible, not hidden inside the final prose.

      Check a claim in the report against a source it actually shows, the way you would check a citation in plain RAG search: does the linked source really say what the report claims, and does the report ever cite a source it does not appear to have opened at all.

      There is a cap you do not see. OpenAI documents a setting a developer can turn to control the total number of tool calls a deep-research run may make before returning a result[3], and every product like it has some version of the same limit. A report that reads thinner than the question deserved, especially one with several parts, may be a run that hit its cap rather than one that ran out of things to find; asking it to keep going, or narrowing the question, is worth trying before trusting a thin answer.

      If a single document already has the answer, upload it and ask directly: deep research is for questions that need several sources found and weighed against each other, not for reading a file you already have.

      Implementation details

      examples/agentic_rag/ is the site’s running example for this page; nothing new was written for it here. It offers the model two tools, search(query), which returns titles and citations but no text, and read(cite), which returns one section’s full text, and loops until the model stops calling tools or a cap is hit. Withholding the text from search is what makes the loop genuinely iterative rather than a slower RAG: the model has to decide, itself, which of the titles it saw are worth opening before it can cite anything with confidence.

      Every tool call and the decision to stop are decided_by: "model": the model’s own output picks the query, picks which section to read, and picks when it has enough. Running a tool and returning its result to the model are always decided_by: "code", the same rule single agent’s example follows. MAX_STEPS (6) and MAX_TOKENS (4000) are the hard caps; when either is hit before the model stops on its own, the code forces one last no-tools call for a final answer, and that forced stop is decided_by: "code": the model never chose to stop, so it is not credited with a decision it did not make.

      examples/agentic_rag/run.py · lines 43–99
      def run(
          question: str,
          model: Model,
          embedder: Embedder | None,
          tracer: Tracer,
          *,
          corpus_dir: Path = DEFAULT_CORPUS_DIR,
          max_steps: int = MAX_STEPS,
          max_tokens: int = MAX_TOKENS,
      ) -> Answer:
          del embedder  # level 5 retrieves through its tools, not a vector index
          sections = load_sections(corpus_dir)
          messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=question)]
          tracer.record(kind="code", decided_by="code", title="Build prompt with tool definitions", detail="search, read")
      
          citations: list[str] = []
          tokens_used = 0
          for _ in range(max_steps):
              completion = model.complete(messages, tools=TOOLS, max_tokens=400)
              tokens_used += completion.tokens_in + completion.tokens_out
      
              if not completion.tool_calls:
                  tracer.record(
                      kind="model",
                      decided_by="model",
                      title="Model stops and answers",
                      detail=completion.text[:200],
                      tokens_in=completion.tokens_in,
                      tokens_out=completion.tokens_out,
                      ms=completion.ms,
                  )
                  return Answer.from_text(completion.text, retrieved_sources=citations)
      
              calls_desc = ", ".join(f"{c.name}({json.dumps(c.arguments, sort_keys=True)})" for c in completion.tool_calls)
              tracer.record(
                  kind="model",
                  decided_by="model",
                  title="Model calls tool(s)",
                  detail=calls_desc,
                  tokens_in=completion.tokens_in,
                  tokens_out=completion.tokens_out,
                  ms=completion.ms,
              )
              turn, calls = assistant_turn(completion, len(messages))
              messages.append(turn)
              for call in calls:
                  result_text, cites = _run_tool(call, sections)
                  citations.extend(cites)
                  tracer.record(kind="code", decided_by="code", title=f"Run tool: {call.name}", detail=result_text[:200])
                  messages.append(tool_result(call, result_text))
      
              if tokens_used >= max_tokens:
                  final = force_final(messages, model, tracer, reason=f"token budget reached: {tokens_used} >= {max_tokens}", max_tokens=400)
                  return Answer.from_text(final.text, retrieved_sources=citations)
      
          final = force_final(messages, model, tracer, reason=f"step cap reached: {max_steps} steps", max_tokens=400)
          return Answer.from_text(final.text, retrieved_sources=citations)

      Run it yourself:

      examples/agentic_rag/README.md · lines 16–16
      python -m examples.agentic_rag --model stub:scripted

      Anthropic’s account of building a production research agent puts a number on what a loop like this costs, and the published sentence carries two figures, not one: “In our data, agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats”[2]. The 4× is the half that belongs on this page: one agent running its own searches. The 15× is for the system of several agents Anthropic was describing, which is level 6, not this one. Both are Anthropic’s numbers, not anything measured here. Anthropic also lists what went wrong in early versions of that system: agents “continuing when they already had sufficient results, using overly verbose search queries, or selecting incorrect tools”[2], the same failures a step cap and a careful stop condition exist to catch here, at a much smaller scale.

      That escalation only pays off once the searches stop depending on each other: when a question splits into independent lines of research that together need more context than one agent can hold, see lead agent and workers for spreading them across several agents instead of running one longer loop.

      When you do not need this

      Try RAG first if one search, over one fixed set of documents, can actually answer the question: most lookups can, and RAG costs one model call every time instead of a number that varies with how hard the question turns out to be.

      Try a fixed multi-step workflow instead of an agent if you already know how many searches a question needs and in what order: a two-step chain that always searches, then always searches again with a refined query, is cheaper and more predictable than a loop when the shape of the task never actually varies.

      Move up to agentic RAG once the next query genuinely depends on what the last one found, so the number of searches cannot be fixed in advance.

      Failure modes

      The loop stops on a thin answer

      How to notice it
      The model decides it has enough after one or two searches when the question actually needed a third, and answers confidently from an incomplete set of sources.
      How to test for it
      Ask a question you know needs sources from more than one document and check the trace: did the model search again after the first result, or answer from what the first search alone returned?

      The loop never stops on its own

      How to notice it
      Anthropic's own account of building a research agent describes early versions "continuing when they already had sufficient results, using overly verbose search queries, or selecting incorrect tools": cost without any added accuracy.
      How to test for it
      Compare the number of searches a question actually needed against the number the trace shows. Extra searches that return the same information as an earlier one are this failure, not thoroughness.

      The cap cuts off a real search partway through

      How to notice it
      The step or token cap is reached before the model was actually done, and the forced final answer reads as complete even though a source it was about to check never got opened.
      How to test for it
      Force a low cap (examples/agentic_rag/run.py's max_steps argument) on a question that needs more searches than the cap allows, and confirm the trace records which cap stopped it rather than presenting the answer as a normal stop.

      A confident source beats a correct one

      How to notice it
      The model settles on the first source that looks authoritative rather than the one that actually answers the question, especially when two sources disagree.
      How to test for it
      Use a conflicting-sources question from evals/corpus/ and check whether the answer notices the conflict or just reports whichever source its search happened to rank first.

      Cost and latency

      Measured: averages over the 60-question run on Muse Glimmer 30B, a model in the Large local (about 30B) class, on one local GPU. Tokens out include the model's hidden reasoning, which it spends before answering. Holds for this model class only.

      4,037Tokens in, per question
      1,247Tokens out, per question
      9.6sWall time, per question
      60Questions in the run
      Compared with RAG (level 2), same modelPer question, RAG (level 2) took 512 tokens in, 687 out and 6.6s on Muse Glimmer 30B; this page took 4,037 in, 1,247 out and 9.6s, on the same 60 questions.

      Most of the extra input is the loop itself: every step sends the whole conversation again, searches and read sections included, so the prompt grows with each round. Anthropic separately reports that in its data agents typically use about 4x more tokens than chat interactions, and multi-agent systems about 15x more than chats; those are Anthropic’s figures, not ones measured here.

      How to Evaluate It

      60 questionslookupmulti-hopnumericunanswerableconflicting sources

      agentic_rag is registered and scored on the same 60-question set as every other technique here: exact or rubric match, citation hit rate, and the count of model-decided steps the trace carries. Multi-hop and conflicting-source questions are where the extra cost is supposed to earn its keep: a multi-hop question needs two sections found and used together, which single-pass RAG often cannot do in one search, and a conflicting-source question needs the loop to notice two retrieved sections disagree rather than stopping after the first one that looks like an answer. A lookup question that RAG already answers in one call is the wrong place to look for agentic RAG’s advantage; if the extra cost does not show up as a better score on multi-hop and conflicting questions specifically, it is not paying for itself.

      Measured result

      55 of 60 correct

      RAG (level 2): 46 of 60 on the same questions, same model.

      On the site's 60-question set, run 09/23/2026 with Muse Glimmer 30B by Meta, a model in the Large local (about 30B) class. Open weights at 4-bit (Q4_K_M), run on one local GPU through Ollama. The tag is a local build of muse-glimmer:30b.

      Correct answers by question kind
      Question kindThis pageRAG (level 2)Share
      Lookup12 of 1212 of 12
      Numeric12 of 1211 of 12
      Conflicting sources11 of 1210 of 12
      Not in the documents12 of 1211 of 12
      Multi-hop8 of 122 of 12
      Retrieval coverage
      90%

      of the sections the questions need reached the prompt

      Citation coverage
      91%

      of the sections the questions need were cited in the answer

      Model-decided steps
      225

      steps where the model chose what happened next

      Empty or ungraded
      0

      answers the score could not read

      Ended by a cap
      38

      questions where the code's step or token budget stopped the loop

      Graded by the same model on 28 rubric questions, the rest by exact match. Checked by a person on 09/23/2026: All 5 answers scored wrong were read, and each is wrong by its rubric or pattern. Three say outright that the section they needed was found but not yet read when the token budget forced an answer; one never reached the warranty terms; and M04 never says the DW-300 has no leak sensor, which its rubric requires. This holds for the Large local (about 30B) class only. Not yet run: Small local (about 8B); Frontier API.

      On this model the extra cost paid for itself where this page said it should: most of the multi-hop questions single-pass RAG missed, the loop answered, because after reading one section it searched again for the next. The other kinds moved less, since RAG already did well on them.

      Most loops did not stop on their own. The token budget in run ended most questions and forced an answer from whatever had been read so far, and three of the answers this run got wrong say so outright: the section they needed had turned up in a search but had not been read yet. Whether a larger budget would answer more is not tested here. The budget is the setting that trades cost for completeness, and a real deployment has to choose it.

      To run it yourself, python scripts/eval_run.py --example agentic_rag --model <spec> --dry projects the cost first; docs/FIRST-LIVE-RUN.md is the full sequence and the checks to read before the score.

      Run it

      What to monitor

      Searches per question and the cap-hit rate (the share of runs that end in a forced final answer). A rising average search count with no change to the questions arriving is worth investigating before it shows up as a cost spike.

      Cost at volume

      Cost per question varies with how many searches it actually takes, unlike RAG's fixed one call. Budget from the cap, not the average, and watch the tail: a handful of hard questions that each use the full cap can cost as much as the rest of a batch combined.

      How it fails in production

      The model keeps searching past the point of diminishing returns on an easy question, or stops one search short on a hard one, and both look identical from outside unless the trace is actually read.

      What to log

      Every query the model chose, every citation returned, every section it read in full, and which cap (if any) ended the run, so a thin or wrong answer traces back to a specific search decision instead of an unexplained gap.

      Try it

      1. Use it

        Give a deep-research mode a question with two parts that live in different sources, and check its shown searches or sources: did it actually run more than one search, and does each part of the answer trace to one it ran?

      2. Build it

        Run python -m examples.agentic_rag --model stub:scripted from the repo root. The model searches, reads the one section its own search turned up, and stops: three turns, one citation. Run it again with --model stub, where the echo is never a tool call, and the loop ends on the first turn having retrieved nothing.

      3. Either lane

        Compare this page's run to RAG's: RAG's five steps are all solid (code-decided); count how many of this run's eight are dashed instead. What does the difference buy, and what does it cost?

      How it connects

      Before, after and instead of this

      Move up when

      • Lead agent and workersThe searches do not depend on each other and there are more of them than one agent's context can carry.

      Often used with

      Decoded in

      Optional: products, tools, and models

      6 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

      In practice

      Resolve conflicting specifications

      An agent finds a manual, notices a later bulletin, searches again, and reconciles the evidence in its answer.

      Out there

      Named products, tools and models

      Products5
      • ChatGPT deep researchOpenAI · research agent
      • Claude ResearchAnthropic · research agent
      • Gemini Deep ResearchGoogle · research agent
      • Grok DeepSearchSpaceXAI · research agent
      • Perplexity Deep ResearchPerplexity · research agent
      Tools1
      • LlamaIndexLlamaIndex · retrieval framework

      Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

      Where this comes from

      Primary sources

      1. Deep Research · Google (accessed 09/19/2026)
      2. How we built our multi-agent research system · Anthropic (accessed 09/19/2026)
      3. Deep research · OpenAI (API documentation) (accessed 09/19/2026)

      Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page