# Write a research brief with citations

_Recipe · needs level 5_

Uses agentic RAG to find sources and a fixed check on every claim against the section it cites. It needs level 5 for the searching; the checking is level 3.


An analyst is asked for a short, sourced brief on a public topic: what changed in a rule this
year, what a public filing actually says, how two published accounts of the same event differ.
Nobody wants prose that sounds right; they want every claim traceable to something they can
open. That is a research brief: search public sources, draft from them, and check that every
claim still says what its source says before anyone reads it.

## Example run

_The web page for this technique includes an interactive step-through of Level 5 · assembled for this recipe. The same steps are described in the sections below._

## Walkthrough

The run above traces one invented topic (a disclosure window, in a made-up advisory) end to
end. Nothing in it is a real rule or a real document.

The agent searches, reads a result in full before relying on it, and searches again once it
notices the first pass covered one side of the question and not the other. It drafts a sentence
per claim with a citation attached. Code then confirms each citation points at something this run
actually opened, and hands the sentence and that one section to the checking prompt.

In the illustrated run the checker finds a sentence whose cited section covers a related but
different point. Code, not the checker, decides what that verdict means: that sentence goes back
with the reason attached, everything else is left untouched, and a round cap decides when a
sentence ships flagged rather than fixed. The narrow question is what makes the verdict
repeatable.

Nothing here pauses for a person, so the limit is worth stating plainly. The check catches a
citation that does not support its sentence. It does not catch a real source nobody searched for,
an editorial judgment about what belongs in the brief, or a topic where "public" turns out not to
mean what the analyst assumed. Only public sources are searched and no private document enters
either prompt's context, which keeps the data handling simple. A checked draft is still not a
reader-ready one. Over agentic RAG alone, the checking adds about one short call per claim.

## What to measure

This recipe's questions are not the site's own document set, so measuring it means building a
small test set for this job specifically: a handful of synthetic public-style topics with a
known-correct citation for every claim a good brief would make, plus a few claims deliberately
paired with the wrong section. Two things matter more than raw fluency. **Citation support
rate**: for a sample of sentences, does the cited section actually contain the claim, checked by
a person against the source. And the checker's own **catch rate** against the deliberately broken
citations, alongside its **false-flag rate** against citations that were correct all along. A
checker that flags everything has a perfect catch rate and is useless. Run the check twice on one
unchanged draft as well: a verdict that moves between runs is not a criterion yet. Nothing here
has been scored; no result file for this recipe exists.

## Variations

- Sample instead of checking every claim, once the support rate on a held-out set shows the full
  pass rarely finds anything.
- Move to [review and debate](/gradient_ascent/techniques/debate-review/) if the errors getting
  through are ones no fixed question would have caught, and budget for a reviewer whose cost per
  brief varies.
- Move to [knowledge graphs](/gradient_ascent/techniques/knowledge-graphs/) if claims start
  needing facts joined across many sources.
- Add [human approval](/gradient_ascent/techniques/human-in-the-loop/) before the brief goes out,
  for a topic where a caught citation is not the only thing worth a second look.
## Design choices

### Why this level, and when to use another approach

[Agentic RAG](/gradient_ascent/techniques/agentic-rag/) runs the search side: the agent decides
what to search for, opens a source before relying on it, and decides for itself when it has
enough to draft from. [Write and check](/gradient_ascent/techniques/evaluator-optimizer/) sits
between the draft and the reader: a second prompt reads one drafted sentence and the section it
cites, and answers one question: does that section say this?

Single-pass [RAG](/gradient_ascent/techniques/rag/) (level 2) is not enough on its own because a
brief's claims rarely come from one search. "What changed" needs at least two queries (the
current rule, and what it replaced), and the second query depends on what the first one turned
up, which is the same reason the RAG page gives for climbing to agentic RAG at all.

Two parts of the checking stay below the model entirely, and should. Plain code confirms that
every citation resolves to a section this run actually opened, which catches an invented
identifier outright and costs nothing. It is the same check write and check's own example runs. Code
also owns the loop: how many times a sentence may come back, and what happens when that cap is
reached with it still unsupported.

What is left for a model is one fixed question, asked the same way every time, and that is what
settles the level. [Review and debate](/gradient_ascent/techniques/debate-review/), level 6, is a
reviewer that is itself an agent: it picks what to check and goes looking with retrieval of its
own. That page's own advice is to try write and check first when the thing you would check is one
fixed, testable question, and "does this section support this sentence" is exactly that question.
So the highest level this job needs is level 5, and it needs it for the searching, not the
checking.

The climb to a reviewing agent has a specific trigger, and it is not volume. It is the failure a
fixed question cannot state: a source nobody thought to search for, or a sentence its cited
section technically supports and still misleads. Catching those needs a reviewer that chooses
what to look at, and the cost stops being one short call per claim and becomes however many turns
it decides to take, up to a cap. [A lead agent and
workers](/gradient_ascent/techniques/orchestrator-workers/) is a different climb, and its condition is that the split itself cannot be written
down in advance. A short brief's split can be: if the sections are known, running them at once is
[parallel calls](/gradient_ascent/techniques/parallelization/), one level down and cheaper.



Last reviewed 2026-09-18.
