Write a research brief with citations
Uses agentic RAG to find sources and a fixed check on every claim against the section it cites. It needs level 5 for the searching; the checking is level 3.
SourcedNeeds level 5
An analyst is asked for a short, sourced brief on a public topic: what changed in a rule this year, what a public filing actually says, how two published accounts of the same event differ. Nobody wants prose that sounds right; they want every claim traceable to something they can open. That is a research brief: search public sources, draft from them, and check that every claim still says what its source says before anyone reads it.
Example run
Optional: inspect the implementation trace
This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.
Write a research brief with citations
An agent searches and drafts; a fixed check asks one question of every citation before it ships.
The run, step by step
This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.
The topic arrives
"Brief: how has the vendor security-disclosure window changed in the last year, and what triggered the change?"
Walkthrough
The run above traces one invented topic (a disclosure window, in a made-up advisory) end to end. Nothing in it is a real rule or a real document.
The agent searches, reads a result in full before relying on it, and searches again once it notices the first pass covered one side of the question and not the other. It drafts a sentence per claim with a citation attached. Code then confirms each citation points at something this run actually opened, and hands the sentence and that one section to the checking prompt.
In the illustrated run the checker finds a sentence whose cited section covers a related but different point. Code, not the checker, decides what that verdict means: that sentence goes back with the reason attached, everything else is left untouched, and a round cap decides when a sentence ships flagged rather than fixed. The narrow question is what makes the verdict repeatable.
Nothing here pauses for a person, so the limit is worth stating plainly. The check catches a citation that does not support its sentence. It does not catch a real source nobody searched for, an editorial judgment about what belongs in the brief, or a topic where “public” turns out not to mean what the analyst assumed. Only public sources are searched and no private document enters either prompt’s context, which keeps the data handling simple. A checked draft is still not a reader-ready one. Over agentic RAG alone, the checking adds about one short call per claim.
What to measure
This recipe’s questions are not the site’s own document set, so measuring it means building a small test set for this job specifically: a handful of synthetic public-style topics with a known-correct citation for every claim a good brief would make, plus a few claims deliberately paired with the wrong section. Two things matter more than raw fluency. Citation support rate: for a sample of sentences, does the cited section actually contain the claim, checked by a person against the source. And the checker’s own catch rate against the deliberately broken citations, alongside its false-flag rate against citations that were correct all along. A checker that flags everything has a perfect catch rate and is useless. Run the check twice on one unchanged draft as well: a verdict that moves between runs is not a criterion yet. Nothing here has been scored; no result file for this recipe exists.
Variations
- Sample instead of checking every claim, once the support rate on a held-out set shows the full pass rarely finds anything.
- Move to review and debate if the errors getting through are ones no fixed question would have caught, and budget for a reviewer whose cost per brief varies.
- Move to knowledge graphs if claims start needing facts joined across many sources.
- Add human approval before the brief goes out, for a topic where a caught citation is not the only thing worth a second look.
Design choices
Why this level, and when to use another approach
Agentic RAG runs the search side: the agent decides what to search for, opens a source before relying on it, and decides for itself when it has enough to draft from. Write and check sits between the draft and the reader: a second prompt reads one drafted sentence and the section it cites, and answers one question: does that section say this?
Single-pass RAG (level 2) is not enough on its own because a brief’s claims rarely come from one search. “What changed” needs at least two queries (the current rule, and what it replaced), and the second query depends on what the first one turned up, which is the same reason the RAG page gives for climbing to agentic RAG at all.
Two parts of the checking stay below the model entirely, and should. Plain code confirms that every citation resolves to a section this run actually opened, which catches an invented identifier outright and costs nothing. It is the same check write and check’s own example runs. Code also owns the loop: how many times a sentence may come back, and what happens when that cap is reached with it still unsupported.
What is left for a model is one fixed question, asked the same way every time, and that is what settles the level. Review and debate, level 6, is a reviewer that is itself an agent: it picks what to check and goes looking with retrieval of its own. That page’s own advice is to try write and check first when the thing you would check is one fixed, testable question, and “does this section support this sentence” is exactly that question. So the highest level this job needs is level 5, and it needs it for the searching, not the checking.
The climb to a reviewing agent has a specific trigger, and it is not volume. It is the failure a fixed question cannot state: a source nobody thought to search for, or a sentence its cited section technically supports and still misleads. Catching those needs a reviewer that chooses what to look at, and the cost stops being one short call per claim and becomes however many turns it decides to take, up to a cap. A lead agent and workers is a different climb, and its condition is that the split itself cannot be written down in advance. A short brief’s split can be: if the sections are known, running them at once is parallel calls, one level down and cheaper.
Techniques this recipe uses
The highest level it needs is level 5.
Write and check
SourcedOne prompt writes, another checks, and the loop repeats until the check passes.
Find out about something across many sources and write it up
This recipe is one worked instance of a kind of job. The reasoning carries over to the others; the subject does not. See the shape.
- A literature review
- Comparing candidate parts or suppliers from their public documentation
- Due diligence on a company
- What a standard requires and how others have met it
- A market overview
Last reviewed 09/18/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page