Level 02 · Added context

Retrieval-augmented generation (RAG)

Searching your documents and giving the results to the model.

Measured

How it works · conceptual architecture

Two paths meet at retrieval.

Indexing prepares the sources. A query selects evidence for this answer.

Step / conditionInformation / relationshipHighlighted box: model
Two paths meet at retrieval.Source documents → indexing → Prepare the index. Prepare the index → searchable sources → Retrieve + select. Question → query → Retrieve + select. Retrieve + select → evidence packet → Generate an answer. Generate an answer → draft → Check or abstain.indexingsearchable sourcesqueryevidence packetdraftASource documentsVersioned text with accessrulesBPrepare the indexSplit, preserve IDs, index thecontentCQuestionWhat the user needs to knowDRetrieve + selectLexical, vector, hybrid;optional rerankEGenerate an answerQuestion + selected evidenceFCheck or abstainSupported claims and validcitations
A
Source documents

Versioned text with access rules

  • indexing B · Prepare the index
B
Prepare the index

Split, preserve IDs, index the content

  • searchable sources D · Retrieve + select
C
Question

What the user needs to know

  • query D · Retrieve + select
D
Retrieve + select

Lexical, vector, hybrid; optional rerank

  • evidence packet E · Generate an answer
E
Generate an answer

Question + selected evidence

  • draft F · Check or abstain
F
Check or abstain

Supported claims and valid citations

    RAG is retrieval-augmented generation, not a guarantee of truth. A missing answer may be a retrieval failure, a source gap, or a generation error; those need different fixes.
    The details that change the design

    Before retrieval

    Apply document permissions. Keep source IDs and revision metadata.

    Before answering

    Check whether the selected evidence covers every part of the question.

    After answering

    A real citation ID can still support the wrong claim. Check entailment as well as existence.

    CHOOSE YOUR PERSPECTIVE

    Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

    GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

    Retrieval-augmented generation (RAG): see it in practice.

    Retrieving external information and placing it in the model's context to support an answer.

    What you’ll walk through

    Follow a question through source retrieval into a grounded answer. Inspect whether the retrieved passages actually support the response, including exceptions and unanswered parts.

    The task in this version

    Is water damage covered by the DW-480 warranty?

    What you’ll learn to check

    Visible query, retrieved passages, grounded answer, source links, and an abstention when evidence is insufficient.

    The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

    Everyday lifeAn authored case with its own evidence, changed condition, and decision.
    The task in this example

    Is water damage covered by the DW-480 warranty?

    Authored case. Select any record below; nothing is sent to a model.
    FOLLOW THE EXAMPLE1 / 6
    Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
    THE VISIBLE WORKStarting evidence
    Source packet
    AUTHORED TEACHING RECORD · NOT A LIVE RUN
    Manual v3 §2: Coverage lasts two years. Manual v3 §4: Water damage is excluded. Question: Is water damage covered for the DW-480? Provenance: fictional manual passages supplied for this example.

    What changed: Duration and exclusions are distinct pieces of evidence.

    WHY THIS MATTERS

    What this case assumes

    The collection may be incomplete or outdated. A citation is useful only when its passage supports the associated claim.

    1 / 6

    Apply this to your project

    Describe your task to your own model and use Retrieval-augmented generation (RAG) as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

    Go deeper: practical guidance, failure modes, and implementation

    Retrieval-augmented generation, or RAG, supplies retrieved information to a model when it generates an answer. Retrieval can use keywords, embeddings, or both; the sources may be your documents or another searchable collection. The 2020 RAG paper describes combining retrieval with generation so the model can draw on external information[1].

    This page teaches a simple, fixed retrieval pipeline: split documents into passages, embed them, retrieve a few relevant passages for the question, and send those passages to one model call. The instruction is to answer from the evidence, but the model can still make unsupported claims. The fixed number of passages and single generation call are choices in this example, not rules that define all RAG systems. Other implementations rerank, rewrite queries, retrieve repeatedly, or combine evidence across documents.

    The simple pipeline sits at level 2 because code determines how context is selected. If a model chooses successive searches, this site calls that agentic RAG. Both approaches augment generation with retrieved information.

    This page is measured: the cost and the score under How to Evaluate It come from a recorded run of this example on a real model, and hold for that model’s class. The step-through just below is still a scripted illustration, and source references do not establish the correctness of every implementation or outcome.

    Optional: inspect the implementation trace

    This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

    Retrieval-augmented generation (RAG)

    Search the documents, give the closest passages to the model, and ask once.

    Level 2 · Added context
    QuestionQuestionEmbed the questionEmbed the questionVector storeVector storeBuild the promptBuild the promptMODELanswers onceanswers onceAnswerAnswerQuestionQuestionEmbed the questionEmbed the questionVector storeVector storeBuild the promptBuild the promptMODELanswers onceanswers onceAnswerAnswer
    0of 1 step so far chosen by the model
    your code chose this stepthe model chose this step

    The run, step by step

    This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

    STEP 01 / 05Your code chose

    The question arrives

    "How long is the warranty on the DW-480,
    and what voids it?"
    0 tokens · 0 ms

    Practical guidance

    Upload files to a chat app’s project or file feature and ask about them. The product may put the files directly into context, retrieve selected passages, or combine both approaches. Uploading a file alone does not tell you which method it uses; check the product documentation.

    Ask questions a handful of passages can answer on their own. “What does the warranty cover” is a focused lookup. “Summarize every change across all our contracts this year” requires much broader coverage: a few highly ranked passages may omit important changes. For that task, check whether the product can systematically cover the collection. Narrower questions and an explicit document checklist make omissions easier to notice.

    Anthropic’s own documentation describes this directly: once a project’s uploaded files approach what the context window can hold, Claude switches into what Anthropic calls RAG mode. Anthropic says that expands how much a project can hold by up to ten times while maintaining response quality[2]; that is the maker’s claim, not a number this site has measured. OpenAI documents the same pattern for its own file search feature: it retrieves passages from uploaded files by keyword and meaning together, and returns an answer with citations to the files it used[3].

    Check the citations every time the product shows them. Open the source it names and confirm the sentence it cites is actually there. A citation that does not obviously support the sentence beside it, or an answer with none at all, is unconfirmed. It does not by itself tell you whether retrieval failed, the model ignored evidence, or the interface omitted the citation.

    RAG can combine facts from multiple documents when the necessary evidence is retrieved and used correctly. A single search can return passages from several files, but questions whose second lookup depends on the first answer may need query decomposition or repeated retrieval. Inspect all required sources and compare a fixed pipeline with an agentic one on the same task.

    Implementation details

    The minimal version of RAG is four fixed steps: chunk the documents, embed the question and every chunk, keep the top few chunks by similarity, and ask the model once with those chunks as its only sources.

    Chunking splits documents into pieces small enough to embed and retrieve individually. The example below chunks by section, since the synthetic document set already has numbered sections; a real document set usually needs its own splitter, tuned so a chunk holds one complete idea rather than cutting a sentence or a table row in half.

    Embedding turns text into a vector, a fixed-length list of numbers, using a model trained so that texts with similar meaning get vectors that point in similar directions; the mechanism, and the search built on it, is embeddings and search. The code below is written against an Embedder interface with two implementations: a deterministic stub for tests, and a real embedding model behind the same interface, so the retrieval logic never has to know which one is running.

    Retrieval scores every chunk’s embedding against the question’s embedding by cosine similarity (how closely the two vectors point in the same direction) and keeps the top k, four by default. _cosine below returns a plain dot product rather than a full cosine, because both embedders return unit vectors, for which the two are the same number. This is the one place a real system usually adds more: a second, more expensive reranking pass over a larger first cut of candidates, scored by a model trained for that job. Cohere and Jina AI both sell one. The example skips reranking to keep the pipeline to four fixed steps.

    Prompt assembly numbers every retrieved chunk, includes its citation (file#section), and the system prompt instructs the model to answer using only those sources and to name which ones it used. Citations are then parsed back out of the model’s answer with a regular expression, so the calling code always knows, in a form it can check automatically, which sources actually contributed to the answer.

    The same four steps work over an engineer’s own documents, not just reference text. Orbeck Power Systems’ SRB-5030 datasheet states one maximum input voltage; a later engineering change notice supersedes it for two of the board’s three revisions, over a capacitor derating rule, and the datasheet is never reissued to say so. One search that retrieves both documents returns an answer with citations a reader can check by hand, in a production test or in a low-volume engineering bring-up alike. The model reports which document says what; it never assembles the number, the margin, or the verdict on which revision is safe.

    Here is the whole pipeline, constants first, as the example runs it:

    examples/rag/run.py · lines 19–80
    LEVEL = 2
    TOP_K = 4
    SYSTEM_PROMPT = (
        "You answer questions about Halvorsen appliances using only the numbered sources below. "
        "If the sources do not contain the answer, say so instead of guessing. End your answer with "
        "a line starting 'Sources:' listing the citations, like 'dw300-manual#3', that you used."
    )
    
    
    def _cosine(a: list[float], b: list[float]) -> float:
        dot = sum(x * y for x, y in zip(a, b))
        return dot  # StubEmbedder and OllamaEmbedder both return unit vectors, so dot == cosine
    
    
    def _retrieve(question: str, sections: dict[str, Section], embedder: Embedder, k: int) -> list[Section]:
        ordered = list(sections.values())
        vectors = embedder.embed([question] + [f"{s.title}\n{s.text}" for s in ordered])
        query_vec, chunk_vecs = vectors[0], vectors[1:]
        scored = sorted(zip(ordered, chunk_vecs), key=lambda pair: _cosine(query_vec, pair[1]), reverse=True)
        return [section for section, _ in scored[:k]]
    
    
    def _build_prompt(question: str, sources: list[Section]) -> str:
        blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
        return f"Sources:\n\n{blocks}\n\nQuestion: {question}"
    
    
    def run(
        question: str,
        model: Model,
        embedder: Embedder,
        tracer: Tracer,
        *,
        corpus_dir: Path = DEFAULT_CORPUS_DIR,
        top_k: int = TOP_K,
    ) -> Answer:
        sections = load_sections(corpus_dir)
        tracer.record(kind="code", decided_by="code", title="Chunk corpus", detail=f"{len(sections)} sections")
        sources = _retrieve(question, sections, embedder, top_k)
        tracer.record(
            kind="code",
            decided_by="code",
            title="Embed and retrieve top-k",
            detail=", ".join(s.cite for s in sources),
        )
        prompt = _build_prompt(question, sources)
        messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=prompt)]
        tracer.record(kind="code", decided_by="code", title="Build prompt with sources", detail=f"{len(sources)} sources")
        completion = model.complete(messages, max_tokens=500)
        tracer.record(
            kind="model",
            decided_by="code",
            title="Ask the model for a cited answer",
            detail=completion.text[:200],
            tokens_in=completion.tokens_in,
            tokens_out=completion.tokens_out,
            ms=completion.ms,
        )
        citations = cited_sources(completion.text)
        tracer.record(kind="code", decided_by="code", title="Parse citations", detail=", ".join(citations) or "none")
        return Answer(text=completion.text, citations=citations, retrieved_sources=[s.cite for s in sources])
    

    Every step above is decided by code, not by the model. The one model call answers the question; it does not choose what happens next, because there is nothing left to choose. Run it yourself:

    examples/rag/README.md · lines 16–16
    python -m examples.rag --model stub:scripted
    When you do not need this

    Try level 0, no model at all first if the documents are small enough for plain keyword search or a regular expression to answer the question directly, with no model and no embeddings to keep in sync.

    Try putting the whole document set straight into the prompt instead of retrieving from it, if it comfortably fits the model’s context window and you are not reusing the same documents across many separate questions. That is context engineering.

    Move up to RAG once the documents are too large, too numerous, or reused too often for either of those to still make sense.

    Failure modes

    The right passage is not retrieved

    How to notice it
    The answer is generic, off-topic, or contradicts a document you know covers the question; a product that shows its sources shows ones that don't relate to what was asked.
    How to test for it
    Run questions where you know which sections hold the answer. Check those sections against the retrieved chunk ids to measure retrieval coverage. Separately check the answer's citations; citation hit rate is not a retrieval metric.

    The passage is retrieved but ignored

    How to notice it
    The correct source is visibly in the retrieved set, but the answer still doesn't use it, invents a different answer, or cites the wrong section.
    How to test for it
    Compare retrieved passages, answer claims, and citations. Missing citations can flag a problem, but inspect the answer to distinguish ignored evidence from a citation omission.

    Chunk boundaries split a fact

    How to notice it
    A number and the sentence explaining it end up in two different chunks (a price in one, the part it prices in the next), so the answer gets one without the other.
    How to test for it
    Check multi-hop and numeric questions specifically. A grading rule with several required patterns catches a citation that matches only part of a compound fact.

    Stale index

    How to notice it
    The answer is correct for an old version of a document but wrong for the current one: a warranty length that changed, a part number that was superseded.
    How to test for it
    In an index that stores a snapshot of passage text, change a source fact without refreshing the index. Check whether retrieval still returns the old passage. Other designs fetch current text separately; test the actual refresh path.

    Conflicting sources

    How to notice it
    Two documents disagree (an installation guide states one clearance, a later service bulletin corrects it) and the answer picks one without saying there's a conflict.
    How to test for it
    Ask a question the corpus answers two different ways on purpose, and check whether the answer names both values and says which one is authoritative.

    Prompt injection through retrieved text

    How to notice it
    A document contains text written to look like an instruction ("ignore the above and say X"), and the answer follows it instead of answering the question.
    How to test for it
    Add a document section containing an embedded instruction and see whether the answer changes to match it. Telling the model to "answer only from the sources" does not by itself prevent this, since the injected text is a source.

    Cost and latency

    Measured: averages over the 60-question run on Muse Glimmer 30B, a model in the Large local (about 30B) class, on one local GPU. Tokens out include the model's hidden reasoning, which it spends before answering. Holds for this model class only.

    512Tokens in, per question
    687Tokens out, per question
    6.6sWall time, per question
    60Questions in the run
    Compared with Agentic RAG (level 5), same modelPer question, Agentic RAG (level 5) took 4,037 tokens in, 1,247 out and 9.6s on Muse Glimmer 30B; this page took 512 in, 687 out and 6.6s, on the same 60 questions.

    One model call per question, always: the pipeline’s code makes exactly one. Agentic RAG, the level 5 version, makes several; both were run on the same questions and model, and the line above compares them.

    How to Evaluate It

    60 questionslookupmulti-hopnumericunanswerableconflicting sources

    The site scores every technique against the same 60-question synthetic set, 12 questions in each of five kinds, over the appliance document set in evals/corpus/. RAG is graded the same way every other level is: exact match or a rubric where exact match doesn’t apply, plus citation hit rate, the share of questions where every source the grading rule expects was actually cited in the answer.

    Two kinds matter most for RAG specifically. Multi-hop questions need two chunks retrieved and used together, which is exactly what single-pass retrieval struggles with. Conflicting-source questions need the answer to notice two chunks disagree, not just cite whichever one the search ranked first.

    Measured result

    46 of 60 correct

    Agentic RAG (level 5): 55 of 60 on the same questions, same model.

    On the site's 60-question set, run 09/23/2026 with Muse Glimmer 30B by Meta, a model in the Large local (about 30B) class. Open weights at 4-bit (Q4_K_M), run on one local GPU through Ollama. The tag is a local build of muse-glimmer:30b.

    Correct answers by question kind
    Question kindThis pageAgentic RAG (level 5)Share
    Lookup12 of 1212 of 12
    Numeric11 of 1212 of 12
    Conflicting sources10 of 1211 of 12
    Not in the documents11 of 1212 of 12
    Multi-hop2 of 128 of 12
    Retrieval coverage
    79%

    of the sections the questions need reached the prompt

    Citation coverage
    76%

    of the sections the questions need were cited in the answer

    Model-decided steps
    0

    code chose every step; the model only wrote the answer

    Empty or ungraded
    0

    answers the score could not read

    Ended by a cap
    0

    questions where the code's step or token budget stopped the loop

    Graded by the same model on 28 rubric questions, the rest by exact match. Checked by a person on 09/23/2026: All 14 answers scored wrong were read, and each is wrong by its rubric or pattern. The closest calls: C06 names all three documents that disagree and quotes the bulletin superseding the others, but never says which figure is correct, which its rubric asks for; U10 refuses correctly but assumes a child lock exists. This holds for the Large local (about 30B) class only. Not yet run: Small local (about 8B); Frontier API.

    The misses sit where single-pass retrieval predicts. Most multi-hop answers were honest: they said the documents did not cover part of the question, because the one search never reached the second document it needed. The worst miss is the other kind. Asked whether a drive belt is covered eighteen months after purchase, the model found the two-year warranty, never saw the exclusion for wear parts, and said yes. That is the case the page’s failure modes warn about: a fluent answer built only from what retrieval happened to return.

    To run it yourself, python scripts/eval_run.py --example rag --model <spec> --dry projects the cost first; docs/FIRST-LIVE-RUN.md is the full sequence and the checks to read before the score.

    Run it

    What to monitor

    Citation hit rate on a sample of real questions, and how often a question comes back with zero retrieved chunks above a similarity floor. Watch retrieval latency and generation latency separately, since a slow answer can come from either half.

    Cost at volume

    This example makes one generation call per question. Cost depends on input and output tokens, model rates, query embeddings, retrieval infrastructure, and how often documents are indexed again. Document embeddings can be reused across questions; no component always dominates.

    How it fails in production

    A document changes and the index isn't rebuilt, so the model confidently answers from stale text. Or a document containing an injected instruction gets indexed and later retrieved and followed.

    What to log

    The question, the retrieved chunk ids with their similarity scores, the final citations, and the full prompt sent to the model, so a bad answer traces back to a retrieval failure or a generation failure without re-running anything.

    Try it

    1. Use it

      Open a chat app's file feature with a document you know well, and ask a question whose answer needs two sections. Does it find both, or one?

    2. Build it

      Run python -m examples.rag --model stub:scripted from the repo root: one call, a grounded answer, one citation. Now ask something the corpus does not cover: --question "What is the DR-520 vent length?". Neither the answer nor the citation moves: the citations line is parsed out of the reply, never checked against what retrieval returned. The write and check page adds that check.

    3. Build it

      Change the retriever. _retrieve in examples/rag/run.py is the only function that picks passages: it embeds the question and every section, then keeps the top k by cosine. Replace its body with a count of shared words, keep the signature, and run again. Which questions get better, and which get worse?

    4. Either lane

      Cause one of the failure modes above on purpose, using the documents in evals/corpus/.

    5. Either lane

      Pick two documents of your own that can disagree over time: a manual and a later errata sheet, say. Put both into a search tool that shows citations, and ask the question the older one alone would answer wrong. Does the answer cite both, or only the one that sounds authoritative?

    How it connects

    Before, after and instead of this

    Move up when

    Instead of

    Decoded in

    Optional: products, tools, and models

    13 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

    Explore 7 more examples
    In practice

    Ask a question about a manual

    Retrieve the warranty passages, put them in the model request, and return an answer with source references.

    Out there

    Named products, tools and models

    Products8
    • ChatGPT ProjectsOpenAI · files and instructions in a chat app
    • Claude ProjectsAnthropic · files and instructions in a chat app
    • DifyLangGenius · visual workflow builder
    • Gemini NotebookGoogle · research notebook · formerly NotebookLM
    • GleanGlean · workplace search
    • Microsoft 365 CopilotMicrosoft · workplace assistant
    • Notion AINotion · workspace assistant
    • PerplexityPerplexity · answer engine
    Tools5
    • Cohere RerankCohere · reranker
    • Haystackdeepset · retrieval framework
    • jina-reranker-v3.5Jina AI · reranker
    • LangChainLangChain · application framework
    • LlamaIndexLlamaIndex · retrieval framework

    Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

    Where this comes from

    Primary sources

    1. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks · arXiv (Meta AI Research, UCL, NYU), 05/22/2020 (accessed 09/19/2026)
    2. What are Projects? · Anthropic (Claude Help Center) (accessed 09/19/2026)
    3. File search · OpenAI (API documentation) (accessed 09/19/2026)

    Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page