# Retrieval-augmented generation (RAG)

_Level 02 · Added context · measured_

Searching your documents and giving the results to the model.

## Conceptual architecture: Two paths meet at retrieval.

Indexing prepares the sources. A query selects evidence for this answer.

- **Source documents:** Versioned text with access rules
- **Prepare the index:** Split, preserve IDs, index the content
- **Question:** What the user needs to know
- **Retrieve + select:** Lexical, vector, hybrid; optional rerank
- **Generate an answer:** Question + selected evidence
- **Check or abstain:** Supported claims and valid citations

Connections:
- Source documents → indexing → Prepare the index
- Prepare the index → searchable sources → Retrieve + select
- Question → query → Retrieve + select
- Retrieve + select → evidence packet → Generate an answer
- Generate an answer → draft → Check or abstain

RAG is retrieval-augmented generation, not a guarantee of truth. A missing answer may be a retrieval failure, a source gap, or a generation error; those need different fixes.
- **Before retrieval:** Apply document permissions. Keep source IDs and revision metadata.
- **Before answering:** Check whether the selected evidence covers every part of the question.
- **After answering:** A real citation ID can still support the wrong claim. Check entailment as well as existence.

## Try this in a recipe
- [Answer a warranty question with evidence](/gradient_ascent/recipes/document-qa.md): Retrieve the relevant policy, answer each part of the question, and distinguish an unknown fact from a retrieval miss.

## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a question through source retrieval into a grounded answer. Inspect whether the retrieved passages actually support the response, including exceptions and unanswered parts.

**Assumptions:** The collection may be incomplete or outdated. A citation is useful only when its passage supports the associated claim.

**Design choices:** Separate retrieval quality from answer quality. Choose whether the evidence supports a direct answer, a qualified answer, or a request for more information.

**Request:** Is water damage covered by the DW-480 warranty?

**Starting evidence:** Manual v3 §2: two-year coverage. Manual v3 §4: water damage excluded.

**Action and control:** Retrieve general coverage and the relevant exclusion before drafting.

**Stage records (authored, not executed):**

### Source packet

Manual v3 §2: Coverage lasts two years.
Manual v3 §4: Water damage is excluded.
Question: Is water damage covered for the DW-480?
Provenance: fictional manual passages supplied for this example.

What changed: Duration and exclusions are distinct pieces of evidence.

### Retrieval plan

Query: DW-480 warranty water damage exclusions.
Required evidence: coverage terms and relevant exclusions.
Version constraint: use the same applicable revision.
Do not infer coverage from duration alone.

What changed: The query is designed around the claim the answer must support.

### Selected passages

Selected: v3 §2 and v3 §4.
Rejected shortcut: §2 by itself.
Claim to support: water-damage coverage.
Supporting passage: §4, not §2.

What changed: Selecting a passage is not enough; its content must bear on the question.

### Answer with support

Water damage is excluded [v3 §4], even during the two-year period [v3 §2].
Claim 1 → exclusion clause.
Claim 2 → duration clause.

What changed: Each claim is paired with the passage that supports it.

### Missing-evidence review

Remove §4 from the packet.
Still established: duration is two years.
No longer established: whether water damage is covered.
Next action: retrieve applicable exclusions or leave coverage unresolved.

What changed: The changed result is caused by an evidence gap, not a different warranty fact.

### Your source contract

Replace: manual with your policies, notes, or records.
Specify: applicable version and freshness.
Check: each consequential claim against its supporting passage.
Escalate: missing or conflicting terms that affect the answer.

What changed: Your documents change; the claim-to-evidence relationship remains.

**Sample result:** Water damage is excluded [v3 §4], even during the two-year period [v3 §2].

**Change something — Retrieve only the general clause:** §2 establishes duration, not water-damage coverage. Request exclusion terms or abstain on coverage.

**Decision:** Can a citation to duration support a water-damage claim?

**Answer:** No; the citation must support the claim.

**Why:** Include a missing exclusion clause and conflicting revisions; retrieval and generation can fail separately.

**Review criteria:** Visible query, retrieved passages, grounded answer, source links, and an abstention when evidence is insufficient.

**Recovery:** When evidence conflicts or does not cover the question, show the gap and search or escalate appropriately. Rewording a confident answer does not repair missing support.

**Adapt it:** Replace the source collection with your manuals, notes, policies, or project records. Set freshness and citation expectations appropriate to the people relying on the answer.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a question through source retrieval into a grounded answer. Inspect whether the retrieved passages actually support the response, including exceptions and unanswered parts.

**Assumptions:** The collection may be incomplete or outdated. A citation is useful only when its passage supports the associated claim.

**Design choices:** Separate retrieval quality from answer quality. Choose whether the evidence supports a direct answer, a qualified answer, or a request for more information.

**Request:** Which settling time applies before measuring this board revision?

**Starting evidence:** Spec rev C: wait 20 ms. Lab note for rev B: wait 5 ms. DUT is rev C.

**Action and control:** Retrieve revision-specific requirements and cite the applicable clause.

**Stage records (authored, not executed):**

### Input record

Spec rev C: wait 20 ms. Lab note for rev B: wait 5 ms. DUT is rev C.

What changed: Establish the facts supplied for this version of the task.

### Design note

Separate retrieval quality from answer quality. Choose whether the evidence supports a direct answer, a qualified answer, or a request for more information.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Retrieve revision-specific requirements and cite the applicable clause.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Use 20 ms for rev C, citing spec C. The older lab note is not the governing requirement.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Check revision, requirement source, units, and quoted support.

If the result falls short:
When evidence conflicts or does not cover the question, show the gap and search or escalate appropriately. Rewording a confident answer does not repair missing support.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Replace the source collection with your manuals, notes, policies, or project records. Set freshness and citation expectations appropriate to the people relying on the answer.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Use 20 ms for rev C, citing spec C. The older lab note is not the governing requirement.

**Change something — Retrieve only the rev B lab note:** The snippet is relevant to settling time but not the target revision. Report missing applicable evidence.

**Decision:** Can a related older note establish the current requirement?

**Answer:** No; retrieve the correct revision.

**Why:** Retrieval relevance does not imply applicability or authority.

**Review criteria:** Check revision, requirement source, units, and quoted support.

**Recovery:** When evidence conflicts or does not cover the question, show the gap and search or escalate appropriately. Rewording a confident answer does not repair missing support.

**Adapt it:** Replace the source collection with your manuals, notes, policies, or project records. Set freshness and citation expectations appropriate to the people relying on the answer.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a question through source retrieval into a grounded answer. Inspect whether the retrieved passages actually support the response, including exceptions and unanswered parts.

**Assumptions:** The collection may be incomplete or outdated. A citation is useful only when its passage supports the associated claim.

**Design choices:** Separate retrieval quality from answer quality. Choose whether the evidence supports a direct answer, a qualified answer, or a request for more information.

**Request:** Summarize current risks for Atlas in this week's status report.

**Starting evidence:** Last week: on track. Current tracker: supplier delivery late. Meeting note: alternative supplier under consideration.

**Action and control:** Retrieve current dated evidence and distinguish a possible mitigation from an approved change.

**Stage records (authored, not executed):**

### Input record

Last week: on track. Current tracker: supplier delivery late. Meeting note: alternative supplier under consideration.

What changed: Establish the facts supplied for this version of the task.

### Design note

Separate retrieval quality from answer quality. Choose whether the evidence supports a direct answer, a qualified answer, or a request for more information.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Retrieve current dated evidence and distinguish a possible mitigation from an approved change.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Atlas has a delivery risk. Alternative sourcing is being considered, not committed. Link both sources.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Check source dates, risk claims, and whether mitigations are approved or merely discussed.

If the result falls short:
When evidence conflicts or does not cover the question, show the gap and search or escalate appropriately. Rewording a confident answer does not repair missing support.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Replace the source collection with your manuals, notes, policies, or project records. Set freshness and citation expectations appropriate to the people relying on the answer.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Atlas has a delivery risk. Alternative sourcing is being considered, not committed. Link both sources.

**Change something — Only last week's report is available:** No fresh evidence found. Current risk status is unknown; do not automatically repeat green.

**Decision:** Does an old green report prove current green status?

**Answer:** No; request current evidence.

**Why:** Prior reports supply continuity, not proof of present conditions.

**Review criteria:** Check source dates, risk claims, and whether mitigations are approved or merely discussed.

**Recovery:** When evidence conflicts or does not cover the question, show the gap and search or escalate appropriately. Rewording a confident answer does not repair missing support.

**Adapt it:** Replace the source collection with your manuals, notes, policies, or project records. Set freshness and citation expectations appropriate to the people relying on the answer.

Retrieval-augmented generation, or RAG, supplies retrieved information to a model when it
generates an answer. Retrieval can use keywords, embeddings, or both; the sources may be your
documents or another searchable collection. The 2020 RAG paper describes combining retrieval
with generation so the model can draw on external information[1].

This page teaches a simple, fixed retrieval pipeline: split documents into passages, embed them,
retrieve a few relevant passages for the question, and send those passages to one model call.
The instruction is to answer from the evidence, but the model can still make unsupported claims.
The fixed number of passages and single generation call are choices in this example, not rules
that define all RAG systems. Other implementations rerank, rewrite queries, retrieve repeatedly,
or combine evidence across documents.

The simple pipeline sits at level 2 because code determines how context is selected. If a model
chooses successive searches, this site calls that [agentic
RAG](/gradient_ascent/techniques/agentic-rag/). Both approaches augment generation with retrieved information.

This page is measured: the cost and the score under How to Evaluate It come from a recorded run of this example on a real model, and hold for that model's class. The step-through just below is still a scripted illustration, and source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 2 · RAG. The same steps are described in the sections below._

## Practical guidance

Upload files to a chat app's project or file feature and ask about them. The product may put the
files directly into context, retrieve selected passages, or combine both approaches. Uploading
a file alone does not tell you which method it uses; check the product documentation.

Ask questions a handful of passages can answer on their own. "What does the warranty cover"
is a focused lookup. "Summarize every change across all our contracts this year" requires much
broader coverage: a few highly ranked passages may omit important changes. For that task, check
whether the product can systematically cover the collection. Narrower questions and an explicit
document checklist make omissions easier to notice.

Anthropic's own documentation describes this directly: once a project's uploaded files approach
what the context window can hold, Claude switches into what Anthropic calls RAG mode. Anthropic
says that expands how much a project can hold by up to ten times while maintaining response
quality[2]; that is the maker's claim, not a number this site has measured. OpenAI
documents the same pattern for its own file search feature: it retrieves passages from uploaded
files by keyword and meaning together, and returns an answer with citations to the files it
used[3].

Check the citations every time the product shows them. Open the source it names and confirm the
sentence it cites is actually there. A citation that does not obviously support the sentence
beside it, or an answer with none at all, is unconfirmed. It does not by itself tell you whether
retrieval failed, the model ignored evidence, or the interface omitted the citation.

RAG can combine facts from multiple documents when the necessary evidence is retrieved and
used correctly. A single search can return passages from several files, but questions whose
second lookup depends on the first answer may need query decomposition or repeated retrieval.
Inspect all required sources and compare a fixed pipeline with an agentic one on the same task.

## Implementation details

The minimal version of RAG is four fixed steps: chunk the documents, embed the question and
every chunk, keep the top few chunks by similarity, and ask the model once with those chunks as
its only sources.

Chunking splits documents into pieces small enough to embed and retrieve individually. The
example below chunks by section, since the synthetic document set already has numbered
sections; a real document set usually needs its own splitter, tuned so a chunk holds one
complete idea rather than cutting a sentence or a table row in half.

Embedding turns text into a vector, a fixed-length list of numbers, using a model trained so
that texts with similar meaning get vectors that point in similar directions; the mechanism, and
the search built on it, is [embeddings and
search](/gradient_ascent/techniques/embeddings-search/). The code below is
written against an `Embedder` interface with two implementations: a deterministic stub for
tests, and a real embedding model behind the same interface, so the retrieval logic never has to
know which one is running.

Retrieval scores every chunk's embedding against the question's embedding by cosine similarity
(how closely the two vectors point in the same direction) and keeps the top `k`, four by
default. `_cosine` below returns a plain dot product rather than a full cosine, because both
embedders return unit vectors, for which the two are the same number. This is the one place a
real system usually adds more: a second, more expensive reranking pass over a larger first cut
of candidates, scored by a model trained for that job. Cohere and Jina AI both sell one. The
example skips reranking to keep the pipeline to four fixed steps.

Prompt assembly numbers every retrieved chunk, includes its citation (`file#section`), and the
system prompt instructs the model to answer using only those sources and to name which ones it
used. Citations are then parsed back out of the model's answer with a regular expression, so
the calling code always knows, in a form it can check automatically, which sources actually
contributed to the answer.

The same four steps work over an engineer's own documents, not just reference text. Orbeck Power
Systems' SRB-5030 datasheet states one maximum input voltage; a later engineering change notice
supersedes it for two of the board's three revisions, over a capacitor derating rule, and the
datasheet is never reissued to say so. One search that retrieves both documents returns an answer
with citations a reader can check by hand, in a production test or in a low-volume engineering
bring-up alike. The model reports which document says what; it never assembles the number, the
margin, or the verdict on which revision is safe.

Here is the whole pipeline, constants first, as the example runs it:

`examples/rag/run.py` (lines 19-80)

```python
LEVEL = 2
TOP_K = 4
SYSTEM_PROMPT = (
    "You answer questions about Halvorsen appliances using only the numbered sources below. "
    "If the sources do not contain the answer, say so instead of guessing. End your answer with "
    "a line starting 'Sources:' listing the citations, like 'dw300-manual#3', that you used."
)

def _cosine(a: list[float], b: list[float]) -> float:
    dot = sum(x * y for x, y in zip(a, b))
    return dot  # StubEmbedder and OllamaEmbedder both return unit vectors, so dot == cosine

def _retrieve(question: str, sections: dict[str, Section], embedder: Embedder, k: int) -> list[Section]:
    ordered = list(sections.values())
    vectors = embedder.embed([question] + [f"{s.title}\n{s.text}" for s in ordered])
    query_vec, chunk_vecs = vectors[0], vectors[1:]
    scored = sorted(zip(ordered, chunk_vecs), key=lambda pair: _cosine(query_vec, pair[1]), reverse=True)
    return [section for section, _ in scored[:k]]

def _build_prompt(question: str, sources: list[Section]) -> str:
    blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
    return f"Sources:\n\n{blocks}\n\nQuestion: {question}"

def run(
    question: str,
    model: Model,
    embedder: Embedder,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    top_k: int = TOP_K,
) -> Answer:
    sections = load_sections(corpus_dir)
    tracer.record(kind="code", decided_by="code", title="Chunk corpus", detail=f"{len(sections)} sections")
    sources = _retrieve(question, sections, embedder, top_k)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Embed and retrieve top-k",
        detail=", ".join(s.cite for s in sources),
    )
    prompt = _build_prompt(question, sources)
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=prompt)]
    tracer.record(kind="code", decided_by="code", title="Build prompt with sources", detail=f"{len(sources)} sources")
    completion = model.complete(messages, max_tokens=500)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model for a cited answer",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    citations = cited_sources(completion.text)
    tracer.record(kind="code", decided_by="code", title="Parse citations", detail=", ".join(citations) or "none")
    return Answer(text=completion.text, citations=citations, retrieved_sources=[s.cite for s in sources])

```

Every step above is decided by code, not by the model. The one model call answers the question;
it does not choose what happens next, because there is nothing left to choose. Run it yourself:

`examples/rag/README.md` (lines 16-16)

```text
python -m examples.rag --model stub:scripted
```

## When you do not need this

Try [level 0, no model at all](/gradient_ascent/techniques/order-zero/) first if the documents are small
enough for plain keyword search or a regular expression to answer the question directly, with no
model and no embeddings to keep in sync.

Try putting the whole document set straight into the prompt instead of retrieving from it, if it
comfortably fits the model's context window and you are not reusing the same documents across
many separate questions. That is [context
engineering](/gradient_ascent/techniques/context-engineering/).

Move up to RAG once the documents are too large, too numerous, or reused too often for either of
those to still make sense.

## Failure modes

### The right passage is not retrieved

- **How to notice it:** The answer is generic, off-topic, or contradicts a document you know covers the question; a product that shows its sources shows ones that don't relate to what was asked.
- **How to test for it:** Run questions where you know which sections hold the answer. Check those sections against the retrieved chunk ids to measure retrieval coverage. Separately check the answer's citations; citation hit rate is not a retrieval metric.

### The passage is retrieved but ignored

- **How to notice it:** The correct source is visibly in the retrieved set, but the answer still doesn't use it, invents a different answer, or cites the wrong section.
- **How to test for it:** Compare retrieved passages, answer claims, and citations. Missing citations can flag a problem, but inspect the answer to distinguish ignored evidence from a citation omission.

### Chunk boundaries split a fact

- **How to notice it:** A number and the sentence explaining it end up in two different chunks (a price in one, the part it prices in the next), so the answer gets one without the other.
- **How to test for it:** Check multi-hop and numeric questions specifically. A grading rule with several required patterns catches a citation that matches only part of a compound fact.

### Stale index

- **How to notice it:** The answer is correct for an old version of a document but wrong for the current one: a warranty length that changed, a part number that was superseded.
- **How to test for it:** In an index that stores a snapshot of passage text, change a source fact without refreshing the index. Check whether retrieval still returns the old passage. Other designs fetch current text separately; test the actual refresh path.

### Conflicting sources

- **How to notice it:** Two documents disagree (an installation guide states one clearance, a later service bulletin corrects it) and the answer picks one without saying there's a conflict.
- **How to test for it:** Ask a question the corpus answers two different ways on purpose, and check whether the answer names both values and says which one is authoritative.

### Prompt injection through retrieved text

- **How to notice it:** A document contains text written to look like an instruction ("ignore the above and say X"), and the answer follows it instead of answering the question.
- **How to test for it:** Add a document section containing an embedded instruction and see whether the answer changes to match it. Telling the model to "answer only from the sources" does not by itself prevent this, since the injected text is a source.

## Cost and latency

_Measured: averages over the 60-question run on Muse Glimmer 30B, a model in the Large local (about 30B) class, on one local GPU. Tokens out include the model's hidden reasoning, which it spends before answering. Holds for this model class only._

- **Tokens in, per question:** 512
- **Tokens out, per question:** 687
- **Wall time, per question:** 6.6s
- **Questions in the run:** 60

**Compared with Agentic RAG (level 5), same model.** Per question, Agentic RAG (level 5) took 4,037 tokens in, 1,247 out and 9.6s on Muse Glimmer 30B; this page took 512 in, 687 out and 6.6s, on the same 60 questions.

One model call per question, always: the pipeline's code makes exactly one. Agentic RAG, the
level 5 version, makes several; both were run on the same questions and model, and the line above
compares them.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

The site scores every technique against the same 60-question synthetic set, 12 questions in
each of five kinds, over the appliance document set in `evals/corpus/`. RAG is graded the same
way every other level is: exact match or a rubric where exact match doesn't apply, plus
**citation hit rate**, the share of questions where every source the grading rule expects was
actually cited in the answer.

Two kinds matter most for RAG specifically. Multi-hop questions need two chunks retrieved and
used together, which is exactly what single-pass retrieval struggles with. Conflicting-source
questions need the answer to notice two chunks disagree, not just cite whichever one the search
ranked first.

### Measured result: Muse Glimmer 30B

**46 of 60 correct** on the site's 60-question set, run 09/23/2026 with Muse Glimmer 30B by Meta, a model in the Large local (about 30B) class. Open weights at 4-bit (Q4_K_M), run on one local GPU through Ollama. The tag is a local build of muse-glimmer:30b.

Agentic RAG (level 5) scored 55 of 60 on the same questions with the same model.

| Question kind | This page | Agentic RAG (level 5), same model |
| --- | --- | --- |
| Lookup | 12 of 12 | 12 of 12 |
| Numeric | 11 of 12 | 12 of 12 |
| Conflicting sources | 10 of 12 | 11 of 12 |
| Not in the documents | 11 of 12 | 12 of 12 |
| Multi-hop | 2 of 12 | 8 of 12 |

- **Retrieval coverage:** 79% of the sections the questions need reached the prompt.
- **Citation coverage:** 76% of the sections the questions need were cited in the answer.
- **Model-decided steps:** 0. Code chose every step; the model only wrote the answer.
- **Empty replies:** 0. **Ungraded answers:** 0.
- **Ended by a cap:** 0 of 60 questions, where the step or token budget in the code stopped the loop and forced an answer.
- **Grader:** the same model, on 28 rubric questions, the rest by exact match. Checked by a person on 09/23/2026: All 14 answers scored wrong were read, and each is wrong by its rubric or pattern. The closest calls: C06 names all three documents that disagree and quotes the bulletin superseding the others, but never says which figure is correct, which its rubric asks for; U10 refuses correctly but assumes a child lock exists.

This holds for the Large local (about 30B) class only. Not yet run: Small local (about 8B); Frontier API.

Result file: https://github.com/reedos/gradient_ascent/blob/main/evals/results/rag/ollama_muse-glimmer_30b-q4_K_M-dflash.json · recorded trace: https://github.com/reedos/gradient_ascent/blob/main/examples/rag/trace.json

The misses sit where single-pass retrieval predicts. Most multi-hop answers were honest: they
said the documents did not cover part of the question, because the one search never reached the
second document it needed. The worst miss is the other kind. Asked whether a drive belt is
covered eighteen months after purchase, the model found the two-year warranty, never saw the
exclusion for wear parts, and said yes. That is the case the page's failure modes warn about: a
fluent answer built only from what retrieval happened to return.

To run it yourself, `python scripts/eval_run.py --example rag --model <spec> --dry` projects the
cost first; `docs/FIRST-LIVE-RUN.md` is the full sequence and the checks to read before the score.

## Run it

**What to monitor.** Citation hit rate on a sample of real questions, and how often a question comes back with zero retrieved chunks above a similarity floor. Watch retrieval latency and generation latency separately, since a slow answer can come from either half.

**Cost at volume.** This example makes one generation call per question. Cost depends on input and output tokens, model rates, query embeddings, retrieval infrastructure, and how often documents are indexed again. Document embeddings can be reused across questions; no component always dominates.

**How it fails in production.** A document changes and the index isn't rebuilt, so the model confidently answers from stale text. Or a document containing an injected instruction gets indexed and later retrieved and followed.

**What to log.** The question, the retrieved chunk ids with their similarity scores, the final citations, and the full prompt sent to the model, so a bad answer traces back to a retrieval failure or a generation failure without re-running anything.

## Try it

1. **Use it.** Open a chat app's file feature with a document you know well, and ask a question whose answer needs two sections. Does it find both, or one?
2. **Build it.** Run python -m examples.rag --model stub:scripted from the repo root: one call, a grounded answer, one citation. Now ask something the corpus does not cover: --question "What is the DR-520 vent length?". Neither the answer nor the citation moves: the citations line is parsed out of the reply, never checked against what retrieval returned. The write and check page adds that check.
3. **Build it.** Change the retriever. _retrieve in examples/rag/run.py is the only function that picks passages: it embeds the question and every section, then keeps the top k by cosine. Replace its body with a count of shared words, keep the signature, and run again. Which questions get better, and which get worse?
4. **Either lane.** Cause one of the failure modes above on purpose, using the documents in evals/corpus/.
5. **Either lane.** Pick two documents of your own that can disagree over time: a manual and a later errata sheet, say. Put both into a search tool that shows citations, and ask the question the older one alone would answer wrong. Does the answer cite both, or only the one that sounds authoritative?


## Sources

1. [Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks](https://arxiv.org/abs/2005.11401) — arXiv (Meta AI Research, UCL, NYU), 2020-05-22 (accessed 2026-09-19)
2. [What are Projects?](https://support.claude.com/en/articles/9517075-what-are-projects) — Anthropic (Claude Help Center) (accessed 2026-09-19)
3. [File search](https://developers.openai.com/api/docs/guides/tools-file-search) — OpenAI (API documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
