The minimal version of RAG is four fixed steps: chunk the documents, embed the question and
every chunk, keep the top few chunks by similarity, and ask the model once with those chunks as
its only sources.
Chunking splits documents into pieces small enough to embed and retrieve individually. The
example below chunks by section, since the synthetic document set already has numbered
sections; a real document set usually needs its own splitter, tuned so a chunk holds one
complete idea rather than cutting a sentence or a table row in half.
Embedding turns text into a vector, a fixed-length list of numbers, using a model trained so
that texts with similar meaning get vectors that point in similar directions; the mechanism, and
the search built on it, is embeddings and
search. The code below is
written against an Embedder interface with two implementations: a deterministic stub for
tests, and a real embedding model behind the same interface, so the retrieval logic never has to
know which one is running.
Retrieval scores every chunk’s embedding against the question’s embedding by cosine similarity
(how closely the two vectors point in the same direction) and keeps the top k, four by
default. _cosine below returns a plain dot product rather than a full cosine, because both
embedders return unit vectors, for which the two are the same number. This is the one place a
real system usually adds more: a second, more expensive reranking pass over a larger first cut
of candidates, scored by a model trained for that job. Cohere and Jina AI both sell one. The
example skips reranking to keep the pipeline to four fixed steps.
Prompt assembly numbers every retrieved chunk, includes its citation (file#section), and the
system prompt instructs the model to answer using only those sources and to name which ones it
used. Citations are then parsed back out of the model’s answer with a regular expression, so
the calling code always knows, in a form it can check automatically, which sources actually
contributed to the answer.
The same four steps work over an engineer’s own documents, not just reference text. Orbeck Power
Systems’ SRB-5030 datasheet states one maximum input voltage; a later engineering change notice
supersedes it for two of the board’s three revisions, over a capacitor derating rule, and the
datasheet is never reissued to say so. One search that retrieves both documents returns an answer
with citations a reader can check by hand, in a production test or in a low-volume engineering
bring-up alike. The model reports which document says what; it never assembles the number, the
margin, or the verdict on which revision is safe.
Here is the whole pipeline, constants first, as the example runs it:
examples/rag/run.py · lines 19–80
LEVEL = 2
TOP_K = 4
SYSTEM_PROMPT = (
"You answer questions about Halvorsen appliances using only the numbered sources below. "
"If the sources do not contain the answer, say so instead of guessing. End your answer with "
"a line starting 'Sources:' listing the citations, like 'dw300-manual#3', that you used."
)
def _cosine(a: list[float], b: list[float]) -> float:
dot = sum(x * y for x, y in zip(a, b))
return dot # StubEmbedder and OllamaEmbedder both return unit vectors, so dot == cosine
def _retrieve(question: str, sections: dict[str, Section], embedder: Embedder, k: int) -> list[Section]:
ordered = list(sections.values())
vectors = embedder.embed([question] + [f"{s.title}\n{s.text}" for s in ordered])
query_vec, chunk_vecs = vectors[0], vectors[1:]
scored = sorted(zip(ordered, chunk_vecs), key=lambda pair: _cosine(query_vec, pair[1]), reverse=True)
return [section for section, _ in scored[:k]]
def _build_prompt(question: str, sources: list[Section]) -> str:
blocks = "\n\n".join(f"[{s.cite}] {s.title}\n{s.text}" for s in sources)
return f"Sources:\n\n{blocks}\n\nQuestion: {question}"
def run(
question: str,
model: Model,
embedder: Embedder,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
top_k: int = TOP_K,
) -> Answer:
sections = load_sections(corpus_dir)
tracer.record(kind="code", decided_by="code", title="Chunk corpus", detail=f"{len(sections)} sections")
sources = _retrieve(question, sections, embedder, top_k)
tracer.record(
kind="code",
decided_by="code",
title="Embed and retrieve top-k",
detail=", ".join(s.cite for s in sources),
)
prompt = _build_prompt(question, sources)
messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=prompt)]
tracer.record(kind="code", decided_by="code", title="Build prompt with sources", detail=f"{len(sources)} sources")
completion = model.complete(messages, max_tokens=500)
tracer.record(
kind="model",
decided_by="code",
title="Ask the model for a cited answer",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
citations = cited_sources(completion.text)
tracer.record(kind="code", decided_by="code", title="Parse citations", detail=", ".join(citations) or "none")
return Answer(text=completion.text, citations=citations, retrieved_sources=[s.cite for s in sources])
Every step above is decided by code, not by the model. The one model call answers the question;
it does not choose what happens next, because there is nothing left to choose. Run it yourself:
examples/rag/README.md · lines 16–16
python -m examples.rag --model stub:scripted