The example below indexes the same synthetic corpus RAG uses, then answers one query two ways
instead of one: by embedding every chunk and ranking them by cosine similarity to the query, and
separately by BM25 keyword score over the same chunks. It reports where the two result sets agree
and where they diverge, instead of answering the question: this level searches, it does not
answer.
What the example embeds with is not an embedding model, and nothing it returns is evidence about
one. StubEmbedder (examples/common/model.py) hashes words into 64 buckets and counts them: a
bag of words with no notion that “quiet” and “dBA” are related unless the words themselves
overlap. Run it on a query that shares real words with the corpus (“DW-300 Normal cycle water
use”) and its top hit matches keyword search’s top hit exactly, which tells you the pipeline
works and nothing about semantics. Run it on “Which dishwasher is quieter, the DW-300 or the
DW-480?” and neither method finds specs-comparison#2, the section that actually gives both
decibel ratings, because the word “quieter” never appears in the corpus at all. A trained model is
what would close that gap, and where one would go is OllamaEmbedder, behind the same Embedder
interface. Read every score below as the shape of the mechanism, not as a result.
examples/embeddings_search/run.py · lines 61–89
def run(
query: str,
model: Model | None,
embedder: Embedder,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
k: int = TOP_K,
) -> Answer:
del model # this level searches; it does not answer
sections = load_sections(corpus_dir)
tracer.record(kind="code", decided_by="code", title="Chunk the corpus", detail=f"{len(sections)} sections")
semantic = _semantic_search(query, sections, embedder, k)
tracer.record(
kind="code",
decided_by="code",
title="Embed the index and rank it by similarity",
detail=", ".join(f"{s.cite}={score:.2f}" for s, score in semantic),
)
keyword = bm25_search(sections, query, k=k)
tracer.record(
kind="code",
decided_by="code",
title="Rank the same query by keyword (BM25)",
detail=", ".join(f"{s.cite}={score:.2f}" for s, score in keyword),
)
summary = _compare(semantic, keyword)
tracer.record(kind="code", decided_by="code", title="Compare the two result sets", detail=summary)
return Answer(text=summary, citations=[s.cite for s, _ in semantic])
The similarity function _semantic_search calls, just above it in the same file, divides by both
vectors’ lengths instead of taking a bare dot product, so the ranking is a true cosine whichever
Embedder is plugged in and not only for one that happens to return unit vectors. Every step is
decided_by: "code": what gets embedded, how many results come back, and how the two result sets
get compared are fixed before anything runs. The example implements neither of the two optional
parts above: no hybrid fusion of the two lists it prints, and no reranking.
Run it yourself:
examples/embeddings_search/README.md · lines 17–17
python -m examples.embeddings_search --model stub --question "DW-300 Normal cycle water use"