# Embeddings and search

_Level 02 · Added context · sourced_

Finding text by meaning instead of by keyword.


## Try this in a recipe
- [Answer a warranty question with evidence](/gradient_ascent/recipes/document-qa.md): Retrieve the relevant policy, answer each part of the question, and distinguish an unknown fact from a retrieval miss.

## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a meaning-based query into candidate results. Examine why a similar passage can be useful for discovery while still being the wrong item, revision, or answer.

**Assumptions:** Similarity scores rank candidates; they do not certify correctness. The collection and its metadata determine what can be found.

**Design choices:** Combine semantic matching with exact identifiers and filters where appropriate. Tune the number of candidates against noise and the cost of missing a useful result.

**Request:** Find troubleshooting guidance for a knocking sound in pump AX-20.

**Starting evidence:** Documents: A, AX-20 knocking; B, AX-30 vibration; C, AX-20 electrical error.

**Action and control:** Use meaning-based matching, then filter for the exact model. Similarity is not compatibility.

**Stage records (authored, not executed):**

### Input record

Documents: A, AX-20 knocking; B, AX-30 vibration; C, AX-20 electrical error.

What changed: Establish the facts supplied for this version of the task.

### Design note

Combine semantic matching with exact identifiers and filters where appropriate. Tune the number of candidates against noise and the cost of missing a useful result.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Use meaning-based matching, then filter for the exact model. Similarity is not compatibility.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Selected: A. B is related but belongs to AX-30; C has the right model but wrong symptom.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Ranked snippets, model filters, relevance judgments, and a case where hybrid search is preferable.

If the result falls short:
If retrieval fails, try alternate wording or an exact lookup and inspect collection coverage. A missing result does not show that the underlying fact is false.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use the pattern for documents, parts, notes, or support records. Decide which identifiers must match exactly and which wording differences should be tolerated.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Selected: A. B is related but belongs to AX-30; C has the right model but wrong symptom.

**Change something — Remove the AX-20 knocking document:** B remains similar but does not establish AX-20 guidance. Report the coverage gap.

**Decision:** Is the nearest semantic match necessarily applicable?

**Answer:** No; inspect identifiers and scope.

**Why:** Contrast semantic similarity with an exact model-number lookup; a similar passage may refer to the wrong product.

**Review criteria:** Ranked snippets, model filters, relevance judgments, and a case where hybrid search is preferable.

**Recovery:** If retrieval fails, try alternate wording or an exact lookup and inspect collection coverage. A missing result does not show that the underlying fact is false.

**Adapt it:** Use the pattern for documents, parts, notes, or support records. Decide which identifiers must match exactly and which wording differences should be tolerated.

An embedding is a list of floating-point numbers standing in for a piece of text. OpenAI's
documentation states what makes that useful: the distance between two embeddings measures how
related the texts are, small distances meaning high relatedness, and it recommends cosine
similarity (how closely two vectors point the same way) for the comparison[1].

Search built on this has three fixed parts: documents are cut into chunks small enough to
retrieve on their own; every chunk is embedded once into a vector index, a database such as
pgvector, Pinecone, Weaviate or Qdrant; and a question is embedded the same way and ranked
against that index. Two more are optional. Hybrid search runs a keyword index alongside the
vectors and merges the two lists, which is what pgvector documents doing with Postgres full-text
search[3]. Reranking re-scores a larger first cut of candidates with a second model:
Cohere describes its rerank models as sorting text by relevance to a query, over results an
existing search already returned[2].

This is the retrieval half of [RAG](/gradient_ascent/techniques/rag/) without the answer. It sits
at level 2 with no model in the loop at all: your code chunks, embeds, ranks and stops.

Sourced, not measured: the claims below are checked against primary sources, but nothing here has
a recorded run or a scored result file, and the example embeds with a bag-of-words stub rather
than a trained model.

_The web page for this technique includes an interactive step-through of Level 2 · Embeddings and search. The same steps are described in the sections below._

## Practical guidance

Test whether a search box you use is running on keywords alone or also on meaning. Search using
a word that is not literally in the document you expect to find, a synonym, a description, or a
rephrasing rather than the document's own term. Typing "quieter" when the document only states a
decibel rating is the shape of the test. A pure keyword search comes back empty or wrong when the
words do not match. A search backed by embeddings, or a hybrid of the two, has a chance of
finding the right result anyway, because it is comparing meaning, not spelling.

Act on what the test shows. Some search tools expose the choice directly, as a toggle between
"keyword" and "smart" or "semantic" search: when one mode gives you a wrong or missing result,
try the other before concluding the tool cannot find the answer at all. Try it in both
directions, since matching on meaning is what finds a passage worded differently than your
question, and matching on literal words is what reliably finds an exact string, a part number, an
order id, a serial number, that a vector comparison has no special reason to rank first. Running
both and merging them is the hybrid arrangement pgvector documents doing with Postgres full-text
search[3].

If a tool has no such toggle and keeps missing rephrased questions, that is not something the
search box lets you fix: the problem is in how the tool was built, and naming that is more useful
than assuming you typed it wrong. Sometimes a maker documents the mechanism plainly: Microsoft
says Microsoft 365 Copilot builds a vectorized semantic index of an organization's files, in
which material with similar meaning sits close together, and that it runs alongside the ordinary
keyword index rather than replacing it[4]. More often nothing says so. You rarely meet
this labeled "embeddings" or "vector search" at all, since it is usually the mechanism under a
product, such as a chat app answering a question about a document you uploaded, rather than the
product itself; the named examples of that are on the [RAG page](/gradient_ascent/techniques/rag/).

## Implementation details

The example below indexes the same synthetic corpus RAG uses, then answers one query two ways
instead of one: by embedding every chunk and ranking them by cosine similarity to the query, and
separately by BM25 keyword score over the same chunks. It reports where the two result sets agree
and where they diverge, instead of answering the question: this level searches, it does not
answer.

What the example embeds with is not an embedding model, and nothing it returns is evidence about
one. `StubEmbedder` (`examples/common/model.py`) hashes words into 64 buckets and counts them: a
bag of words with no notion that "quiet" and "dBA" are related unless the words themselves
overlap. Run it on a query that shares real words with the corpus ("DW-300 Normal cycle water
use") and its top hit matches keyword search's top hit exactly, which tells you the pipeline
works and nothing about semantics. Run it on "Which dishwasher is quieter, the DW-300 or the
DW-480?" and neither method finds `specs-comparison#2`, the section that actually gives both
decibel ratings, because the word "quieter" never appears in the corpus at all. A trained model is
what would close that gap, and where one would go is `OllamaEmbedder`, behind the same `Embedder`
interface. Read every score below as the shape of the mechanism, not as a result.

`examples/embeddings_search/run.py` (lines 61-89)

```python
def run(
    query: str,
    model: Model | None,
    embedder: Embedder,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    k: int = TOP_K,
) -> Answer:
    del model  # this level searches; it does not answer
    sections = load_sections(corpus_dir)
    tracer.record(kind="code", decided_by="code", title="Chunk the corpus", detail=f"{len(sections)} sections")
    semantic = _semantic_search(query, sections, embedder, k)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Embed the index and rank it by similarity",
        detail=", ".join(f"{s.cite}={score:.2f}" for s, score in semantic),
    )
    keyword = bm25_search(sections, query, k=k)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Rank the same query by keyword (BM25)",
        detail=", ".join(f"{s.cite}={score:.2f}" for s, score in keyword),
    )
    summary = _compare(semantic, keyword)
    tracer.record(kind="code", decided_by="code", title="Compare the two result sets", detail=summary)
    return Answer(text=summary, citations=[s.cite for s, _ in semantic])
```

The similarity function `_semantic_search` calls, just above it in the same file, divides by both
vectors' lengths instead of taking a bare dot product, so the ranking is a true cosine whichever
`Embedder` is plugged in and not only for one that happens to return unit vectors. Every step is
`decided_by: "code"`: what gets embedded, how many results come back, and how the two result sets
get compared are fixed before anything runs. The example implements neither of the two optional
parts above: no hybrid fusion of the two lists it prints, and no reranking.

Run it yourself:

`examples/embeddings_search/README.md` (lines 17-17)

```text
python -m examples.embeddings_search --model stub --question "DW-300 Normal cycle water use"
```

## When you do not need this

Try [level 0, no model at all](/gradient_ascent/techniques/order-zero/), plain keyword search, first if your
questions reliably use the same words as the documents: a part number, an exact phrase, a
serial number. It is simpler, needs no index to keep in sync, and often wins outright on exact
identifiers.

Move to embeddings and search once questions are worded differently than the source text (a
synonym, a paraphrase, a description instead of the term the document uses) which is exactly
where keyword matching stops working.

## Failure modes

### Chunk boundaries split a fact

- **How to notice it:** A number and the sentence explaining it end up in two different chunks, so a search that finds one chunk misses the other half of the answer.
- **How to test for it:** Check whether a fact and the context it needs to be understood ever sit in the same chunk. If a chunk boundary regularly falls in the middle of one idea, the chunking, not the search, is the problem.

### Query and index embedded with different models

- **How to notice it:** Every result comes back with a low, flat similarity score and none of them look related to the query, even for an easy question.
- **How to test for it:** Confirm the model id used to build the index matches the model id used to embed the query. Two different embedding models do not share a vector space, even at the same number of dimensions.

### Exact identifiers get lost in semantic-only search

- **How to notice it:** A search for a part number, an order id, or a serial number returns plausible-looking but wrong results, because nothing in the corpus is a closer semantic match than something else.
- **How to test for it:** Search for a known exact identifier with the semantic path alone, then with keyword search alone. If keyword search wins outright, the system needs the hybrid combination, not a better embedding model.

### Under-trained or low-dimensional embeddings blur unrelated content together

- **How to notice it:** Results include documents with no topical connection to the query at all, not just imperfect ones.
- **How to test for it:** Run a query with almost no literal word overlap with the target passage and see what comes back. This repo's own stub embedder shows the failure directly: a hashing bag of words with only 64 buckets collides often enough that its "semantic" results are sometimes worse than plain keyword search on the same query.

### Stale index

- **How to notice it:** A source document changes and search keeps returning the old text, since the index was built at write time, not read time.
- **How to test for it:** Change a document without rebuilding the index and search for the changed fact. The old embedding is still what gets compared.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 0
- **Chunks indexed:** 79
- **Embedding calls:** 1 (batched)
- **Wall time:** ~40ms

**Compared with RAG (level 2, same corpus).** The same retrieval work RAG does, without the one model call RAG adds afterward to turn the results into an answer.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

`scripts/eval_run.py` will not score this example: it returns a comparison between two result
sets rather than an answer the site's 60-question set can grade, so asking the runner for a
score prints that reason and stops. What it would still be measured on, the same way the
retrieval half of RAG is, is
**citation hit rate**: whether the top-k results for a question's kind actually include the
section the question's grading rule expects. A retrieval-only technique like this one should be
judged on that, kind by kind, separately from whatever answers it up to.

No result file exists for retrieval quality on this technique yet (see `docs/EVALS.md`).

## Run it

**What to monitor.** The distribution of top-result similarity scores across real queries. A growing share of queries with a low top score usually means the index and the query are drifting apart, not that the questions got harder.

**Cost at volume.** Indexing cost scales with corpus size and happens once (or incrementally, as documents change); query-time cost scales with query volume, one small embedding call per query. Reindexing the whole corpus on every change, instead of only the changed documents, is the usual way this gets expensive.

**How it fails in production.** The corpus changes but the index is not rebuilt, so search keeps returning stale text. Or the embedding model gets swapped for a newer one without reindexing everything, so old and new vectors sit in the same index and are no longer comparable to each other.

**What to log.** The query, the top-k results and their similarity scores from each method if running hybrid, the embedding model id and index build date, so a bad result traces back to a stale index or a model mismatch without re-running anything.

## Try it

1. **Use it.** Search a tool you already use for something using a word that does not literally appear in the document you expect to find. Does it still find it, or does it come back empty?
2. **Build it.** Run python -m examples.embeddings_search --model stub --question "Which dishwasher is quieter, the DW-300 or the DW-480?" and compare the two result sets in the output. Neither one finds specs-comparison#2, the section that actually answers this, because the word "quieter" never appears in the corpus.
3. **Either lane.** Change TOP_K from 3 to 6 in examples/embeddings_search/run.py and rerun a query from above. Does the overlap between the semantic and keyword result sets grow?


## Sources

1. [Vector embeddings](https://developers.openai.com/api/docs/guides/embeddings) — OpenAI (API documentation) (accessed 2026-09-19)
2. [Cohere's Rerank Model](https://docs.cohere.com/docs/rerank) — Cohere (documentation) (accessed 2026-09-19)
3. [pgvector](https://github.com/pgvector/pgvector) — pgvector (GitHub README) (accessed 2026-09-19)
4. [Semantic indexing for Microsoft Copilot](https://learn.microsoft.com/en-us/microsoftsearch/semantic-index-for-copilot) — Microsoft (Microsoft Learn) (accessed 2026-09-19)


Last reviewed 2026-09-19.
