Level 02 · Added context

Embeddings and search

Finding text by meaning instead of by keyword.

Sourced

Concept at a glance

SequenceConceptual illustration
Find nearby meanings in a shared space.Query + passages leads to Similarity search. Similarity search leads to Ranked passages. Similarity finds candidates; it does not prove that a passage answers the question.Query + passagesEncode as vectorsSimilarity searchFind nearby vectorsRanked passagesRead the best matchesFind nearby meanings in a shared space.Query + passages leads to Similarity search. Similarity search leads to Ranked passages. Similarity finds candidates; it does not prove that a passage answers the question.Query + passagesEncode as vectorsSimilarity searchFind nearby vectorsRanked passagesRead the best matches
Read the connections in words
  • Query + passages → Similarity search: Find nearby vectors.
  • Similarity search → Ranked passages: Read the best matches.
Key idea

Similarity finds candidates; it does not prove that a passage answers the question.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Embeddings and search: see it in practice.

Representing content numerically and retrieving similar items, often combined with keyword or metadata search.

What you’ll walk through

Follow a meaning-based query into candidate results. Examine why a similar passage can be useful for discovery while still being the wrong item, revision, or answer.

The task in this version

Find troubleshooting guidance for a knocking sound in pump AX-20.

What you’ll learn to check

Ranked snippets, model filters, relevance judgments, and a case where hybrid search is preferable.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Find troubleshooting guidance for a knocking sound in pump AX-20.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Documents: A, AX-20 knocking; B, AX-30 vibration; C, AX-20 electrical error.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Similarity scores rank candidates; they do not certify correctness. The collection and its metadata determine what can be found.

1 / 6

Apply this to your project

Describe your task to your own model and use Embeddings and search as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

An embedding is a list of floating-point numbers standing in for a piece of text. OpenAI’s documentation states what makes that useful: the distance between two embeddings measures how related the texts are, small distances meaning high relatedness, and it recommends cosine similarity (how closely two vectors point the same way) for the comparison[1].

Search built on this has three fixed parts: documents are cut into chunks small enough to retrieve on their own; every chunk is embedded once into a vector index, a database such as pgvector, Pinecone, Weaviate or Qdrant; and a question is embedded the same way and ranked against that index. Two more are optional. Hybrid search runs a keyword index alongside the vectors and merges the two lists, which is what pgvector documents doing with Postgres full-text search[3]. Reranking re-scores a larger first cut of candidates with a second model: Cohere describes its rerank models as sorting text by relevance to a query, over results an existing search already returned[2].

This is the retrieval half of RAG without the answer. It sits at level 2 with no model in the loop at all: your code chunks, embeds, ranks and stops.

Sourced, not measured: the claims below are checked against primary sources, but nothing here has a recorded run or a scored result file, and the example embeds with a bag-of-words stub rather than a trained model.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Embeddings and search

Index the corpus once, then rank it two ways for the same query: by embedding, and by keyword.

Level 2 · Added context
QueryQueryChunk the corpusChunk the corpusVector indexVector indexRank by similarityRank by similarityRank by keywordRank by keywordCompared result setsCompared result setsQueryQueryChunk the corpusChunk the corpusVector indexVector indexRank by similarityRank by similarityRank by keywordRank by keywordCompared result setsCompared result sets
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

The query arrives

"DW-300 Normal cycle water use"
0 tokens · 0 ms

Practical guidance

Test whether a search box you use is running on keywords alone or also on meaning. Search using a word that is not literally in the document you expect to find, a synonym, a description, or a rephrasing rather than the document’s own term. Typing “quieter” when the document only states a decibel rating is the shape of the test. A pure keyword search comes back empty or wrong when the words do not match. A search backed by embeddings, or a hybrid of the two, has a chance of finding the right result anyway, because it is comparing meaning, not spelling.

Act on what the test shows. Some search tools expose the choice directly, as a toggle between “keyword” and “smart” or “semantic” search: when one mode gives you a wrong or missing result, try the other before concluding the tool cannot find the answer at all. Try it in both directions, since matching on meaning is what finds a passage worded differently than your question, and matching on literal words is what reliably finds an exact string, a part number, an order id, a serial number, that a vector comparison has no special reason to rank first. Running both and merging them is the hybrid arrangement pgvector documents doing with Postgres full-text search[3].

If a tool has no such toggle and keeps missing rephrased questions, that is not something the search box lets you fix: the problem is in how the tool was built, and naming that is more useful than assuming you typed it wrong. Sometimes a maker documents the mechanism plainly: Microsoft says Microsoft 365 Copilot builds a vectorized semantic index of an organization’s files, in which material with similar meaning sits close together, and that it runs alongside the ordinary keyword index rather than replacing it[4]. More often nothing says so. You rarely meet this labeled “embeddings” or “vector search” at all, since it is usually the mechanism under a product, such as a chat app answering a question about a document you uploaded, rather than the product itself; the named examples of that are on the RAG page.

Implementation details

The example below indexes the same synthetic corpus RAG uses, then answers one query two ways instead of one: by embedding every chunk and ranking them by cosine similarity to the query, and separately by BM25 keyword score over the same chunks. It reports where the two result sets agree and where they diverge, instead of answering the question: this level searches, it does not answer.

What the example embeds with is not an embedding model, and nothing it returns is evidence about one. StubEmbedder (examples/common/model.py) hashes words into 64 buckets and counts them: a bag of words with no notion that “quiet” and “dBA” are related unless the words themselves overlap. Run it on a query that shares real words with the corpus (“DW-300 Normal cycle water use”) and its top hit matches keyword search’s top hit exactly, which tells you the pipeline works and nothing about semantics. Run it on “Which dishwasher is quieter, the DW-300 or the DW-480?” and neither method finds specs-comparison#2, the section that actually gives both decibel ratings, because the word “quieter” never appears in the corpus at all. A trained model is what would close that gap, and where one would go is OllamaEmbedder, behind the same Embedder interface. Read every score below as the shape of the mechanism, not as a result.

examples/embeddings_search/run.py · lines 61–89
def run(
    query: str,
    model: Model | None,
    embedder: Embedder,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    k: int = TOP_K,
) -> Answer:
    del model  # this level searches; it does not answer
    sections = load_sections(corpus_dir)
    tracer.record(kind="code", decided_by="code", title="Chunk the corpus", detail=f"{len(sections)} sections")
    semantic = _semantic_search(query, sections, embedder, k)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Embed the index and rank it by similarity",
        detail=", ".join(f"{s.cite}={score:.2f}" for s, score in semantic),
    )
    keyword = bm25_search(sections, query, k=k)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Rank the same query by keyword (BM25)",
        detail=", ".join(f"{s.cite}={score:.2f}" for s, score in keyword),
    )
    summary = _compare(semantic, keyword)
    tracer.record(kind="code", decided_by="code", title="Compare the two result sets", detail=summary)
    return Answer(text=summary, citations=[s.cite for s, _ in semantic])

The similarity function _semantic_search calls, just above it in the same file, divides by both vectors’ lengths instead of taking a bare dot product, so the ranking is a true cosine whichever Embedder is plugged in and not only for one that happens to return unit vectors. Every step is decided_by: "code": what gets embedded, how many results come back, and how the two result sets get compared are fixed before anything runs. The example implements neither of the two optional parts above: no hybrid fusion of the two lists it prints, and no reranking.

Run it yourself:

examples/embeddings_search/README.md · lines 17–17
python -m examples.embeddings_search --model stub --question "DW-300 Normal cycle water use"
When you do not need this

Try level 0, no model at all, plain keyword search, first if your questions reliably use the same words as the documents: a part number, an exact phrase, a serial number. It is simpler, needs no index to keep in sync, and often wins outright on exact identifiers.

Move to embeddings and search once questions are worded differently than the source text (a synonym, a paraphrase, a description instead of the term the document uses) which is exactly where keyword matching stops working.

Failure modes

Chunk boundaries split a fact

How to notice it
A number and the sentence explaining it end up in two different chunks, so a search that finds one chunk misses the other half of the answer.
How to test for it
Check whether a fact and the context it needs to be understood ever sit in the same chunk. If a chunk boundary regularly falls in the middle of one idea, the chunking, not the search, is the problem.

Query and index embedded with different models

How to notice it
Every result comes back with a low, flat similarity score and none of them look related to the query, even for an easy question.
How to test for it
Confirm the model id used to build the index matches the model id used to embed the query. Two different embedding models do not share a vector space, even at the same number of dimensions.

Exact identifiers get lost in semantic-only search

How to notice it
A search for a part number, an order id, or a serial number returns plausible-looking but wrong results, because nothing in the corpus is a closer semantic match than something else.
How to test for it
Search for a known exact identifier with the semantic path alone, then with keyword search alone. If keyword search wins outright, the system needs the hybrid combination, not a better embedding model.

Under-trained or low-dimensional embeddings blur unrelated content together

How to notice it
Results include documents with no topical connection to the query at all, not just imperfect ones.
How to test for it
Run a query with almost no literal word overlap with the target passage and see what comes back. This repo's own stub embedder shows the failure directly: a hashing bag of words with only 64 buckets collides often enough that its "semantic" results are sometimes worse than plain keyword search on the same query.

Stale index

How to notice it
A source document changes and search keeps returning the old text, since the index was built at write time, not read time.
How to test for it
Change a document without rebuilding the index and search for the changed fact. The old embedding is still what gets compared.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

0Model calls, one question
79Chunks indexed
1 (batched)Embedding calls
~40msWall time
Compared with RAG (level 2, same corpus)The same retrieval work RAG does, without the one model call RAG adds afterward to turn the results into an answer.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

scripts/eval_run.py will not score this example: it returns a comparison between two result sets rather than an answer the site’s 60-question set can grade, so asking the runner for a score prints that reason and stops. What it would still be measured on, the same way the retrieval half of RAG is, is citation hit rate: whether the top-k results for a question’s kind actually include the section the question’s grading rule expects. A retrieval-only technique like this one should be judged on that, kind by kind, separately from whatever answers it up to.

No result file exists for retrieval quality on this technique yet (see docs/EVALS.md).

Run it

What to monitor

The distribution of top-result similarity scores across real queries. A growing share of queries with a low top score usually means the index and the query are drifting apart, not that the questions got harder.

Cost at volume

Indexing cost scales with corpus size and happens once (or incrementally, as documents change); query-time cost scales with query volume, one small embedding call per query. Reindexing the whole corpus on every change, instead of only the changed documents, is the usual way this gets expensive.

How it fails in production

The corpus changes but the index is not rebuilt, so search keeps returning stale text. Or the embedding model gets swapped for a newer one without reindexing everything, so old and new vectors sit in the same index and are no longer comparable to each other.

What to log

The query, the top-k results and their similarity scores from each method if running hybrid, the embedding model id and index build date, so a bad result traces back to a stale index or a model mismatch without re-running anything.

Try it

  1. Use it

    Search a tool you already use for something using a word that does not literally appear in the document you expect to find. Does it still find it, or does it come back empty?

  2. Build it

    Run python -m examples.embeddings_search --model stub --question "Which dishwasher is quieter, the DW-300 or the DW-480?" and compare the two result sets in the output. Neither one finds specs-comparison#2, the section that actually answers this, because the word "quieter" never appears in the corpus.

  3. Either lane

    Change TOP_K from 3 to 6 in examples/embeddings_search/run.py and rerun a query from above. Does the overlap between the semantic and keyword result sets grow?

How it connects

Before, after and instead of this

Move up when

Pages that need this one

Decoded in

Optional: products, tools, and models

15 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 9 more examples
In practice

Find differently worded instructions

Search for “water use” and retrieve passages about “consumption,” then inspect the matches.

Out there

Named products, tools and models

Products1
  • Microsoft 365 Copilot semantic indexMicrosoft · vector index behind a workplace assistant
Tools9
  • ChromaChroma · vector database
  • Cohere RerankCohere · reranker
  • FAISSMeta · vector search library
  • jina-reranker-v3.5Jina AI · reranker
  • MilvusZilliz · vector database
  • pgvectoropen source · vector search in Postgres
  • PineconePinecone · vector database
  • QdrantQdrant · vector database
  • WeaviateWeaviate · vector database
Models5
  • BGEBAAI · open embedding model
  • Cohere Embed v4Cohere · embedding model · formerly Cohere Embed, generic
  • Nomic Embed Text v2Nomic · open embedding model · formerly Nomic Embed
  • text-embedding-3OpenAI · embedding model
  • Voyage Embed 4Voyage AI · embedding model · formerly Voyage embeddings

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Vector embeddings · OpenAI (API documentation) (accessed 09/19/2026)
  2. Cohere's Rerank Model · Cohere (documentation) (accessed 09/19/2026)
  3. pgvector · pgvector (GitHub README) (accessed 09/19/2026)
  4. Semantic indexing for Microsoft Copilot · Microsoft (Microsoft Learn) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page