Level 02

Added context

Add relevant documents, retrieved passages, or stored information to the current request. This supplies context beyond the model’s training without retraining it. Missing, stale, or misleading material can still produce a poor answer.

Level 02

Who decides the next stepYou or your software select the information supplied to the model.
Techniques

What is at this level

Supply relevant information

Context engineering

Sourced

Deciding what goes into the request, and caching the parts that repeat.

Embeddings and search

Sourced

Finding text by meaning instead of by keyword.

Retrieval-augmented generation (RAG)

Measured

Searching your documents and giving the results to the model.

Knowledge graphs and GraphRAG

Sourced

Storing facts as entities and relations, for questions that span several documents.

Memory

Sourced

Keeping information from one conversation to the next.

Upgrade conditions

When something here is not enough

Each line names the failure that justifies moving to a higher level.

Embeddings and search → Retrieval-augmented generation (RAG)

The person asking wants an answer with the passages behind it, not a ranked list of passages to read themselves.

Retrieval-augmented generation (RAG) → Knowledge graphs and GraphRAG

The answer needs facts joined across documents, or you must show where each fact came from.

Retrieval-augmented generation (RAG) → Agentic RAG and deep research

One retrieval is not enough and the next query depends on what the last one found.

Knowledge graphs and GraphRAG → Agentic RAG and deep research

Which entity to follow next depends on what the last hop returned, so the queries cannot be written in advance.

Memory → Long-running tasks

What has to survive between sessions is work still in progress, not facts about the person you are talking to.

Recipes

Jobs that top out here

Each one needs this level and no higher, and says why.

Level 1 + Level 2

Answer questions about a set of documents

Uses RAG, structured output and an eval set. Level 2 is enough because a single search answers most questions.

This example uses level 2
↗
Level 1 + Level 2

Answer questions from a datasheet, a test spec and a change notice

Retrieval over the documents an engineer already has, answered with citations that can be checked. The case that matters is a change notice contradicting the datasheet on one number, where the right answer depends on the board revision. Level 2 is enough because one search finds the passage.

This example uses level 2
↗
Out there

Named products, tools and models

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Products that work this way12
  • ChatGPT memoryOpenAI · memory in a chat app
  • ChatGPT ProjectsOpenAI · files and instructions in a chat app
  • Claude memoryAnthropic · memory in a chat app
  • Claude ProjectsAnthropic · files and instructions in a chat app
  • DifyLangGenius · visual workflow builder
  • Gemini NotebookGoogle · research notebook · formerly NotebookLM
  • GleanGlean · workplace search
  • Glean Enterprise GraphGlean · knowledge graph inside a workplace search product
  • Microsoft 365 CopilotMicrosoft · workplace assistant
  • Microsoft 365 Copilot semantic indexMicrosoft · vector index behind a workplace assistant
  • Notion AINotion · workspace assistant
  • PerplexityPerplexity · answer engine
Tools for building it20
  • ChromaChroma · vector database
  • Cohere RerankCohere · reranker
  • Deep AgentsLangChain · agent harness
  • FAISSMeta · vector search library
  • GraphRAGMicrosoft · knowledge-graph retrieval
  • Haystackdeepset · retrieval framework
  • jina-reranker-v3.5Jina AI · reranker
  • LangChainLangChain · application framework
  • LazyGraphRAGMicrosoft · knowledge-graph retrieval
  • LettaLetta · agents with long-term memory · formerly MemGPT
  • LightRAGopen source · knowledge-graph retrieval
  • LlamaIndexLlamaIndex · retrieval framework
  • Mem0Mem0 · memory layer
  • MilvusZilliz · vector database
  • Neo4jNeo4j · graph database
  • pgvectoropen source · vector search in Postgres
  • PineconePinecone · vector database
  • QdrantQdrant · vector database
  • WeaviateWeaviate · vector database
  • ZepZep · memory layer
Models5
  • BGEBAAI · open embedding model
  • Cohere Embed v4Cohere · embedding model · formerly Cohere Embed, generic
  • Nomic Embed Text v2Nomic · open embedding model · formerly Nomic Embed
  • text-embedding-3OpenAI · embedding model
  • Voyage Embed 4Voyage AI · embedding model · formerly Voyage embeddings
Frontier

What is still unsolved here

Open problems at this level, what people are trying, and the source each rests on. Read 09/19/2026. This block ages faster than the rest of the page, and nothing in it predicts which approach wins.

Which embedding model you pick changes which questions your search answers, and no model is best at everything. Choosing one is still a matter of testing on your own documents rather than reading a ranking.

What people are trying

Running a retrieval evaluation over your own document set before committing, and combining keyword search with embedding search so the cases either one misses are covered by the other.

Where it bites: Embeddings and search

A model that advertises a large context window can still fail to use a fact sitting inside it, once the fact does not share wording with the question. The window is what the model will accept, not what it will reliably find.

What people are trying

Benchmarks built to strip the literal word overlap between question and answer, so a model cannot pass by matching text. On the building side, keeping the context smaller and chosen on purpose rather than filling the window because it is there.

Where it bites: Context engineering

Retrieval cuts invention down and does not end it, and most checkers score a whole answer at once. An answer that is right except for one made-up sentence scores like a right answer, which is the case a reader most needs flagged.

What people are trying

Checking groundedness sentence by sentence, for instance by dropping each retrieved passage in turn and watching whether the model's confidence in that sentence collapses. Accuracy at the span level is well short of the answer level.

Where it bites: Retrieval-augmented generation (RAG)

When a memory store holds two facts that contradict each other, an old preference and the one that replaced it, there is no settled rule for which should govern the answer. A system can retrieve the stale one and still answer correctly, or retrieve the right one and answer wrongly, and an outcome score cannot tell those apart.

What people are trying

Evaluations that inject controlled conflicts into synthetic multi-session histories and score retrieval separately from the final answer, so it is visible which half failed.

Where it bites: Memory

← Level 01 · Direct prompting

Pages at this level last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page