# Level 02 · Added context

_Supply relevant information_

Add relevant documents, retrieved passages, or stored information to the current request. This supplies context beyond the model’s training without retraining it. Missing, stale, or misleading material can still produce a poor answer.


## Who decides the next step

You or your software select the information supplied to the model.


## What is at this level

- [Context engineering](/gradient_ascent/techniques/context-engineering/) (sourced): Deciding what goes into the request, and caching the parts that repeat.
- [Embeddings and search](/gradient_ascent/techniques/embeddings-search/) (sourced): Finding text by meaning instead of by keyword.
- [Retrieval-augmented generation (RAG)](/gradient_ascent/techniques/rag/) (measured): Searching your documents and giving the results to the model.
- [Knowledge graphs and GraphRAG](/gradient_ascent/techniques/knowledge-graphs/) (sourced): Storing facts as entities and relations, for questions that span several documents.
- [Memory](/gradient_ascent/techniques/memory/) (sourced): Keeping information from one conversation to the next.

## Upgrade conditions

- **Embeddings and search → Retrieval-augmented generation (RAG):** The person asking wants an answer with the passages behind it, not a ranked list of passages to read themselves.
- **Retrieval-augmented generation (RAG) → Knowledge graphs and GraphRAG:** The answer needs facts joined across documents, or you must show where each fact came from.
- **Retrieval-augmented generation (RAG) → Agentic RAG and deep research:** One retrieval is not enough and the next query depends on what the last one found.
- **Knowledge graphs and GraphRAG → Agentic RAG and deep research:** Which entity to follow next depends on what the last hop returned, so the queries cannot be written in advance.
- **Memory → Long-running tasks:** What has to survive between sessions is work still in progress, not facts about the person you are talking to.

## Named products, tools and models


### Products

- ChatGPT memory — OpenAI · memory in a chat app
- ChatGPT Projects — OpenAI · files and instructions in a chat app
- Claude memory — Anthropic · memory in a chat app
- Claude Projects — Anthropic · files and instructions in a chat app
- Dify — LangGenius · visual workflow builder
- Gemini Notebook — Google · research notebook
- Glean — Glean · workplace search
- Glean Enterprise Graph — Glean · knowledge graph inside a workplace search product
- Microsoft 365 Copilot — Microsoft · workplace assistant
- Microsoft 365 Copilot semantic index — Microsoft · vector index behind a workplace assistant
- Notion AI — Notion · workspace assistant
- Perplexity — Perplexity · answer engine

### Tools

- Chroma — Chroma · vector database
- Cohere Rerank — Cohere · reranker
- Deep Agents — LangChain · agent harness
- FAISS — Meta · vector search library
- GraphRAG — Microsoft · knowledge-graph retrieval
- Haystack — deepset · retrieval framework
- jina-reranker-v3.5 — Jina AI · reranker
- LangChain — LangChain · application framework
- LazyGraphRAG — Microsoft · knowledge-graph retrieval
- Letta — Letta · agents with long-term memory
- LightRAG — open source · knowledge-graph retrieval
- LlamaIndex — LlamaIndex · retrieval framework
- Mem0 — Mem0 · memory layer
- Milvus — Zilliz · vector database
- Neo4j — Neo4j · graph database
- pgvector — open source · vector search in Postgres
- Pinecone — Pinecone · vector database
- Qdrant — Qdrant · vector database
- Weaviate — Weaviate · vector database
- Zep — Zep · memory layer

### Models

- BGE — BAAI · open embedding model
- Cohere Embed v4 — Cohere · embedding model
- Nomic Embed Text v2 — Nomic · open embedding model
- text-embedding-3 — OpenAI · embedding model
- Voyage Embed 4 — Voyage AI · embedding model

## What is still unsolved at this level

_As of 09/19/2026. This block ages faster than the rest of the page._


### Embeddings and search

Which embedding model you pick changes which questions your search answers, and no model is best at everything. Choosing one is still a matter of testing on your own documents rather than reading a ranking.

**What people are trying:** Running a retrieval evaluation over your own document set before committing, and combining keyword search with embedding search so the cases either one misses are covered by the other.

- [MTEB: Massive Text Embedding Benchmark](https://arxiv.org/abs/2210.07316) · arXiv · read 09/19/2026: "We find that no particular text embedding method dominates across all tasks."

### Context engineering

A model that advertises a large context window can still fail to use a fact sitting inside it, once the fact does not share wording with the question. The window is what the model will accept, not what it will reliably find.

**What people are trying:** Benchmarks built to strip the literal word overlap between question and answer, so a model cannot pass by matching text. On the building side, keeping the context smaller and chosen on purpose rather than filling the window because it is there.

- [NoLiMa: Long-Context Evaluation Beyond Literal Matching](https://arxiv.org/abs/2502.05167) · arXiv · read 09/19/2026: "performance degrades significantly as context length increases"

### Retrieval-augmented generation (RAG)

Retrieval cuts invention down and does not end it, and most checkers score a whole answer at once. An answer that is right except for one made-up sentence scores like a right answer, which is the case a reader most needs flagged.

**What people are trying:** Checking groundedness sentence by sentence, for instance by dropping each retrieved passage in turn and watching whether the model's confidence in that sentence collapses. Accuracy at the span level is well short of the answer level.

- [Detecting Hallucinations in Retrieval-Augmented Generation through Grounding-Aware Sensitivity by Perturbation (GASP)](https://arxiv.org/abs/2607.04223) · arXiv · read 09/19/2026: "Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why."

### Memory

When a memory store holds two facts that contradict each other, an old preference and the one that replaced it, there is no settled rule for which should govern the answer. A system can retrieve the stale one and still answer correctly, or retrieve the right one and answer wrongly, and an outcome score cannot tell those apart.

**What people are trying:** Evaluations that inject controlled conflicts into synthetic multi-session histories and score retrieval separately from the final answer, so it is visible which half failed.

- [MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts](https://arxiv.org/abs/2605.20926) · arXiv · read 09/19/2026: "providing limited insight into how systems retrieve and rank temporally valid, factually correct, and contextually applicable memory evidence under conflicting alternatives"
