Level 02 · Added context

Memory

Keeping information from one conversation to the next.

Sourced

Concept at a glance

Carry selected information into a later conversation.

SequenceConceptual illustration
Carry selected information into a later conversation.Conversation now leads to Memory store. Memory store leads to Conversation later. Memory is a store you maintain and retrieve from, not unlimited context.Conversation nowChoose what to rememberMemory storeSave, update, or forgetConversation laterRetrieve relevant memoriesCarry selected information into a later conversation.Conversation now leads to Memory store. Memory store leads to Conversation later. Memory is a store you maintain and retrieve from, not unlimited context.Conversation nowChoose what to rememberMemory storeSave, update, or forgetConversation laterRetrieve relevant memories
Read the connections in words
  • Conversation now → Memory store: Save, update, or forget.
  • Memory store → Conversation later: Retrieve relevant memories.
Key idea

Memory is a store you maintain and retrieve from, not unlimited context.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Memory: see it in practice.

Persisting selected information across interactions and retrieving it when relevant.

What you’ll walk through

Follow a preference or prior fact from one interaction into a later task. See when reusing it helps and when correction, expiry, or a different context should override it.

The task in this version

Remember that I prefer planning calls after 3 pm.

What you’ll learn to check

A memory card with origin, update, removal, and a new-session test showing what was actually loaded.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Everyday lifeAn authored case with its own evidence, changed condition, and decision.
The task in this example

Remember that I prefer planning calls after 3 pm.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
User explicitly states a recurring preference. No preference was previously stored.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Stored information needs a scope and a way to correct it. A past preference is not necessarily a permanent rule or authorization.

1 / 6

Apply this to your project

Describe your task to your own model and use Memory as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Memory is keeping information from one conversation to the next. A chat has no memory of its own beyond the messages inside it, so a product that seems to remember you is writing facts down somewhere and reading them back into a later, otherwise unrelated conversation.

Two different things get called memory. A summary is a compressed account of what happened, cheap to reread, but whatever it leaves out is gone. A record is closer to the raw fact (a purchase on a date, a preference stated once) kept on its own and searched later. Letta draws that line in its own product: unlike memory blocks, “archival memory fragments cannot be pinned to the context window, and must be queried on-demand via tools”[3]. Mem0, which sells memory infrastructure for other people’s products rather than a chat app, states the record approach outright: memories accumulate, nothing is overwritten, and a search returns the relevant ones[2].

Memory sits at level 2, context: every write, search and deletion here is your code’s, and the model only answers with what it was handed. This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Memory

Write facts across conversations, recall the relevant ones, forget on request.

Level 2 · Added context
Facts from earlier turnsFacts fromearlier turnsWrite to memoryWrite to memoryForget on requestForget on requestMemory storeMemory storeMODELanswers onceanswers onceAnswerAnswerFacts from earlier turnsFacts fromearlier turnsWrite to memoryWrite to memoryForget on requestForget on requestMemory storeMemory storeMODELanswers onceanswers onceAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

Facts from earlier conversations arrive

"Owns a Halvorsen DW-300, purchased 2024-03-15."
"The DW-300 is in a rental property."
"Prefers email for service updates."
"Also owns a DR-520 dryer, purchased 2025-01-10."
0 tokens · 0 ms

Practical guidance

Open the chat app’s memory settings. Three things are worth reading there, in order, and each one tells you what to do next. Several chat apps have the feature now; the ones this site has checked are listed under Out there at the foot of this page, and they do not all behave the same way, so check the one you actually use rather than assuming what you learned from another.

What is stored. Anthropic documents Claude’s answer: everything Claude memory holds is listed by topic in the settings, and you can open any topic to read it, edit it, or delete that one on its own; by default it excludes personal or sensitive subject matter such as health, race, religious beliefs, politics and gender identity unless you turn that on[1]. Read the list once. If an entry is wrong or stale, edit or delete that entry rather than resetting everything.

Who can see it. In Claude, memory is per person: Anthropic says an account’s owners cannot view or edit an individual user’s memories, and each project keeps its own separate memory space[1]. That is one maker’s design, not an industry rule. Check this setting specifically before you say something in front of a shared or work account; if it does not say memory is private to you, assume it is not.

How to get rid of it. Two controls that sound similar are not. Anthropic documents pausing as keeping existing memories while stopping new ones, and resetting as permanently deleting all of them, with no undo[1]. Pause for a clean slate going forward without losing what is already useful; reset only when you want it all gone. And deleting a conversation does not delete what it taught: Anthropic states that when a conversation expires or is deleted, the memory entries generated from it are not removed[1], so delete those separately if you want the fact itself gone, not just the transcript.

None of this is a reason to avoid the feature. A memory that quietly gets everything right looks, from outside, exactly like one that got something wrong a year ago and nobody has read since, which is the reason to open the settings page once rather than never.

Implementation details

The example below is a small store with the three operations that matter: write, recall, and forget. It is handed a list of facts already stated: the extraction step a real product runs to decide what is worth keeping is its own model call this example does not duplicate, since the knowledge graphs example already shows that shape (one model call, always made, that fills in content rather than choosing what happens next). What this example shows instead is what happens after writing: recall ranks stored facts by cosine similarity to a new question, the same mechanism as embeddings and search and the same bag-of-words stub embedder, just over a handful of personal facts instead of a document corpus. The scores it prints show the mechanism and say nothing about how a real embedding model would rank them. forget removes an entry outright; nothing recalled afterward can include it again, which is the whole reason the operation exists rather than being another kind of write.

examples/memory/run.py · lines 71–106
def run(
    question: str,
    model: Model,
    embedder: Embedder,
    tracer: Tracer,
    *,
    facts: list[str] | None = None,
    forget_ids: list[str] | None = None,
    k: int = TOP_K,
) -> Answer:
    store = MemoryStore(embedder)
    written = [store.write(fact) for fact in (facts or [])]
    tracer.record(kind="code", decided_by="code", title="Write facts to memory", detail=f"{len(written)} entries: {', '.join(written) or 'none'}")
    forgotten = [eid for eid in (forget_ids or []) if store.forget(eid)]
    tracer.record(kind="code", decided_by="code", title="Forget requested entries", detail=", ".join(forgotten) or "none")
    recalled = store.recall(question, k=k)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Recall memories relevant to the question",
        detail=", ".join(f"{e.id}={score:.2f}" for e, score in recalled) or "none recalled",
    )
    context = "\n".join(f"- {e.text}" for e, _ in recalled) or "(no relevant memories)"
    prompt = f"Remembered facts:\n{context}\n\nQuestion: {question}"
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=prompt)]
    completion = model.complete(messages, max_tokens=300)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Answer using recalled memory",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    return Answer(text=completion.text, citations=[e.id for e, _ in recalled])

Ask about a dishwasher’s warranty with all four sample facts in memory, and recall surfaces the purchase date and the fact that the unit is in a rental property: both relevant, since rental use caps the warranty at 90 days regardless of the standard term. Forget the rental-property fact first and ask again: the recalled set and the citations both change, and the model is answering a different question, because that fact no longer exists anywhere the code can reach. Run it yourself:

examples/memory/README.md · lines 15–15
python -m examples.memory --model stub:scripted

Every step is decided_by: "code": the code always writes what it is given, always searches, always forgets what it is told to, and asks the model once at the end with whatever recall turned up.

When you do not need this

Try context engineering first if everything relevant is already in the current conversation. Memory only matters for information that has to survive after a conversation ends: a single session never needs it.

Move to memory once a product needs to answer questions using something a user said in a different, earlier conversation, not just the one open right now.

Memory is one of the techniques behind the answer people in conversation, looking things up and taking small actions job shape, wherever a good reply depends on what an earlier conversation already said.

Failure modes

Everything gets written and nothing gets pruned

How to notice it
The memory store grows without bound, recall gets slower, and old, stale, or contradicted facts start outranking current ones for no reason a user can see.
How to test for it
Check whether anything ever gets removed automatically, and whether a fact that was later corrected by the user still shows up in recall.

Recall surfaces a plausible but wrong memory

How to notice it
An answer confidently uses a fact that sounds related to the question but is not the one that actually applies, the same failure mode embedding-based search has generally.
How to test for it
Ask a question with two stored facts that are superficially similar but say different things, and check which one recall actually returns.

A deleted fact keeps influencing answers anyway

How to notice it
A user deletes a memory, but an answer still reflects it, because the fact was already folded into a summary, a cached embedding, or a derived record that deletion never touched.
How to test for it
Delete a fact after it has already been used once, then ask a new question that only the deleted fact could answer. If the old answer still comes through in any form, deletion is not reaching everywhere the fact was copied to.

Memory written in one context leaks into another

How to notice it
A fact stated in one setting (a work project, a shared account) surfaces in an unrelated one where it does not belong, especially on a shared or team plan.
How to test for it
Write a fact under one context or project and check whether it is recalled from a different, unrelated one that should not have access to it.

No relevant memory, but the model answers as if there were

How to notice it
Recall returns nothing useful, and the answer states something confidently anyway instead of saying it does not know.
How to test for it
Ask a question with no relevant fact in memory at all, and confirm the answer says so rather than guessing.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls, one question
4Memories written
3Memories recalled
~1.4sWall time
Compared with RAG (level 2, a document corpus instead of a personal fact store)Recall here runs over a handful of memories instead of a whole document corpus, so the search itself is close to free. The ongoing cost this example does not show is deciding what is worth writing in the first place, which a real product spends a model call on for every conversation, not just once.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

The site’s shared 60-question set is asked within a single sitting over one document set, so it does not test what memory is actually for: a fact stated in one session and needed again in a separate, later one. A fair eval for this technique would need its own question set, written as pairs across sessions: a fact stated in session one, a question in session two that depends on recalling it, and a check that a fact explicitly forgotten between sessions no longer affects the answer.

So scripts/eval_run.py will not score this example: asking it to prints that reason and stops, rather than returning a number measured on the wrong thing (see docs/EVALS.md). The three measurements that would mean something here are recall of a fact written in an earlier session, the share of answers that use a recalled fact when one applies, and a check that a forgotten fact never reaches an answer again.

Run it

What to monitor

How many memories exist per user and how fast that count grows, recall hit rate on real questions, and the count of forget requests that succeed versus the count still pending. A pending forget older than a few minutes is worth a page in its own right.

Cost at volume

Storage and recall both scale with memory count per user, not with question volume; a store that never prunes or expires anything gets slower and less relevant over time even if nothing is technically wrong with it.

How it fails in production

A user asks to forget something and the entry disappears from the visible list, but a summary or a cached copy made before the deletion keeps influencing answers. Or a memory written for one context surfaces in an unrelated one on a shared account.

What to log

What was written and when, what was recalled for a given question and its relevance score, what was forgotten and when, and whether a forget request actually reached every place the fact had been copied to.

Try it

  1. Use it

    Open a chat app's memory settings and read what it says it has stored about you. Delete one entry, start a new conversation, and ask something that entry would have answered. Does the answer still reflect it?

  2. Build it

    Run python -m examples.memory --model stub:scripted from the repo root. Recall returns m0, m1 and m2, and the answer says no: the dishwasher is in a rental property, which caps coverage at 90 days from a purchase date long past. Ask about the dryer instead (--question "Is my dryer still under warranty?") and the recalled line becomes m3, m0, m1 while the answer does not move. Recall is the only part of this run that responds to what you asked: the reply is a fixture for the first question.

  3. Either lane

    Cause a failure mode above on purpose, using the synthetic facts in examples/memory/__main__.py.

How it connects

Before, after and instead of this

Move up when

  • Long-running tasksWhat has to survive between sessions is work still in progress, not facts about the person you are talking to.

Often used with

Pages that need this one

Optional: products, tools, and models

6 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Remember a project preference

Save the preferred units from an earlier conversation and retrieve them when preparing the next report.

Out there

Named products, tools and models

Products2
  • ChatGPT memoryOpenAI · memory in a chat app
  • Claude memoryAnthropic · memory in a chat app
Tools4
  • Deep AgentsLangChain · agent harness
  • LettaLetta · agents with long-term memory · formerly MemGPT
  • Mem0Mem0 · memory layer
  • ZepZep · memory layer

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Use Claude's chat search and memory to build on previous context · Anthropic (Claude Help Center) (accessed 09/19/2026)
  2. mem0ai/mem0 · Mem0 (GitHub README) (accessed 09/19/2026)
  3. Archival memory · Letta (documentation) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page