The example below is a small store with the three operations that matter: write, recall, and
forget. It is handed a list of facts already stated: the extraction step a real product runs
to decide what is worth keeping is its own model call this example does not duplicate, since
the knowledge graphs example already shows that
shape (one model call, always made, that fills in content rather than choosing what happens
next). What this example shows instead is what happens after writing: recall ranks stored facts
by cosine similarity to a new question, the same mechanism as
embeddings and search and the same bag-of-words
stub embedder, just over a handful of personal facts instead of a document corpus. The scores
it prints show the mechanism and say nothing about how a real embedding model would rank them.
forget removes an entry outright; nothing recalled afterward can include it again, which is the
whole reason the operation exists rather than being another kind of write.
examples/memory/run.py · lines 71–106
def run(
question: str,
model: Model,
embedder: Embedder,
tracer: Tracer,
*,
facts: list[str] | None = None,
forget_ids: list[str] | None = None,
k: int = TOP_K,
) -> Answer:
store = MemoryStore(embedder)
written = [store.write(fact) for fact in (facts or [])]
tracer.record(kind="code", decided_by="code", title="Write facts to memory", detail=f"{len(written)} entries: {', '.join(written) or 'none'}")
forgotten = [eid for eid in (forget_ids or []) if store.forget(eid)]
tracer.record(kind="code", decided_by="code", title="Forget requested entries", detail=", ".join(forgotten) or "none")
recalled = store.recall(question, k=k)
tracer.record(
kind="code",
decided_by="code",
title="Recall memories relevant to the question",
detail=", ".join(f"{e.id}={score:.2f}" for e, score in recalled) or "none recalled",
)
context = "\n".join(f"- {e.text}" for e, _ in recalled) or "(no relevant memories)"
prompt = f"Remembered facts:\n{context}\n\nQuestion: {question}"
messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=prompt)]
completion = model.complete(messages, max_tokens=300)
tracer.record(
kind="model",
decided_by="code",
title="Answer using recalled memory",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
return Answer(text=completion.text, citations=[e.id for e, _ in recalled])
Ask about a dishwasher’s warranty with all four sample facts in memory, and recall surfaces the
purchase date and the fact that the unit is in a rental property: both relevant, since rental
use caps the warranty at 90 days regardless of the standard term. Forget the rental-property fact
first and ask again: the recalled set and the citations both change, and the model is answering a
different question, because that fact no longer exists anywhere the code can reach. Run it
yourself:
examples/memory/README.md · lines 15–15
python -m examples.memory --model stub:scripted
Every step is decided_by: "code": the code always writes what it is given, always searches,
always forgets what it is told to, and asks the model once at the end with whatever recall
turned up.