# Context engineering

_Level 02 · Added context · sourced_

Deciding what goes into the request, and caching the parts that repeat.


## Try this in a recipe
- [Answer a warranty question with evidence](/gradient_ascent/recipes/document-qa.md): Retrieve the relevant policy, answer each part of the question, and distinguish an unknown fact from a retrieval miss.

## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow the selection of information for a particular request. Compare what is available with what is actually included, then see how source selection changes an otherwise similar response.

**Assumptions:** Some available material may be old, irrelevant, or untrusted. Including everything can bury the evidence needed for this decision.

**Design choices:** Choose sources by authority, relevance, freshness, and size. Preserve unresolved questions and provenance when compressing information.

**Request:** Reply about warranty coverage using the current manual, not old case notes.

**Starting evidence:** Question: Does water damage qualify? Manual v3: excluded. Old v1 note: sometimes covered.

**Action and control:** Select v3's relevant clause and identify the old note as superseded. Show what enters the model request.

**Stage records (authored, not executed):**

### Input record

Question: Does water damage qualify? Manual v3: excluded. Old v1 note: sometimes covered.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose sources by authority, relevance, freshness, and size. Preserve unresolved questions and provenance when compressing information.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Select v3's relevant clause and identify the old note as superseded. Show what enters the model request.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Context: question + v3 exclusion + version metadata. Answer: water damage is excluded under the supplied current manual.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

A visible request-context tray, version labels, included/excluded evidence, and answer differences.

If the result falls short:
If a response uses the wrong source, inspect the assembled request first. Restore the missing evidence or correct source precedence before rewriting the answer.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to a project folder, personal research collection, or reporting workspace. Define which sources take precedence and what must survive summarization in your setting.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Context: question + v3 exclusion + version metadata. Answer: water damage is excluded under the supplied current manual.

**Change something — Drop the version metadata:** Two conflicting statements lack a reliable precedence rule. Flag the conflict and seek the applicable version.

**Decision:** Does adding every document necessarily improve the answer?

**Answer:** No; relevance and precedence matter.

**Why:** Compare a relevant excerpt with a stale manual and distracting history; context size is not context quality.

**Review criteria:** A visible request-context tray, version labels, included/excluded evidence, and answer differences.

**Recovery:** If a response uses the wrong source, inspect the assembled request first. Restore the missing evidence or correct source precedence before rewriting the answer.

**Adapt it:** Apply this to a project folder, personal research collection, or reporting workspace. Define which sources take precedence and what must survive summarization in your setting.


## Guided worked example · Everyday life

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow the selection of information for a particular request. Compare what is available with what is actually included, then see how source selection changes an otherwise similar response.

**Assumptions:** Some available material may be old, irrelevant, or untrusted. Including everything can bury the evidence needed for this decision.

**Design choices:** Choose sources by authority, relevance, freshness, and size. Preserve unresolved questions and provenance when compressing information.

**Request:** Plan a weekend trip using my current preferences and the latest opening hours.

**Starting evidence:** Current preference: avoid stairs. Old trip: hiking was welcome. Museum hours: Saturday only.

**Action and control:** Select current access needs and opening hours; mark the old preference as superseded for this trip.

**Stage records (authored, not executed):**

### Input record

Current preference: avoid stairs. Old trip: hiking was welcome. Museum hours: Saturday only.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose sources by authority, relevance, freshness, and size. Preserve unresolved questions and provenance when compressing information.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Select current access needs and opening hours; mark the old preference as superseded for this trip.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Draft suggests a Saturday museum visit after confirming step-free access; no Sunday opening assumed.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Inspect which dated facts and preferences actually enter the request.

If the result falls short:
If a response uses the wrong source, inspect the assembled request first. Restore the missing evidence or correct source precedence before rewriting the answer.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to a project folder, personal research collection, or reporting workspace. Define which sources take precedence and what must survive summarization in your setting.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Draft suggests a Saturday museum visit after confirming step-free access; no Sunday opening assumed.

**Change something — Load only the old itinerary:** The draft may favor stairs and outdated hours. Restore current constraints before recommending activities.

**Decision:** Should an old successful itinerary override current needs?

**Answer:** No; current context controls this task.

**Why:** A past success is reference material, not a permanent preference or current fact.

**Review criteria:** Inspect which dated facts and preferences actually enter the request.

**Recovery:** If a response uses the wrong source, inspect the assembled request first. Restore the missing evidence or correct source precedence before rewriting the answer.

**Adapt it:** Apply this to a project folder, personal research collection, or reporting workspace. Define which sources take precedence and what must survive summarization in your setting.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow the selection of information for a particular request. Compare what is available with what is actually included, then see how source selection changes an otherwise similar response.

**Assumptions:** Some available material may be old, irrelevant, or untrusted. Including everything can bury the evidence needed for this decision.

**Design choices:** Choose sources by authority, relevance, freshness, and size. Preserve unresolved questions and provenance when compressing information.

**Request:** Prepare a patch against the installed driver API, not a newer release.

**Starting evidence:** Project pins driver 2.4. Search result describes 3.0. Local 2.4 docs use measure_voltage, not sample_voltage.

**Action and control:** Select version-matched docs and relevant calling code; exclude incompatible API instructions.

**Stage records (authored, not executed):**

### Input record

Project pins driver 2.4. Search result describes 3.0. Local 2.4 docs use measure_voltage, not sample_voltage.

What changed: Establish the facts supplied for this version of the task.

### Design note

Choose sources by authority, relevance, freshness, and size. Preserve unresolved questions and provenance when compressing information.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Select version-matched docs and relevant calling code; exclude incompatible API instructions.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Patch plan uses measure_voltage from 2.4 and leaves dependency upgrades out of scope.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Cross-check function names and signatures against the pinned version.

If the result falls short:
If a response uses the wrong source, inspect the assembled request first. Restore the missing evidence or correct source precedence before rewriting the answer.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to a project folder, personal research collection, or reporting workspace. Define which sources take precedence and what must survive summarization in your setting.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Patch plan uses measure_voltage from 2.4 and leaves dependency upgrades out of scope.

**Change something — Trim the version pin from the next request:** The model may mix 3.0 syntax into a 2.4 project. Reload the version constraint before editing.

**Decision:** Is the newest documentation always the right context?

**Answer:** No; match the deployed dependency version.

**Why:** Context selection must preserve operational constraints as well as topical relevance.

**Review criteria:** Cross-check function names and signatures against the pinned version.

**Recovery:** If a response uses the wrong source, inspect the assembled request first. Restore the missing evidence or correct source precedence before rewriting the answer.

**Adapt it:** Apply this to a project folder, personal research collection, or reporting workspace. Define which sources take precedence and what must survive summarization in your setting.

Context engineering is deciding what goes into the request: which instructions, which examples,
which documents, and how much of the conversation so far. The model knows only what it was
trained on and what the request contains, so everything else is a choice your code makes, on
every call.

Two things about that choice matter beyond fitting the content in. Order affects cost: Anthropic
and OpenAI both cache a matching prefix of a request and bill the reused part at a lower rate, and
both tell you to put the content that never changes first and the content that changes every call
last[1][2]. Length affects quality on its own: Chroma's research across 18
models reports that models do not use their context uniformly, and that performance grows
increasingly unreliable as input length grows, even on simple tasks[3]. A longer request
is not free just because it still fits.

Context engineering sits at level 2, context. Nothing here searches and nothing is decided by the
model: your code fixes what goes in, and in what order, before the request is sent.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

_The web page for this technique includes an interactive step-through of Level 2 · Context engineering. The same steps are described in the sections below._

## Practical guidance

The feature to find is the one that saves instructions and files across conversations instead of
inside one. Chat apps name it differently: a project, a space, custom instructions, a saved
assistant. Look in the sidebar or the settings for anything that says the material will apply to
every new conversation. Anthropic's Claude Projects keeps a project's files and instructions
available to every conversation inside it: that pinned material is the stable part of the
request, and what you type each turn is the part that changes[5]. Other makers' chat
apps have the same feature; the ones this site has checked are listed under Out there at the foot
of this page.

Pin three things, and nothing else.

1. Who the answer is for and what shape it takes. "You are drafting for a twelve-person insurance
   brokerage. Plain American English, short paragraphs, no bullet lists unless I ask for them."
2. The rules that never change. "Never state a premium, a deadline or a policy number that is not
   in the attached documents. If a document does not say, write 'not stated' rather than
   estimating it."
3. The reference files themselves: the style guide, the current rate sheet, the standard letter.

Then type only what changed. "Draft the renewal letter for the account in today's file. It
renews 3/14/2027."

To check it worked, open a brand new conversation inside the project and ask something only the
pinned material can answer, such as what your style guide says about bullet lists. An answer with
nothing pasted in means the pinned material is reaching the model. No answer usually means what
you wrote went into one conversation rather than into the project, which is a different box in
most products.

Two moves when the answers get worse instead of better. Start a fresh conversation rather than
continuing one that has run for hours: more in the request measurably costs answer quality, not
just money[3], and a new conversation keeps the pinned material while dropping the
accumulated back-and-forth. And stop adding files once they stop fitting. Past that point a
product stops sending all of it and starts searching instead, which is
[RAG](/gradient_ascent/techniques/rag/); Claude Projects makes that switch on its own.

None of this is worth setting up for a question you will ask once, with nothing to reuse. Type
the question.

## Implementation details

The example below builds one request from parts that change at different rates: system
instructions and the reference documents never change between calls; conversation history and the
question change every call. It puts the parts that never change first and the parts that change
last, which is the ordering a caching backend needs to reuse the front of the request[1]
[2]. Nothing here searches for anything (the whole synthetic document set goes in, the
same way regardless of the question) so this is what [RAG](/gradient_ascent/techniques/rag/)'s
"put it all in the window" alternative actually looks like in code.

The static block is built once per call from every document in `evals/corpus/`, concatenated in a
fixed, alphabetical order so its bytes are identical from one question to the next. A
`token_budget` caps the whole request; what is left after the static block is what the
conversation history gets to use. When history does not fit, the code drops the oldest turns
first and never touches the documents, since the documents are the part later calls can still
reuse from cache. This is the one place the code makes a real decision, and it is a fixed rule,
not a judgment call: nothing here is decided by the model.

Ordering the request is necessary for caching but not sufficient, and the two makers cited above
differ in a way worth reading before you count on a discount. OpenAI documents caching as on by
default for supported models, matching a prefix automatically once it clears a minimum cacheable
length; that minimum is 1,024 tokens on its newest models and different on earlier ones, and
reused tokens are billed at a reduced rate, discounted up to 90%[2]. Anthropic requires
you to mark the end of the reusable prefix with a `cache_control` field, has a per-model minimum
below which a marked prompt is silently not cached, and prices a cache read as a fraction of its
base input rate (a tenth for most models, lower still for its newest ones) against a write that
costs more than an uncached call[1]. Those are the makers' own published terms, not
numbers this site has measured, and the write premium is why caching pays off over repeated calls
rather than on the first one.

`examples/context_engineering/run.py` (lines 58-97)

```python
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    history: list[Message] | None = None,
    token_budget: int = TOKEN_BUDGET,
) -> Answer:
    del embedder  # nothing is retrieved: the whole document set goes in, or none of it does
    static_block, static_tokens = _static_block(corpus_dir)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Assemble the static, cache-friendly block",
        detail="system instructions + the full document set, same on every call",
        tokens_in=static_tokens,
    )
    kept_history = _fit_history(history or [], max(token_budget - static_tokens, 0), tracer)
    dynamic = "\n".join(f"{m.role}: {m.content}" for m in kept_history)
    user_content = static_block + (f"\n\n{dynamic}" if dynamic else "") + f"\n\nQuestion: {question}"
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=user_content)]
    tracer.record(
        kind="code",
        decided_by="code",
        title="Build the final prompt",
        detail=f"documents first, then {len(kept_history)} history turn(s), then the question last",
    )
    completion = model.complete(messages, max_tokens=400)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model once with the full context",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    return Answer.from_text(completion.text, retrieved_sources=list(load_sections(corpus_dir)))
```

Every step is `decided_by: "code"`: what goes in, in what order, and what gets cut when it does
not fit are all fixed before the model ever sees the request. Run it yourself:

`examples/context_engineering/README.md` (lines 13-13)

```text
python -m examples.context_engineering --model stub:scripted
```

## When you do not need this

Skip this and write a single prompt directly if there is nothing beyond the question itself to
include: no documents, no reusable instructions, no history worth keeping. That is plain
[chat](/gradient_ascent/techniques/chat/) or [prompt
engineering](/gradient_ascent/techniques/prompt-engineering/).

Retrieve instead once the material outgrows the window. This page and
[RAG](/gradient_ascent/techniques/rag/) are the two answers to the same question, and four
things separate them. Size: Anthropic's engineering team says include a knowledge base under
about 200,000 tokens in full and reach for retrieval as it grows past that[4]. Quality:
a request that still fits can still answer worse, since performance grows less reliable as input
length grows[3], while retrieval keeps the request short whatever the corpus does.
Cost: everything you send is billed on every call, discounted to a tenth of the input rate on a
cache hit[1] but never to nothing, where a search bills only the few passages it
returns. And what you give up by retrieving is the guarantee that the answer's source was in the
request at all: a search can miss the passage, and sending everything cannot.

## Failure modes

### A cache that never hits

- **How to notice it:** Every call costs and takes as much as the first one, even though most of the request is the same material as last time.
- **How to test for it:** Look at what sits before the first part that changes between calls: a timestamp, a random id, or a reordered document list at or near the front invalidates the cached prefix every call. Then check the two things ordering cannot fix: whether the provider needs caching switched on explicitly for this request, and whether the prefix clears the provider's minimum cacheable length, below which nothing is cached and no error is returned.

### Context rot: worse answers from a request that still fits

- **How to notice it:** A question the model could answer easily in a short prompt gets a wrong, vague, or lower-confidence answer once the request grows, with nothing over the model's stated context limit.
- **How to test for it:** Ask the same question with a small slice of the material and with the full set included. A large gap between the two, on a question the full set does not need, is context rot rather than a missing fact.

### Silent truncation drops the fact that mattered

- **How to notice it:** The answer is confidently wrong or generic, and the one passage that would have answered it correctly was cut to fit the budget without anyone noticing.
- **How to test for it:** Log what was actually cut, not just that a cut happened. Rerun a failing question with the budget doubled and see whether the answer changes.

### Prompt injection through included material

- **How to notice it:** A document contains text written to look like an instruction, and the answer follows it instead of answering the question.
- **How to test for it:** Add a section containing an embedded instruction to the included material and see whether the answer changes to match it.

### Stale static content

- **How to notice it:** The static block was built once and reused across many calls; a source document changes and the answer keeps reflecting the old text.
- **How to test for it:** Change a document after the static block has been assembled, without rebuilding it, and ask a question the change affects.

## Cost and latency

_Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example._

- **Model calls, one question:** 1
- **Tokens in, first call:** ~7,020
- **Tokens out:** ~64
- **Wall time:** ~1.6s

**Compared with RAG (level 2, retrieving from the same 12 documents).** About 3.8 times the tokens of the illustrated RAG run, since nothing is retrieved: the whole document set goes in on every call, cache discount aside.

## How to Evaluate It

_Scored on 60 questions across kinds: lookup, multi-hop, numeric, unanswerable, conflicting sources._

This technique would be scored against the same 60-question synthetic set as every other level,
over the document set in `evals/corpus/`. Because nothing is retrieved, every source is present
for every question kind, so multi-hop and conflicting-source questions would test whether the
model can still find and join facts inside a long request, not whether the right passage was
fetched. Lookup and numeric questions are the ones most exposed to context rot: a fact the model
could state easily on its own gets buried in material that has nothing to do with the question.

No result file exists for this technique yet (see `docs/EVALS.md`), so this page cannot say a
number for any of it. Run `python scripts/eval_run.py --example context_engineering --model
<spec> --dry` to project the cost of a real run first. Expect a large number: this is the level
that sends the whole document set on every question.

## Run it

**What to monitor.** Cache hit rate on the static block, and tokens billed at the full input rate versus the cached rate. A hit rate that drops for no reason usually means something changed in the part of the request meant to stay fixed.

**Cost at volume.** Dominated by tokens in, most of which is the same material sent again on every call; the caching discount is what keeps that affordable at volume, so a change that breaks caching is a cost regression even if nothing else about the answers changes.

**How it fails in production.** Someone adds a per-request value, such as a timestamp or a reordered list, ahead of the static block and the cache stops hitting. Or the reference material grows past the context window and starts getting silently cut, dropping whichever fact happened to be last.

**What to log.** The token count of the static block versus the dynamic part, whether the call was a cache hit, what (if anything) was trimmed to fit the budget, and the full assembled prompt, so a bad answer traces back to a missing fact rather than a model mistake.

## Try it

1. **Use it.** Open a chat app's project or custom-instructions feature. Put a document there instead of pasting it into a message, ask about it, then ask something unrelated in the same project. Does it still see the document?
2. **Build it.** Run python -m examples.context_engineering --model stub:scripted from the repo root, then lower TOKEN_BUDGET in examples/context_engineering/run.py below the static block's 6,934 tokens. The history budget floors at zero, nothing is cut, and the prompt goes out over budget: it governs history, not the block the cache depends on.
3. **Either lane.** Cause a failure mode above on purpose, with the documents in evals/corpus/.


## Sources

1. [Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) — Anthropic (Claude Platform documentation) (accessed 2026-09-19)
2. [Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching) — OpenAI (API documentation) (accessed 2026-09-19)
3. [Context Rot: How Increasing Input Tokens Impacts LLM Performance](https://www.trychroma.com/research/context-rot) — Chroma Research, 2025-07-14 (accessed 2026-09-19)
4. [Contextual Retrieval](https://www.anthropic.com/engineering/contextual-retrieval) — Anthropic (Engineering blog) (accessed 2026-09-19)
5. [What are Projects?](https://support.claude.com/en/articles/9517075-what-are-projects) — Anthropic (Claude Help Center) (accessed 2026-09-19)


Last reviewed 2026-09-19.
