Level 02 · Added context

Context engineering

Deciding what goes into the request, and caching the parts that repeat.

Sourced

Concept at a glance

Choose what the model can see this time.

SequenceConceptual illustration
Choose what the model can see this time.Available material leads to Assemble context. Assemble context leads to Model request. The model works from the context you actually send, not everything your system knows.Available materialInstructions, files, historyAssemble contextSelect, order, fit, cacheModel requestOnly the chosen materialChoose what the model can see this time.Available material leads to Assemble context. Assemble context leads to Model request. The model works from the context you actually send, not everything your system knows.Available materialInstructions, files, historyAssemble contextSelect, order, fit, cacheModel requestOnly the chosen material
Read the connections in words
  • Available material → Assemble context: Select, order, fit, cache.
  • Assemble context → Model request: Only the chosen material.
Key idea

The model works from the context you actually send, not everything your system knows.

CHOOSE YOUR PERSPECTIVE

Same concept, different task and consequences. Switching starts a fresh walkthrough; prior answers and approvals do not carry over.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Context engineering: see it in practice.

Selecting and arranging the instructions, reference material, and history included in each model request.

What you’ll walk through

Follow the selection of information for a particular request. Compare what is available with what is actually included, then see how source selection changes an otherwise similar response.

The task in this version

Reply about warranty coverage using the current manual, not old case notes.

What you’ll learn to check

A visible request-context tray, version labels, included/excluded evidence, and answer differences.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Reply about warranty coverage using the current manual, not old case notes.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Question: Does water damage qualify? Manual v3: excluded. Old v1 note: sometimes covered.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Some available material may be old, irrelevant, or untrusted. Including everything can bury the evidence needed for this decision.

1 / 6

Apply this to your project

Describe your task to your own model and use Context engineering as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Context engineering is deciding what goes into the request: which instructions, which examples, which documents, and how much of the conversation so far. The model knows only what it was trained on and what the request contains, so everything else is a choice your code makes, on every call.

Two things about that choice matter beyond fitting the content in. Order affects cost: Anthropic and OpenAI both cache a matching prefix of a request and bill the reused part at a lower rate, and both tell you to put the content that never changes first and the content that changes every call last[1][2]. Length affects quality on its own: Chroma’s research across 18 models reports that models do not use their context uniformly, and that performance grows increasingly unreliable as input length grows, even on simple tasks[3]. A longer request is not free just because it still fits.

Context engineering sits at level 2, context. Nothing here searches and nothing is decided by the model: your code fixes what goes in, and in what order, before the request is sent.

This page provides primary references and illustrative examples. The examples are scripted, not measured model runs; source references do not establish the correctness of every implementation or outcome.

Optional: inspect the implementation trace

This separate, scripted trace illustrates the code example discussed in the implementation details. Use it to inspect individual steps, inputs, and outputs.

Context engineering

Assemble one request from fixed parts, cache-friendly order first, and ask once.

Level 2 · Added context
Question + historyQuestion + historyBuild the static blockBuild the static blockFit history to budgetFit history to budgetBuild the final promptBuild the final promptMODELanswers onceanswers onceAnswerAnswerQuestion + historyQuestion + historyBuild the static blockBuild the static blockFit history to budgetFit history to budgetBuild the final promptBuild the final promptMODELanswers onceanswers onceAnswerAnswer
0of 1 step so far chosen by the model
your code chose this stepthe model chose this step

The run, step by step

This is a scripted illustration, not a recorded model run. Its timing and token figures are not measured performance.

STEP 01 / 05Your code chose

The question and the conversation so far arrive

"What is the DW-300's Normal cycle water use?"
2 prior turns of history
0 tokens · 0 ms

Practical guidance

The feature to find is the one that saves instructions and files across conversations instead of inside one. Chat apps name it differently: a project, a space, custom instructions, a saved assistant. Look in the sidebar or the settings for anything that says the material will apply to every new conversation. Anthropic’s Claude Projects keeps a project’s files and instructions available to every conversation inside it: that pinned material is the stable part of the request, and what you type each turn is the part that changes[5]. Other makers’ chat apps have the same feature; the ones this site has checked are listed under Out there at the foot of this page.

Pin three things, and nothing else.

  1. Who the answer is for and what shape it takes. “You are drafting for a twelve-person insurance brokerage. Plain American English, short paragraphs, no bullet lists unless I ask for them.”
  2. The rules that never change. “Never state a premium, a deadline or a policy number that is not in the attached documents. If a document does not say, write ‘not stated’ rather than estimating it.”
  3. The reference files themselves: the style guide, the current rate sheet, the standard letter.

Then type only what changed. “Draft the renewal letter for the account in today’s file. It renews 3/14/2027.”

To check it worked, open a brand new conversation inside the project and ask something only the pinned material can answer, such as what your style guide says about bullet lists. An answer with nothing pasted in means the pinned material is reaching the model. No answer usually means what you wrote went into one conversation rather than into the project, which is a different box in most products.

Two moves when the answers get worse instead of better. Start a fresh conversation rather than continuing one that has run for hours: more in the request measurably costs answer quality, not just money[3], and a new conversation keeps the pinned material while dropping the accumulated back-and-forth. And stop adding files once they stop fitting. Past that point a product stops sending all of it and starts searching instead, which is RAG; Claude Projects makes that switch on its own.

None of this is worth setting up for a question you will ask once, with nothing to reuse. Type the question.

Implementation details

The example below builds one request from parts that change at different rates: system instructions and the reference documents never change between calls; conversation history and the question change every call. It puts the parts that never change first and the parts that change last, which is the ordering a caching backend needs to reuse the front of the request[1] [2]. Nothing here searches for anything (the whole synthetic document set goes in, the same way regardless of the question) so this is what RAG’s “put it all in the window” alternative actually looks like in code.

The static block is built once per call from every document in evals/corpus/, concatenated in a fixed, alphabetical order so its bytes are identical from one question to the next. A token_budget caps the whole request; what is left after the static block is what the conversation history gets to use. When history does not fit, the code drops the oldest turns first and never touches the documents, since the documents are the part later calls can still reuse from cache. This is the one place the code makes a real decision, and it is a fixed rule, not a judgment call: nothing here is decided by the model.

Ordering the request is necessary for caching but not sufficient, and the two makers cited above differ in a way worth reading before you count on a discount. OpenAI documents caching as on by default for supported models, matching a prefix automatically once it clears a minimum cacheable length; that minimum is 1,024 tokens on its newest models and different on earlier ones, and reused tokens are billed at a reduced rate, discounted up to 90%[2]. Anthropic requires you to mark the end of the reusable prefix with a cache_control field, has a per-model minimum below which a marked prompt is silently not cached, and prices a cache read as a fraction of its base input rate (a tenth for most models, lower still for its newest ones) against a write that costs more than an uncached call[1]. Those are the makers’ own published terms, not numbers this site has measured, and the write premium is why caching pays off over repeated calls rather than on the first one.

examples/context_engineering/run.py · lines 58–97
def run(
    question: str,
    model: Model,
    embedder: Embedder | None,
    tracer: Tracer,
    *,
    corpus_dir: Path = DEFAULT_CORPUS_DIR,
    history: list[Message] | None = None,
    token_budget: int = TOKEN_BUDGET,
) -> Answer:
    del embedder  # nothing is retrieved: the whole document set goes in, or none of it does
    static_block, static_tokens = _static_block(corpus_dir)
    tracer.record(
        kind="code",
        decided_by="code",
        title="Assemble the static, cache-friendly block",
        detail="system instructions + the full document set, same on every call",
        tokens_in=static_tokens,
    )
    kept_history = _fit_history(history or [], max(token_budget - static_tokens, 0), tracer)
    dynamic = "\n".join(f"{m.role}: {m.content}" for m in kept_history)
    user_content = static_block + (f"\n\n{dynamic}" if dynamic else "") + f"\n\nQuestion: {question}"
    messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=user_content)]
    tracer.record(
        kind="code",
        decided_by="code",
        title="Build the final prompt",
        detail=f"documents first, then {len(kept_history)} history turn(s), then the question last",
    )
    completion = model.complete(messages, max_tokens=400)
    tracer.record(
        kind="model",
        decided_by="code",
        title="Ask the model once with the full context",
        detail=completion.text[:200],
        tokens_in=completion.tokens_in,
        tokens_out=completion.tokens_out,
        ms=completion.ms,
    )
    return Answer.from_text(completion.text, retrieved_sources=list(load_sections(corpus_dir)))

Every step is decided_by: "code": what goes in, in what order, and what gets cut when it does not fit are all fixed before the model ever sees the request. Run it yourself:

examples/context_engineering/README.md · lines 13–13
python -m examples.context_engineering --model stub:scripted
When you do not need this

Skip this and write a single prompt directly if there is nothing beyond the question itself to include: no documents, no reusable instructions, no history worth keeping. That is plain chat or prompt engineering.

Retrieve instead once the material outgrows the window. This page and RAG are the two answers to the same question, and four things separate them. Size: Anthropic’s engineering team says include a knowledge base under about 200,000 tokens in full and reach for retrieval as it grows past that[4]. Quality: a request that still fits can still answer worse, since performance grows less reliable as input length grows[3], while retrieval keeps the request short whatever the corpus does. Cost: everything you send is billed on every call, discounted to a tenth of the input rate on a cache hit[1] but never to nothing, where a search bills only the few passages it returns. And what you give up by retrieving is the guarantee that the answer’s source was in the request at all: a search can miss the passage, and sending everything cannot.

Failure modes

A cache that never hits

How to notice it
Every call costs and takes as much as the first one, even though most of the request is the same material as last time.
How to test for it
Look at what sits before the first part that changes between calls: a timestamp, a random id, or a reordered document list at or near the front invalidates the cached prefix every call. Then check the two things ordering cannot fix: whether the provider needs caching switched on explicitly for this request, and whether the prefix clears the provider's minimum cacheable length, below which nothing is cached and no error is returned.

Context rot: worse answers from a request that still fits

How to notice it
A question the model could answer easily in a short prompt gets a wrong, vague, or lower-confidence answer once the request grows, with nothing over the model's stated context limit.
How to test for it
Ask the same question with a small slice of the material and with the full set included. A large gap between the two, on a question the full set does not need, is context rot rather than a missing fact.

Silent truncation drops the fact that mattered

How to notice it
The answer is confidently wrong or generic, and the one passage that would have answered it correctly was cut to fit the budget without anyone noticing.
How to test for it
Log what was actually cut, not just that a cut happened. Rerun a failing question with the budget doubled and see whether the answer changes.

Prompt injection through included material

How to notice it
A document contains text written to look like an instruction, and the answer follows it instead of answering the question.
How to test for it
Add a section containing an embedded instruction to the included material and see whether the answer changes to match it.

Stale static content

How to notice it
The static block was built once and reused across many calls; a source document changes and the answer keeps reflecting the old text.
How to test for it
Change a document after the static block has been assembled, without rebuilding it, and ask a question the change affects.

Cost and latency

Illustrative, not measured: no result file exists for this technique yet, so every number below is a worked example. See docs/EVALS.md for how a real run is recorded and scored.

1Model calls, one question
~7,020Tokens in, first call
~64Tokens out
~1.6sWall time
Compared with RAG (level 2, retrieving from the same 12 documents)About 3.8 times the tokens of the illustrated RAG run, since nothing is retrieved: the whole document set goes in on every call, cache discount aside.

How to Evaluate It

60 questionslookupmulti-hopnumericunanswerableconflicting sources

This technique would be scored against the same 60-question synthetic set as every other level, over the document set in evals/corpus/. Because nothing is retrieved, every source is present for every question kind, so multi-hop and conflicting-source questions would test whether the model can still find and join facts inside a long request, not whether the right passage was fetched. Lookup and numeric questions are the ones most exposed to context rot: a fact the model could state easily on its own gets buried in material that has nothing to do with the question.

No result file exists for this technique yet (see docs/EVALS.md), so this page cannot say a number for any of it. Run python scripts/eval_run.py --example context_engineering --model <spec> --dry to project the cost of a real run first. Expect a large number: this is the level that sends the whole document set on every question.

Run it

What to monitor

Cache hit rate on the static block, and tokens billed at the full input rate versus the cached rate. A hit rate that drops for no reason usually means something changed in the part of the request meant to stay fixed.

Cost at volume

Dominated by tokens in, most of which is the same material sent again on every call; the caching discount is what keeps that affordable at volume, so a change that breaks caching is a cost regression even if nothing else about the answers changes.

How it fails in production

Someone adds a per-request value, such as a timestamp or a reordered list, ahead of the static block and the cache stops hitting. Or the reference material grows past the context window and starts getting silently cut, dropping whichever fact happened to be last.

What to log

The token count of the static block versus the dynamic part, whether the call was a cache hit, what (if anything) was trimmed to fit the budget, and the full assembled prompt, so a bad answer traces back to a missing fact rather than a model mistake.

Try it

  1. Use it

    Open a chat app's project or custom-instructions feature. Put a document there instead of pasting it into a message, ask about it, then ask something unrelated in the same project. Does it still see the document?

  2. Build it

    Run python -m examples.context_engineering --model stub:scripted from the repo root, then lower TOKEN_BUDGET in examples/context_engineering/run.py below the static block's 6,934 tokens. The history budget floors at zero, nothing is cut, and the prompt goes out over budget: it governs history, not the block the cache depends on.

  3. Either lane

    Cause a failure mode above on purpose, with the documents in evals/corpus/.

How it connects

Before, after and instead of this

Instead of

Decoded in

Optional: products, tools, and models

3 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Prepare a design review

Send the current requirements and relevant schematic notes, rather than every file and every old conversation.

Out there

Named products, tools and models

Products2
  • ChatGPT ProjectsOpenAI · files and instructions in a chat app
  • Claude ProjectsAnthropic · files and instructions in a chat app
Tools1
  • Deep AgentsLangChain · agent harness

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Prompt caching · Anthropic (Claude Platform documentation) (accessed 09/19/2026)
  2. Prompt caching · OpenAI (API documentation) (accessed 09/19/2026)
  3. Context Rot: How Increasing Input Tokens Impacts LLM Performance · Chroma Research, 07/14/2025 (accessed 09/19/2026)
  4. Contextual Retrieval · Anthropic (Engineering blog) (accessed 09/19/2026)
  5. What are Projects? · Anthropic (Claude Help Center) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page