The example below builds one request from parts that change at different rates: system
instructions and the reference documents never change between calls; conversation history and the
question change every call. It puts the parts that never change first and the parts that change
last, which is the ordering a caching backend needs to reuse the front of the request[1]
[2]. Nothing here searches for anything (the whole synthetic document set goes in, the
same way regardless of the question) so this is what RAG’s
“put it all in the window” alternative actually looks like in code.
The static block is built once per call from every document in evals/corpus/, concatenated in a
fixed, alphabetical order so its bytes are identical from one question to the next. A
token_budget caps the whole request; what is left after the static block is what the
conversation history gets to use. When history does not fit, the code drops the oldest turns
first and never touches the documents, since the documents are the part later calls can still
reuse from cache. This is the one place the code makes a real decision, and it is a fixed rule,
not a judgment call: nothing here is decided by the model.
Ordering the request is necessary for caching but not sufficient, and the two makers cited above
differ in a way worth reading before you count on a discount. OpenAI documents caching as on by
default for supported models, matching a prefix automatically once it clears a minimum cacheable
length; that minimum is 1,024 tokens on its newest models and different on earlier ones, and
reused tokens are billed at a reduced rate, discounted up to 90%[2]. Anthropic requires
you to mark the end of the reusable prefix with a cache_control field, has a per-model minimum
below which a marked prompt is silently not cached, and prices a cache read as a fraction of its
base input rate (a tenth for most models, lower still for its newest ones) against a write that
costs more than an uncached call[1]. Those are the makers’ own published terms, not
numbers this site has measured, and the write premium is why caching pays off over repeated calls
rather than on the first one.
examples/context_engineering/run.py · lines 58–97
def run(
question: str,
model: Model,
embedder: Embedder | None,
tracer: Tracer,
*,
corpus_dir: Path = DEFAULT_CORPUS_DIR,
history: list[Message] | None = None,
token_budget: int = TOKEN_BUDGET,
) -> Answer:
del embedder # nothing is retrieved: the whole document set goes in, or none of it does
static_block, static_tokens = _static_block(corpus_dir)
tracer.record(
kind="code",
decided_by="code",
title="Assemble the static, cache-friendly block",
detail="system instructions + the full document set, same on every call",
tokens_in=static_tokens,
)
kept_history = _fit_history(history or [], max(token_budget - static_tokens, 0), tracer)
dynamic = "\n".join(f"{m.role}: {m.content}" for m in kept_history)
user_content = static_block + (f"\n\n{dynamic}" if dynamic else "") + f"\n\nQuestion: {question}"
messages = [Message(role="system", content=SYSTEM_PROMPT), Message(role="user", content=user_content)]
tracer.record(
kind="code",
decided_by="code",
title="Build the final prompt",
detail=f"documents first, then {len(kept_history)} history turn(s), then the question last",
)
completion = model.complete(messages, max_tokens=400)
tracer.record(
kind="model",
decided_by="code",
title="Ask the model once with the full context",
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
return Answer.from_text(completion.text, retrieved_sources=list(load_sections(corpus_dir)))
Every step is decided_by: "code": what goes in, in what order, and what gets cut when it does
not fit are all fixed before the model ever sees the request. Run it yourself:
examples/context_engineering/README.md · lines 13–13
python -m examples.context_engineering --model stub:scripted