Primary sources
- Prompt caching · Anthropic (Claude Platform Docs) (accessed 09/19/2026)
- Batch API · OpenAI (API documentation) (accessed 09/19/2026)
Spending fewer tokens and less time for the same result: caching, batching, smaller models, shorter context.
Sourced
Concept at a glance
Check quality after each change; cheaper requests are not useful if they fail the task.
A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.
Reducing resource use or latency while preserving an explicitly defined level of quality.
Follow a proposed saving through a quality and validity check. Inspect whether reuse, a smaller model, or fewer calls reduces expense without changing the result people rely on.
Reuse a policy answer only when valid for this user.
Cache-hit/miss cases, scope/version keys, labeled sample costs, latency distribution, and a quality floor.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Reuse a policy answer only when valid for this user.
Authored case. Select any record below; nothing is sent to a model.What changed: Establish the facts supplied for this version of the task.
A cheaper response is not a saving if it creates more correction work. Cached content may be invalid for another user, source version, or time.
Describe your task to your own model and use Cost optimization as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
Cost optimization is choosing, deliberately, which of several ways to spend fewer tokens and less time fits a system, not applying all of them by default. Preserve the required outcome, automation, and human effort when comparing designs. A lower model bill is not a saving if it hands unwanted work back to the user. The decision worksheet classifies a candidate design; it does not establish the cheapest suitable solution. Possible levers include: a smaller model for questions that don’t need a large one (routing), less text in the prompt (context engineering), reusing a prefix instead of resending it (prompt caching), deferring work that doesn’t need an instant answer (batching), capping how long an answer or a loop is allowed to run, and caching a whole answer rather than only a prompt prefix.
This topic is not a level on the ladder; it applies at every level, the way ops does, and every lever below is one of the levers ops names in general. This page goes one level deeper into each one: what it actually saves, and what it risks.
This page is sourced, not measured: every figure below is a maker’s own published rate, read off their page on the date given, and no lever here has been pulled on this site’s own traffic and scored.
Before questioning any line on an AI bill, check whether you are even paying for the right amount of tool for the job: the worksheet walks through exactly that question, and a smaller, cheaper setup that still gets the work done beats optimizing anything built on top of the wrong one. Answer it first.
Once that is settled, two more levers show up as discounts on a vendor’s own pricing page, and both are worth understanding before a bill surprises you. Caching charges more to store something reusable, then charges less every time it gets reused: on the Claude API as its prompt-caching page reads on September 19, 2026, a cache write costs “1.25 times the base input tokens price” for a five-minute cache and “2 times” for a one-hour one, while a cache read costs “0.1 times the base input tokens price”[1]. Those are Anthropic’s figures for Anthropic’s models on that date and say nothing about any other provider, but the shape holds generally: a tool that resends the same instructions or documents in full on every call, instead of reusing them, is paying (and likely charging you) full price for material that has not changed since the last call.
Batching trades a wait for a discount. OpenAI’s Batch API offers a “50% cost discount compared to synchronous APIs” for work that “completes within 24 hours (and often more quickly)” instead of right away[2]. That is a good trade for a nightly report nobody is watching load, and the wrong one for anything a person is waiting on right now.
The check that tells you whether either lever is actually saving anything: ask for the bill’s own breakdown, not just the total. Put it to the vendor in writing: “Show me, per call, which requests hit a cache and which were batched, against which were charged the full rate.” If they cannot produce that split, the discount may not be reaching you even if the feature exists on paper. And if your own usage is occasional rather than steady, the honest answer is that neither lever is worth chasing: caching and batching both pay off on volume, and a handful of questions a week will not reach the point where either matters.
examples/ops/ already has the cost estimator this page’s levers get measured against: it reads
recorded trace files and a price table the caller supplies, and reports cost and latency per
question and per level. Nothing here builds a second one; the levers above all show up as the
same two numbers this estimator already reports; they just change what a trace, or a price table,
looks like going in.
def _cost(model_id: str, tokens_in: int, tokens_out: int, prices: PriceTable) -> float | None:
price = prices.get(model_id)
if price is None:
return None
per_in, per_out = price
return round(tokens_in / 1000 * per_in + tokens_out / 1000 * per_out, 6)The arithmetic is one multiply-and-add per model id: tokens in against the input price, tokens
out against the output price, both per 1,000 tokens. Every lever above changes one of those four
numbers rather than the formula: caching changes which price a call’s input tokens are billed
at, routing changes which row of the price table applies at all, and shorter context changes
tokens_in directly.
No maker’s price is written into this repository anywhere, on purpose: a real price changes without notice and differs enough between providers that hard-coding one would go stale silently and read as this site’s own claim about a real number rather than a caller’s. The caller supplies and dates its own table instead:
def load_price_table(path: Path) -> PriceTable:
"""A price table file: `{"<model_id>": {"in_per_1k": ..., "out_per_1k": ...}, ...}`, USD per
1,000 tokens. The caller states where these numbers came from; nothing here checks that."""
raw = json.loads(path.read_text(encoding="utf-8"))
return {model_id: (float(entry["in_per_1k"]), float(entry["out_per_1k"])) for model_id, entry in raw.items()}Measuring cost per successful task, not per call, is what keeps a cheaper-looking lever honest: a
smaller model that answers three times before getting a question right costs more than a larger
one that answers once, even though every individual call was cheaper. That is exactly what an
eval scored on outcomes, not on call count, is built to
catch: the site’s own citation_hit_rate and score_overall fields, read alongside tokens_in
and tokens_out from the same result file, are what a real version of this comparison would use.
The same two numbers read differently depending who is asking, and this estimator does not care which: a caller divides them however its own setting counts cost. A production line watching 4,000 units a day divides by units and by day, the way ops’s own cost strip does, so a fraction of a cent a call is worth chasing. An engineer validating five prototype boards before a design review divides by runs instead: five calls in an afternoon almost never read a cached prefix a second time inside its window, so caching rarely earns back its own write at that volume, and the number worth comparing against is the time the drafting saved, not the token price. A person writing one measurement’s uncertainty budget into a report is not dividing by anything: if a model drafts the surrounding prose at all, its cost is rounding error against the instrument’s own calibration, and the only cost that matters is being wrong.
tests/test_example_ops.py checks this arithmetic directly and is not mine to extend, but it is
worth reading here: 1,000 input and 1,000 output tokens against a $1/$2-per-1,000-token made-up
table comes to exactly $3.00, and a model id missing from the table reports usd: None, never
$0.00: an unpriced call costs something; the estimator just cannot say how much, which is the
same distinction that matters when a lever changes which model answered and the price table
hasn’t caught up yet.
Avoid adding optimization machinery until there is enough real, sustained traffic for the engineering cost of adding one to be smaller than what it saves: a cache write costs more than a plain call and only earns that back on a second read; a router is a second system to build and keep in sync with whatever it is choosing between.
Move up to a specific lever once its own condition is genuinely true: cache once the same content is read again inside its window, batch once real work can wait a day, route once a real share of questions are easy enough for a cheaper model to answer as well as the current one does.
At high enough volume the question stops being which lever and becomes whether to pay per token at all; running a model locally trades the bill for hardware and the upkeep of a service you now operate.
Each level has one lever that pays before any of the others. The list below is that lever, not a catalog; ops covers what each level costs to run.
max_steps exists for.Cost per successful task, not per call, broken out by which lever touched a given request (cached, batched, routed to a smaller model) so a change to one lever's savings does not hide inside an aggregate that also includes requests it never touched.
Every lever here trades a bit of engineering complexity for a real discount, per the maker figures quoted above; the crossover point where that trade is worth making is a function of real, sustained volume, not of how the levers look on paper.
A lever's condition stops being true without anyone noticing (a cache's content stops repeating, a batch queue starts holding requests someone is actually waiting on) and the saving silently becomes a cost or a complaint instead.
Which lever, if any, touched each call; tokens in and out; cache hit or miss; whether a call was batched or synchronous; and which model actually answered, so cost per successful task can be reconstructed after the fact rather than only estimated in advance.
Pick a task you currently do with a single, most-capable chat model. Using the worksheet's own questions, decide whether a cheaper level or a smaller model would still pass: then actually try it once and compare.
Run python -m examples.ops --demo from the repo root and read the RAG row's tokens_in (1,850). Using the cache-write and cache-read multipliers quoted above, work out by hand how many reads inside a five-minute window it takes before caching that prefix has paid for its own write.
Pick one technique's Cost and latency strip elsewhere on this site. Name the one lever from this page that would cut its cost the most, and the one condition (named in this page's own failure modes) that would have to hold for that lever to actually pay off.
3 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.
Hosted gateway for model calls
Maker’s documentation Checked 09/18/2026Tracing and cost tracking
Maker’s documentation Checked 09/18/2026Tracing and cost tracking
Maker’s documentation Checked 09/18/2026Cache stable material or use a smaller model for easy cases, then rerun the quality checks.
Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page