Topics at every level

Cost optimization

Spending fewer tokens and less time for the same result: caching, batching, smaller models, shorter context.

Sourced

Concept at a glance

Spend less where the task allows it.

Decision pathsConceptual illustration
Spend less where the task allows it.Measured workload leads to Reuse work. Measured workload leads to Use less context. Measured workload leads to Smaller model. Check quality after each change; cheaper requests are not useful if they fail the task.Measured workloadFind where the cost comesfromReuse workCache or batch requestsUse less contextKeep the relevant materialSmaller modelUse it where quality holdsSpend less where the task allows it.Measured workload leads to Reuse work. Measured workload leads to Use less context. Measured workload leads to Smaller model. Check quality after each change; cheaper requests are not useful if they fail the task.Measured workloadFind where the cost comesfromReuse workCache or batch requestsUse less contextKeep the relevant materialSmaller modelUse it where quality holds
Read the connections in words
  • Measured workload → Reuse work: Cache or batch requests.
  • Measured workload → Use less context: Keep the relevant material.
  • Measured workload → Smaller model: Use it where quality holds.
Key idea

Check quality after each change; cheaper requests are not useful if they fail the task.

A focused business & team operations example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Cost optimization: see it in practice.

Reducing resource use or latency while preserving an explicitly defined level of quality.

What you’ll walk through

Follow a proposed saving through a quality and validity check. Inspect whether reuse, a smaller model, or fewer calls reduces expense without changing the result people rely on.

The task in this version

Reuse a policy answer only when valid for this user.

What you’ll learn to check

Cache-hit/miss cases, scope/version keys, labeled sample costs, latency distribution, and a quality floor.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Business & team operationsAn authored case with its own evidence, changed condition, and decision.
The task in this example

Reuse a policy answer only when valid for this user.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Cache key: policy v3 and employee role. New caller: contractor with different access.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

A cheaper response is not a saving if it creates more correction work. Cached content may be invalid for another user, source version, or time.

1 / 6

Apply this to your project

Describe your task to your own model and use Cost optimization as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Cost optimization is choosing, deliberately, which of several ways to spend fewer tokens and less time fits a system, not applying all of them by default. Preserve the required outcome, automation, and human effort when comparing designs. A lower model bill is not a saving if it hands unwanted work back to the user. The decision worksheet classifies a candidate design; it does not establish the cheapest suitable solution. Possible levers include: a smaller model for questions that don’t need a large one (routing), less text in the prompt (context engineering), reusing a prefix instead of resending it (prompt caching), deferring work that doesn’t need an instant answer (batching), capping how long an answer or a loop is allowed to run, and caching a whole answer rather than only a prompt prefix.

This topic is not a level on the ladder; it applies at every level, the way ops does, and every lever below is one of the levers ops names in general. This page goes one level deeper into each one: what it actually saves, and what it risks.

This page is sourced, not measured: every figure below is a maker’s own published rate, read off their page on the date given, and no lever here has been pulled on this site’s own traffic and scored.

Practical guidance

Before questioning any line on an AI bill, check whether you are even paying for the right amount of tool for the job: the worksheet walks through exactly that question, and a smaller, cheaper setup that still gets the work done beats optimizing anything built on top of the wrong one. Answer it first.

Once that is settled, two more levers show up as discounts on a vendor’s own pricing page, and both are worth understanding before a bill surprises you. Caching charges more to store something reusable, then charges less every time it gets reused: on the Claude API as its prompt-caching page reads on September 19, 2026, a cache write costs “1.25 times the base input tokens price” for a five-minute cache and “2 times” for a one-hour one, while a cache read costs “0.1 times the base input tokens price”[1]. Those are Anthropic’s figures for Anthropic’s models on that date and say nothing about any other provider, but the shape holds generally: a tool that resends the same instructions or documents in full on every call, instead of reusing them, is paying (and likely charging you) full price for material that has not changed since the last call.

Batching trades a wait for a discount. OpenAI’s Batch API offers a “50% cost discount compared to synchronous APIs” for work that “completes within 24 hours (and often more quickly)” instead of right away[2]. That is a good trade for a nightly report nobody is watching load, and the wrong one for anything a person is waiting on right now.

The check that tells you whether either lever is actually saving anything: ask for the bill’s own breakdown, not just the total. Put it to the vendor in writing: “Show me, per call, which requests hit a cache and which were batched, against which were charged the full rate.” If they cannot produce that split, the discount may not be reaching you even if the feature exists on paper. And if your own usage is occasional rather than steady, the honest answer is that neither lever is worth chasing: caching and batching both pay off on volume, and a handful of questions a week will not reach the point where either matters.

Implementation details

examples/ops/ already has the cost estimator this page’s levers get measured against: it reads recorded trace files and a price table the caller supplies, and reports cost and latency per question and per level. Nothing here builds a second one; the levers above all show up as the same two numbers this estimator already reports; they just change what a trace, or a price table, looks like going in.

examples/ops/run.py · lines 66–71
def _cost(model_id: str, tokens_in: int, tokens_out: int, prices: PriceTable) -> float | None:
    price = prices.get(model_id)
    if price is None:
        return None
    per_in, per_out = price
    return round(tokens_in / 1000 * per_in + tokens_out / 1000 * per_out, 6)

The arithmetic is one multiply-and-add per model id: tokens in against the input price, tokens out against the output price, both per 1,000 tokens. Every lever above changes one of those four numbers rather than the formula: caching changes which price a call’s input tokens are billed at, routing changes which row of the price table applies at all, and shorter context changes tokens_in directly.

No maker’s price is written into this repository anywhere, on purpose: a real price changes without notice and differs enough between providers that hard-coding one would go stale silently and read as this site’s own claim about a real number rather than a caller’s. The caller supplies and dates its own table instead:

examples/ops/run.py · lines 59–63
def load_price_table(path: Path) -> PriceTable:
    """A price table file: `{"<model_id>": {"in_per_1k": ..., "out_per_1k": ...}, ...}`, USD per
    1,000 tokens. The caller states where these numbers came from; nothing here checks that."""
    raw = json.loads(path.read_text(encoding="utf-8"))
    return {model_id: (float(entry["in_per_1k"]), float(entry["out_per_1k"])) for model_id, entry in raw.items()}

Measuring cost per successful task, not per call, is what keeps a cheaper-looking lever honest: a smaller model that answers three times before getting a question right costs more than a larger one that answers once, even though every individual call was cheaper. That is exactly what an eval scored on outcomes, not on call count, is built to catch: the site’s own citation_hit_rate and score_overall fields, read alongside tokens_in and tokens_out from the same result file, are what a real version of this comparison would use.

The same two numbers read differently depending who is asking, and this estimator does not care which: a caller divides them however its own setting counts cost. A production line watching 4,000 units a day divides by units and by day, the way ops’s own cost strip does, so a fraction of a cent a call is worth chasing. An engineer validating five prototype boards before a design review divides by runs instead: five calls in an afternoon almost never read a cached prefix a second time inside its window, so caching rarely earns back its own write at that volume, and the number worth comparing against is the time the drafting saved, not the token price. A person writing one measurement’s uncertainty budget into a report is not dividing by anything: if a model drafts the surrounding prose at all, its cost is rounding error against the instrument’s own calibration, and the only cost that matters is being wrong.

tests/test_example_ops.py checks this arithmetic directly and is not mine to extend, but it is worth reading here: 1,000 input and 1,000 output tokens against a $1/$2-per-1,000-token made-up table comes to exactly $3.00, and a model id missing from the table reports usd: None, never $0.00: an unpriced call costs something; the estimator just cannot say how much, which is the same distinction that matters when a lever changes which model answered and the price table hasn’t caught up yet.

When you do not need this

Avoid adding optimization machinery until there is enough real, sustained traffic for the engineering cost of adding one to be smaller than what it saves: a cache write costs more than a plain call and only earns that back on a second read; a router is a second system to build and keep in sync with whatever it is choosing between.

Move up to a specific lever once its own condition is genuinely true: cache once the same content is read again inside its window, batch once real work can wait a day, route once a real share of questions are easy enough for a cheaper model to answer as well as the current one does.

At high enough volume the question stops being which lever and becomes whether to pay per token at all; running a model locally trades the bill for hardware and the upkeep of a service you now operate.

Failure modes

A cache write outnumbers its reads

How to notice it
The bill goes up after adding caching, not down, because the content being cached is rarely if ever read a second time inside its window, so every call pays the higher write price with none of the cheaper reads to offset it.
How to test for it
Track cache hit rate as its own number, separate from total spend; a lever that is supposed to save money but shows a falling hit rate is this failure, not a fluke.

Shorter context drops the passage a later question needs

How to notice it
Trimming context to save tokens removes a passage that looked unnecessary for the question it was trimmed against, but turns out to be exactly what a later, different question needed.
How to test for it
Run the same trimmed context against a held-out set of questions it was not tuned against, not only the ones used to decide what to cut.

A cheaper model answers wrong and nobody notices the extra cost of getting it right

How to notice it
Routing to a smaller model looks like a savings in the per-call numbers, but the smaller model needs a retry or a correction more often, so the true cost per successful task is higher than the per-call price suggested.
How to test for it
Score cost per successful task, not per call, the way this page's "measure by outcome, not by call count" point argues: a lever that wins on the wrong denominator is not actually a saving.

Batching a request that actually needed an instant answer

How to notice it
Something gets routed to a batch queue that a person was actually waiting on, so the published discount is real but comes with a wait of hours that nobody agreed to on their behalf.
How to test for it
Check whether anything currently batched has a person waiting on its specific result, not just whether the aggregate batch completion time looks acceptable.

An output length cap truncates a correct answer

How to notice it
A hard cap on output tokens set to save cost cuts off an answer mid-sentence or mid-list on the questions that genuinely needed the extra length, and the truncation reads as a wrong answer rather than an incomplete one.
How to test for it
Run the cap against the longest legitimate answers in a question set, not only the typical case, and check whether any of them get cut rather than finish short.

At each level

Each level has one lever that pays before any of the others. The list below is that lever, not a catalog; ops covers what each level costs to run.

  • Conventional software: nothing is billed per token, so no lever below applies. Worth listing because it is the cheapest answer the worksheet can reach.
  • Direct prompting: model choice, and nothing else. Nothing repeats yet, so there is nothing to cache; nothing is deferred, so there is nothing to batch.
  • Added context: what you put in the window, because you pay for it on every single call. Cutting it and caching the part that cannot be cut are the same lever seen from two sides: see context engineering.
  • Workflows: the step that runs on every question whether it needs to or not. Find it first, and only then ask whether it could be batched or skipped.
  • Tool use: not a lever but a budgeting rule. The model decides whether a tool call happens, so the same question costs one call or two, and an estimate that assumes the cheaper case is wrong for some fraction of traffic you do not control.
  • Agent loops: the ceiling, before anything else. A loop with no fixed length has no natural stopping cost, which is what a single agent’s max_steps exists for.
  • Teams of Agents: seat assignment. Every lever multiplies by the number of agents, so the largest single saving is which model sits in which seat.
  • Always-on agents: how often it wakes up. A system that runs whether or not there is anything to do spends most of its budget deciding there is nothing to do.

Practices

  • Settle the level before touching a lever. The worksheet decides whether the system is one level too high, which is worth more than every lever on this page put together.
  • Divide by successful tasks, never by calls. A lever that wins on the wrong denominator is not a saving, and the two numbers move in opposite directions often enough to matter.
  • Give each lever its own condition and check it on a schedule: a cache needs a hit rate, a batch queue needs nobody waiting, a router needs a real share of easy questions.
  • Keep prices out of the code. A dated table supplied by the caller goes stale visibly; a hard-coded number goes stale silently and reads as a claim about a real price.
  • Enforce the levers once, at the gateway, rather than reimplementing caching and budgets in every application that calls a model.

Run it

What to monitor

Cost per successful task, not per call, broken out by which lever touched a given request (cached, batched, routed to a smaller model) so a change to one lever's savings does not hide inside an aggregate that also includes requests it never touched.

Cost at volume

Every lever here trades a bit of engineering complexity for a real discount, per the maker figures quoted above; the crossover point where that trade is worth making is a function of real, sustained volume, not of how the levers look on paper.

How it fails in production

A lever's condition stops being true without anyone noticing (a cache's content stops repeating, a batch queue starts holding requests someone is actually waiting on) and the saving silently becomes a cost or a complaint instead.

What to log

Which lever, if any, touched each call; tokens in and out; cache hit or miss; whether a call was batched or synchronous; and which model actually answered, so cost per successful task can be reconstructed after the fact rather than only estimated in advance.

Try it

  1. Use it

    Pick a task you currently do with a single, most-capable chat model. Using the worksheet's own questions, decide whether a cheaper level or a smaller model would still pass: then actually try it once and compare.

  2. Build it

    Run python -m examples.ops --demo from the repo root and read the RAG row's tokens_in (1,850). Using the cache-write and cache-read multipliers quoted above, work out by hand how many reads inside a five-minute window it takes before caching that prefix has paid for its own write.

  3. Either lane

    Pick one technique's Cost and latency strip elsewhere on this site. Name the one lever from this page that would cut its cost the most, and the one condition (named in this page's own failure modes) that would have to hold for that lever to actually pay off.

How it connects

Before, after and instead of this

Optional: products, tools, and models

3 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Reduce repeated work

Cache stable material or use a smaller model for easy cases, then rerun the quality checks.

Out there

Named products, tools and models

Tools3
  • Cloudflare AI GatewayCloudflare · hosted gateway for model calls
  • HeliconeHelicone · tracing and cost tracking
  • LangfuseClickHouse · tracing and cost tracking

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Prompt caching · Anthropic (Claude Platform Docs) (accessed 09/19/2026)
  2. Batch API · OpenAI (API documentation) (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page