# Cost optimization

_Topics at every level · sourced_

Spending fewer tokens and less time for the same result: caching, batching, smaller models, shorter context.


## Guided worked example · Business & team operations

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a proposed saving through a quality and validity check. Inspect whether reuse, a smaller model, or fewer calls reduces expense without changing the result people rely on.

**Assumptions:** A cheaper response is not a saving if it creates more correction work. Cached content may be invalid for another user, source version, or time.

**Design choices:** Measure the full task cost and compare alternatives on representative outcomes. Cache only when the key covers the conditions that make reuse valid.

**Request:** Reuse a policy answer only when valid for this user.

**Starting evidence:** Cache key: policy v3 and employee role. New caller: contractor with different access.

**Action and control:** Check version and access scope before declaring a cache hit.

**Stage records (authored, not executed):**

### Input record

Cache key: policy v3 and employee role. New caller: contractor with different access.

What changed: Establish the facts supplied for this version of the task.

### Design note

Measure the full task cost and compare alternatives on representative outcomes. Cache only when the key covers the conditions that make reuse valid.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Check version and access scope before declaring a cache hit.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Cache miss: retrieve the contractor's authorized policy. Savings cannot justify unauthorized content.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Cache-hit/miss cases, scope/version keys, labeled sample costs, latency distribution, and a quality floor.

If the result falls short:
When a saving introduces errors, narrow its scope or fall back to the reliable path. Reassess the assumptions rather than assuming every request needs the expensive route.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Apply this to repeated answers, extraction, or agent loops. Optimize the dominant cost in your workload and define an acceptable quality floor first.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Cache miss: retrieve the contractor's authorized policy. Savings cannot justify unauthorized content.

**Change something — Policy becomes v4 for the same role:** Bypass or invalidate v3. Cached wording is not proof of current correctness.

**Decision:** Should a cache hit ignore version changes?

**Answer:** No; verify freshness and scope.

**Why:** A cache can serve stale or unauthorized content; lower average cost can hide worse tail latency or errors.

**Review criteria:** Cache-hit/miss cases, scope/version keys, labeled sample costs, latency distribution, and a quality floor.

**Recovery:** When a saving introduces errors, narrow its scope or fall back to the reliable path. Reassess the assumptions rather than assuming every request needs the expensive route.

**Adapt it:** Apply this to repeated answers, extraction, or agent loops. Optimize the dominant cost in your workload and define an acceptable quality floor first.

Cost optimization is choosing, deliberately, which of several ways to spend fewer tokens and less
time fits a system, not applying all of them by default. Preserve the required outcome, automation,
and human effort when comparing designs. A lower model bill is not a saving if it hands unwanted
work back to the user. The [decision worksheet](/gradient_ascent/worksheet/) classifies a candidate
design; it does not establish the cheapest suitable solution. Possible levers include: a smaller model for
questions that don't need a large one ([routing](/gradient_ascent/techniques/routing/)), less
text in the prompt ([context engineering](/gradient_ascent/techniques/context-engineering/)),
reusing a prefix instead of resending it (prompt caching), deferring work that doesn't need an
instant answer (batching), capping how long an answer or a loop is allowed to run, and caching a
whole answer rather than only a prompt prefix.

This topic is not a level on the ladder; it applies at every level, the way
[ops](/gradient_ascent/techniques/ops/) does, and every lever below is one of the levers ops
names in general. This page goes one level deeper into each one: what it actually saves, and
what it risks.

This page is sourced, not measured: every figure below is a maker's own published rate, read off
their page on the date given, and no lever here has been pulled on this site's own traffic and
scored.

## Practical guidance

Before questioning any line on an AI bill, check whether you are even paying for the right amount
of tool for the job: the [worksheet](/gradient_ascent/worksheet/) walks through exactly that
question, and a smaller, cheaper setup that still gets the work done beats optimizing anything
built on top of the wrong one. Answer it first.

Once that is settled, two more levers show up as discounts on a vendor's own pricing page, and
both are worth understanding before a bill surprises you. Caching charges more to store something
reusable, then charges less every time it gets reused: on the Claude API as its prompt-caching
page reads on September 19, 2026, a cache write costs "1.25 times the base input tokens price"
for a five-minute cache and "2 times" for a one-hour one, while a cache read costs "0.1 times the
base input tokens price"[1]. Those are Anthropic's figures for Anthropic's models on
that date and say nothing about any other provider, but the shape holds generally: a tool that
resends the same instructions or documents in full on every call, instead of reusing them, is
paying (and likely charging you) full price for material that has not changed since the last
call.

Batching trades a wait for a discount. OpenAI's Batch API offers a "50% cost discount compared to
synchronous APIs" for work that "completes within 24 hours (and often more quickly)" instead of
right away[2]. That is a good trade for a nightly report nobody is watching load, and the
wrong one for anything a person is waiting on right now.

The check that tells you whether either lever is actually saving anything: ask for the bill's own
breakdown, not just the total. Put it to the vendor in writing: "Show me, per call, which requests
hit a cache and which were batched, against which were charged the full rate." If they cannot
produce that split, the discount may not be reaching you even if the feature exists on paper. And
if your own
usage is occasional rather than steady, the honest answer is that neither lever is worth chasing:
caching and batching both pay off on volume, and a handful of questions a week will not reach the
point where either matters.

## Implementation details

`examples/ops/` already has the cost estimator this page's levers get measured against: it reads
recorded trace files and a price table the caller supplies, and reports cost and latency per
question and per level. Nothing here builds a second one; the levers above all show up as the
same two numbers this estimator already reports; they just change what a trace, or a price table,
looks like going in.

`examples/ops/run.py` (lines 66-71)

```python
def _cost(model_id: str, tokens_in: int, tokens_out: int, prices: PriceTable) -> float | None:
    price = prices.get(model_id)
    if price is None:
        return None
    per_in, per_out = price
    return round(tokens_in / 1000 * per_in + tokens_out / 1000 * per_out, 6)
```

The arithmetic is one multiply-and-add per model id: tokens in against the input price, tokens
out against the output price, both per 1,000 tokens. Every lever above changes one of those four
numbers rather than the formula: caching changes which price a call's input tokens are billed
at, routing changes which row of the price table applies at all, and shorter context changes
`tokens_in` directly.

No maker's price is written into this repository anywhere, on purpose: a real price changes
without notice and differs enough between providers that hard-coding one would go stale silently
and read as this site's own claim about a real number rather than a caller's. The caller supplies
and dates its own table instead:

`examples/ops/run.py` (lines 59-63)

```python
def load_price_table(path: Path) -> PriceTable:
    """A price table file: `{"<model_id>": {"in_per_1k": ..., "out_per_1k": ...}, ...}`, USD per
    1,000 tokens. The caller states where these numbers came from; nothing here checks that."""
    raw = json.loads(path.read_text(encoding="utf-8"))
    return {model_id: (float(entry["in_per_1k"]), float(entry["out_per_1k"])) for model_id, entry in raw.items()}
```

Measuring cost per successful task, not per call, is what keeps a cheaper-looking lever honest: a
smaller model that answers three times before getting a question right costs more than a larger
one that answers once, even though every individual call was cheaper. That is exactly what an
[eval](/gradient_ascent/techniques/evals/) scored on outcomes, not on call count, is built to
catch: the site's own `citation_hit_rate` and `score_overall` fields, read alongside `tokens_in`
and `tokens_out` from the same result file, are what a real version of this comparison would use.

The same two numbers read differently depending who is asking, and this estimator does not care
which: a caller divides them however its own setting counts cost. A production line watching
4,000 units a day divides by units and by day, the way [ops](/gradient_ascent/techniques/ops/)'s
own cost strip does, so a fraction of a cent a call is worth chasing. An engineer validating five
prototype boards before a design review divides by runs instead: five calls in an afternoon
almost never read a cached prefix a second time inside its window, so caching rarely earns back
its own write at that volume, and the number worth comparing against is the time the drafting
saved, not the token price. A person writing one measurement's uncertainty budget into a report
is not dividing by anything: if a model drafts the surrounding prose at all, its cost is rounding
error against the instrument's own calibration, and the only cost that matters is being wrong.

`tests/test_example_ops.py` checks this arithmetic directly and is not mine to extend, but it is
worth reading here: 1,000 input and 1,000 output tokens against a $1/$2-per-1,000-token made-up
table comes to exactly $3.00, and a model id missing from the table reports `usd: None`, never
$0.00: an unpriced call costs something; the estimator just cannot say how much, which is the
same distinction that matters when a lever changes which model answered and the price table
hasn't caught up yet.

## When you do not need this

Avoid adding optimization machinery until there is enough real,
sustained traffic for the engineering cost of adding one to be smaller than what it saves: a
cache write costs more than a plain call and only earns that back on a second read; a router is a
second system to build and keep in sync with whatever it is choosing between.

Move up to a specific lever once its own condition is genuinely true: cache once the same content
is read again inside its window, batch once real work can wait a day, route once a real share of
questions are easy enough for a cheaper model to answer as well as the current one does.

At high enough volume the question stops being which lever and becomes whether to pay per token
at all; [running a model locally](/gradient_ascent/techniques/local-inference/) trades the bill
for hardware and the upkeep of a service you now operate.

## Failure modes

### A cache write outnumbers its reads

- **How to notice it:** The bill goes up after adding caching, not down, because the content being cached is rarely if ever read a second time inside its window, so every call pays the higher write price with none of the cheaper reads to offset it.
- **How to test for it:** Track cache hit rate as its own number, separate from total spend; a lever that is supposed to save money but shows a falling hit rate is this failure, not a fluke.

### Shorter context drops the passage a later question needs

- **How to notice it:** Trimming context to save tokens removes a passage that looked unnecessary for the question it was trimmed against, but turns out to be exactly what a later, different question needed.
- **How to test for it:** Run the same trimmed context against a held-out set of questions it was not tuned against, not only the ones used to decide what to cut.

### A cheaper model answers wrong and nobody notices the extra cost of getting it right

- **How to notice it:** Routing to a smaller model looks like a savings in the per-call numbers, but the smaller model needs a retry or a correction more often, so the true cost per successful task is higher than the per-call price suggested.
- **How to test for it:** Score cost per successful task, not per call, the way this page's "measure by outcome, not by call count" point argues: a lever that wins on the wrong denominator is not actually a saving.

### Batching a request that actually needed an instant answer

- **How to notice it:** Something gets routed to a batch queue that a person was actually waiting on, so the published discount is real but comes with a wait of hours that nobody agreed to on their behalf.
- **How to test for it:** Check whether anything currently batched has a person waiting on its specific result, not just whether the aggregate batch completion time looks acceptable.

### An output length cap truncates a correct answer

- **How to notice it:** A hard cap on output tokens set to save cost cuts off an answer mid-sentence or mid-list on the questions that genuinely needed the extra length, and the truncation reads as a wrong answer rather than an incomplete one.
- **How to test for it:** Run the cap against the longest legitimate answers in a question set, not only the typical case, and check whether any of them get cut rather than finish short.

## At each level

Each level has one lever that pays before any of the others. The list below is that lever, not a
catalog; [ops](/gradient_ascent/techniques/ops/) covers what each level costs to run.

- [Conventional software](/gradient_ascent/levels/0/): nothing is billed per token, so no lever below
  applies. Worth listing because it is the cheapest answer the worksheet can reach.
- [Direct prompting](/gradient_ascent/levels/1/): model choice, and nothing else. Nothing repeats yet,
  so there is nothing to cache; nothing is deferred, so there is nothing to batch.
- [Added context](/gradient_ascent/levels/2/): what you put in the window, because you pay for it on
  every single call. Cutting it and caching the part that cannot be cut are the same lever seen
  from two sides: see [context engineering](/gradient_ascent/techniques/context-engineering/).
- [Workflows](/gradient_ascent/levels/3/): the step that runs on every question whether it
  needs to or not. Find it first, and only then ask whether it could be batched or skipped.
- [Tool use](/gradient_ascent/levels/4/): not a lever but a budgeting rule. The model decides
  whether a tool call happens, so the same question costs one call or two, and an estimate that
  assumes the cheaper case is wrong for some fraction of traffic you do not control.
- [Agent loops](/gradient_ascent/levels/5/): the ceiling, before anything else. A loop with no fixed
  length has no natural stopping cost, which is what
  [a single agent](/gradient_ascent/techniques/single-agent/)'s `max_steps` exists for.
- [Teams of Agents](/gradient_ascent/levels/6/): seat assignment. Every lever multiplies by the
  number of agents, so the largest single saving is which model sits in which seat.
- [Always-on agents](/gradient_ascent/levels/7/): how often it wakes up. A system that runs
  whether or not there is anything to do spends most of its budget deciding there is nothing to
  do.

## Practices

- Settle the level before touching a lever. The
  [worksheet](/gradient_ascent/worksheet/) decides whether the system is one level too high,
  which is worth more than every lever on this page put together.
- Divide by successful tasks, never by calls. A lever that wins on the wrong denominator is not
  a saving, and the two numbers move in opposite directions often enough to matter.
- Give each lever its own condition and check it on a schedule: a cache needs a hit rate, a batch
  queue needs nobody waiting, a router needs a real share of easy questions.
- Keep prices out of the code. A dated table supplied by the caller goes stale visibly; a
  hard-coded number goes stale silently and reads as a claim about a real price.
- Enforce the levers once, at [the gateway](/gradient_ascent/techniques/ai-gateways/), rather
  than reimplementing caching and budgets in every application that calls a model.

## Run it

**What to monitor.** Cost per successful task, not per call, broken out by which lever touched a given
  request (cached, batched, routed to a smaller model) so a change to one lever's savings does
  not hide inside an aggregate that also includes requests it never touched.

**Cost at volume.** Every lever here trades a bit of engineering complexity for a real discount, per
  the maker figures quoted above; the crossover point where that trade is worth making is a
  function of real, sustained volume, not of how the levers look on paper.

**How it fails in production.** A lever's condition stops being true without anyone noticing (a cache's
  content stops repeating, a batch queue starts holding requests someone is actually waiting on)
  and the saving silently becomes a cost or a complaint instead.

**What to log.** Which lever, if any, touched each call; tokens in and out; cache hit or miss; whether
  a call was batched or synchronous; and which model actually answered, so cost per successful
  task can be reconstructed after the fact rather than only estimated in advance.

## Try it

1. **Use it.** Pick a task you currently do with a single, most-capable chat model. Using the worksheet's own questions, decide whether a cheaper level or a smaller model would still pass: then actually try it once and compare.
2. **Build it.** Run python -m examples.ops --demo from the repo root and read the RAG row's tokens_in (1,850). Using the cache-write and cache-read multipliers quoted above, work out by hand how many reads inside a five-minute window it takes before caching that prefix has paid for its own write.
3. **Either lane.** Pick one technique's Cost and latency strip elsewhere on this site. Name the one lever from this page that would cut its cost the most, and the one condition (named in this page's own failure modes) that would have to hold for that lever to actually pay off.


## Sources

1. [Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) — Anthropic (Claude Platform Docs) (accessed 2026-09-19)
2. [Batch API](https://developers.openai.com/api/docs/guides/batch) — OpenAI (API documentation) (accessed 2026-09-19)


Last reviewed 2026-09-19.
