# Operations

_Topics at every level · sourced_

Cost, speed, monitoring and running models on your own hardware.


## Try this in a recipe
- [Resume a monitor without duplicating alerts](/gradient_ascent/recipes/nightly-monitor.md): Process a stock event, save a local outbox record, and prove that replaying the same event does not create another alert.

## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a change from a candidate release into monitored use and a possible rollback. Inspect how operational decisions depend on user outcomes, not just whether the service responds.

**Assumptions:** Model, prompt, retrieval, and dependency changes can all alter behavior. Versions and rollback options must be known.

**Design choices:** Roll out proportionately to risk, compare relevant quality and reliability signals, and retain a known working path. A small internal tool may need a simpler process.

**Request:** Roll out a support update gradually and recover from regression.

**Starting evidence:** Release R2 serves a fictional 5% segment. Servers are healthy but reviewed unsupported answers increase.

**Action and control:** Monitor quality alongside availability and apply the predefined rollback rule.

**Stage records (authored, not executed):**

### Input record

Release R2 serves a fictional 5% segment. Servers are healthy but reviewed unsupported answers increase.

What changed: Establish the facts supplied for this version of the task.

### Design note

Roll out proportionately to risk, compare relevant quality and reliability signals, and retain a known working path. A small internal tool may need a simpler process.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Monitor quality alongside availability and apply the predefined rollback rule.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

Hold expansion and simulate rollback. Preserve request/configuration records for investigation.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Release manifest, staged rollout, quality alert, rollback event, and incident record.

If the result falls short:
When outcomes regress, limit exposure, identify the changed component, and roll back or degrade gracefully. Retrying requests is not a cure for a systematic quality regression.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use this for any maintained AI feature. Choose release stages and response thresholds around your workload, users, and recovery cost.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** Hold expansion and simulate rollback. Preserve request/configuration records for investigation.

**Change something — Monitor only server errors:** The quality regression is invisible despite successful HTTP responses. Add outcome-based signals.

**Decision:** Does a healthy service imply correct answers?

**Answer:** No; health and quality differ.

**Why:** A quality regression can occur without a server error; plan rollback, capacity limits, and degraded operation.

**Review criteria:** Release manifest, staged rollout, quality alert, rollback event, and incident record.

**Recovery:** When outcomes regress, limit exposure, identify the changed component, and roll back or degrade gracefully. Retrying requests is not a cure for a systematic quality regression.

**Adapt it:** Use this for any maintained AI feature. Choose release stages and response thresholds around your workload, users, and recovery cost.

Every technique on this site has a cost strip because a working system has to keep working at
volume, not just answer one question correctly. Ops is the set of levers for that: pay for less
by caching what repeats and batching what can wait, keep latency down by choosing the smallest
model that still passes, watch rate limits before they turn into a queue of failed requests, and
know what a system is doing well enough to say why it costs what it costs. Running a model on
your own hardware trades a per-call bill for hardware and upkeep you own; whether that trade pays
depends on volume, on how sensitive the data is, and on whether a model small enough to run there
does the job well enough.

None of this changes what a technique does; it changes what it costs to run enough times to
matter. Every "Cost and latency" strip on this site measures one question on its own, which is
exactly the part these levers do not settle.

This page is sourced, not measured: what each lever does comes from primary sources, and no
production trace exists for this site, so what follows describes the levers and shows how to read
what one already has.

## Practical guidance

When a product you pay for by usage is slow, expensive, or flaky, three questions to the vendor
find out why faster than guessing at settings on your own side. Send them together, in writing,
rather than asking one at a time over a call.

Does it cache anything that repeats? Anthropic's own documentation describes prompt caching as
reusing "specific prefixes in your prompts" at a fraction of the price: a cache write costs "1.25
times" the ordinary rate for a five-minute cache, "2 times" for a one-hour one, and a cache read
costs "0.1 times the base input tokens price," with a footnote naming two current models whose
reads are cheaper still[2]. A product built around a repeated system prompt or a shared
document set that is not doing this is paying, and probably charging you, full price on every
call.

Does it batch anything that does not need an instant answer? OpenAI's Batch API documents "50%
cost discount compared to synchronous APIs" for results that complete "within 24 hours (and often
more quickly)" instead of immediately[3]. Ask whether any of what you pay for is
eligible for that and simply is not using it.

When it slows down or goes quiet, ask what specifically was hit: "Which limit did we hit, requests,
input tokens, or output tokens, and when does it clear?" Anthropic's own API is rate-limited on
requests, input tokens and output tokens per minute, with a "token bucket algorithm" so capacity
"is continuously replenished up to your maximum limit, rather than being reset at fixed
intervals", and a 429 error naming which limit was exceeded[1]. A well-built product can
answer that question specifically; one that just says the system is having issues has not looked,
or is not telling you.

If the honest answer is that the product runs on hardware someone owns instead of paying per call,
that is its own trade, not a strict upgrade: Ollama, one tool built for this, states "Nothing you
run locally ever leaves your machine"[5], in exchange for buying and maintaining the
machine it runs on, and a ceiling on how large a model that machine can actually hold.

## Implementation details

The example is a small reporting tool: it reads trace files in the exact shape
`examples/common/trace.py` writes, and a price table the caller supplies, and reports cost and
latency per question and per level. No price is written into the code: a maker's per-token price
is specific to one model and changes without notice, so hard-coding one here would eventually be
wrong and would read as this site's own claim about a real price rather than the caller's.

`examples/ops/run.py` (lines 74-87)

```python
def cost_for_trace(trace: dict, prices: PriceTable) -> QuestionCost:
    tokens_in = sum(step["tokens_in"] for step in trace["steps"])
    tokens_out = sum(step["tokens_out"] for step in trace["steps"])
    ms = sum(step["ms"] for step in trace["steps"])
    usd = _cost(trace["model_id"], tokens_in, tokens_out, prices)
    return QuestionCost(
        example=trace["example"],
        level=trace["level"],
        model_id=trace["model_id"],
        tokens_in=tokens_in,
        tokens_out=tokens_out,
        ms=ms,
        usd=usd,
    )
```

A trace's model id might not be in the price table at all: a model retired since the table was
built, a typo, a local model with no per-token price because nothing is billed for it. The
estimator reports that as `usd: None`, not as free, and names every unpriced model id it found so
the gap is visible instead of silently zeroed out:

`examples/ops/run.py` (lines 94-107)

```python
def summarize_by_level(costs: list[QuestionCost]) -> list[LevelSummary]:
    summaries = []
    for level in sorted({c.level for c in costs}):
        subset = [c for c in costs if c.level == level]
        priced = [c.usd for c in subset if c.usd is not None]
        summaries.append(
            LevelSummary(
                level=level,
                n=len(subset),
                mean_usd=_mean(priced) if priced else None,
                mean_ms=_mean([c.ms for c in subset]),
            )
        )
    return summaries
```

`python -m examples.ops --demo` needs no files: it writes two synthetic trace files with a real
`Tracer` (one shaped like a level 1 chat call, one like the level 2 RAG run this site's own `rag`
page illustrates) and prices them against a small table that is explicitly made up for the demo,
kept apart from `examples.ops.run` itself so nothing about the reusable code depends on it. A real
report would point `--traces` at recorded `trace.json` files and `--prices` at a table built from
a maker's current, dated price page instead.

`tests/test_example_ops.py` uses only made-up prices (`fake-small`, `fake-big`), never a real
vendor's figures, and checks the arithmetic directly: 1,000 input and 1,000 output tokens against
a $1/$2-per-1,000-token table comes to exactly $3.00, a trace with an unpriced model id reports
`None` rather than $0, and `summarize_by_level` groups and averages correctly across several
traces at the same level. That is the whole measurement for this example: it answers no question,
so the site's 60-question set does not score it, and `docs/EVALS.md` says so with the reason.

This site's `trace.json` is its own small format, built for stepping through one recorded run on
a page, not for production monitoring across every call a system makes. A real deployment
generally reaches for a shared standard instead: OpenTelemetry describes itself as "vendor- and
tool-agnostic", an "observability framework and toolkit" for producing traces, metrics and logs
that many different backends can read, and says that "The backend (storage) and the frontend
(visualization) of telemetry data are intentionally left to other tools"[4]. These are
the same tokens-in, tokens-out and wall-time fields this example reads out of a trace file, but
emitted in a shape a tracing backend already knows how to store, query and alert on.

## When you do not need this

Skip caching, batching and a routing layer before there is enough traffic for any of them to pay
for themselves. A cache write costs more than a plain call and only earns that back once the same
content is read again inside its window; a routing layer is a second system to build, test and
keep in sync with whatever it is choosing between. Below some real volume, the simplest version
(one model, called directly, priced as is) costs less in engineering time than any of these
levers saves in tokens.

Skip building a tracing format of your own, too. This site's own `trace.json` exists to step
through one recorded run on a page; a real deployment should reach for a standard such as
OpenTelemetry from the start rather than growing its own and migrating later.

Add a lever once traffic is real and sustained: caching once the same content is genuinely read
again inside its cache window, batching once a real share of requests do not need an instant
answer, and routing once questions arrive that a cheaper model would demonstrably have answered
just as well as the expensive one.

## Failure modes

### A cache that quietly stops paying off

- **How to notice it:** The bill creeps up over weeks with no single request failing or slowing down, because a cache's TTL started expiring between requests that used to land inside it.
- **How to test for it:** Track cache hit rate as its own metric, not just total spend; a hit rate that drifts down with no code change is this failure, and a spend total alone will not show it until much later.

### An unpriced model id goes unnoticed

- **How to notice it:** A cost report understates the real bill because a model id (retired, mistyped, or newly added) has no entry in the price table and gets silently treated as free instead of flagged.
- **How to test for it:** Check a report for a named list of unpriced model ids, the way this page's own example's summarize_by_level does, rather than trusting a total that a missing price can quietly shrink.

### A retry storm looks like the product being slow

- **How to notice it:** Requests line up and time out during a traffic spike, and from the outside it reads as the product being generally unreliable rather than a specific rate limit being hit.
- **How to test for it:** Check whether a rate-limit error is logged with which limit it hit, separately from an ordinary timeout; if the two look the same in the logs, a slow period cannot be told apart from an unrelated one.

### Local hardware sized for the wrong model

- **How to notice it:** A self-hosted setup that ran a smaller model comfortably starts missing its latency target once a task needs a larger local model to pass the same eval, and nothing about the original sizing accounted for that trade.
- **How to test for it:** Run the actual eval the product needs to pass on the smallest local model that could plausibly work before committing to hardware, not just on whichever model happened to be handy.

### Routing sends the hard question to the cheap model

- **How to notice it:** A router built to send easy requests to a cheaper model occasionally misjudges a hard one as easy, and the wrong-sized model answers it badly with no separate signal that routing, not the model itself, made the mistake.
- **How to test for it:** Score accuracy broken out by which model actually answered, not only by question kind; a gap between the router's intended difficulty split and the model that actually handled a question is this failure.

## At each level

- [Conventional software](/gradient_ascent/levels/0/): there are no tokens and no per-call bill, the same
  zero row [level 0](/gradient_ascent/techniques/order-zero/)'s own cost strip reports; cost is
  the compute you already pay for, and the ops questions below don't really start yet.
- [Direct prompting](/gradient_ascent/levels/1/): one question's cost and latency is the floor every
  model-calling level's cost strip on this site compares against, the baseline
  [chat](/gradient_ascent/techniques/chat/)'s own page shows.
- [Added context](/gradient_ascent/levels/2/): how much you put in the window can cost more than the
  question itself does; caching the part that repeats across questions is where the first real
  savings show up, which is the whole subject of
  [context engineering](/gradient_ascent/techniques/context-engineering/).
- [Workflows](/gradient_ascent/levels/3/): the same fixed step run thousands of times a day,
  with no need for an answer inside the next minute (one step of
  [prompt chaining](/gradient_ascent/techniques/prompt-chaining/), run unconditionally on every
  question) is exactly the case batching was built for.
- [Tool use](/gradient_ascent/levels/4/): whether a question needs a tool at all is the model's own
  call, so cost swings between one call and two for the same kind of question
  ([function calling](/gradient_ascent/techniques/function-calling/)'s own cost strip shows that
  split directly) and a cost estimate here has to budget for both, not just the typical case.
- [Agent loops](/gradient_ascent/levels/5/): a loop with no fixed number of steps is the hardest
  thing on the ladder to put a ceiling on; a hard cap on steps or tokens matters here the way it
  matters to [a single agent](/gradient_ascent/techniques/single-agent/)'s own `max_steps` and
  `max_tokens`, or to `--budget-tokens` on this site's own eval runner.
- [Teams of Agents](/gradient_ascent/levels/6/): several models working one task multiplies
  every per-call number by however many agents are in the team, the way
  [a lead and its workers](/gradient_ascent/techniques/orchestrator-workers/) cost several times
  a single call for the same question. This is exactly where routing a smaller model to the
  easy seats and a stronger one to the hard seat earns its keep.
- [Always-on agents](/gradient_ascent/levels/7/): a system that runs continuously has a cost
  that is a rate, not a number per question ([an
  always-on assistant](/gradient_ascent/techniques/agent-teammates/) spends one call every tick whether or not anything gets proposed)
  so what's worth alerting on shifts from "did this one thing cost too much" to "is this hour
  costing more than a normal hour does."

## Practices

- Cache what repeats, batch what can wait, and route by difficulty: the three levers that move a
  bill, each with its own break-even.
  [Cost optimization](/gradient_ascent/techniques/cost-optimization/) works through all three.
- Measure cost and latency broken out by level or technique, not only as one aggregate number, so
  a change to one part of a system doesn't hide in an average across all the others.
- Log enough on every call to answer "why did this cost what it cost" by reading a record instead
  of reproducing the call, and alert on a rate rather than a running total for anything that runs
  continuously. [Observability](/gradient_ascent/techniques/observability/) is what to record and
  in what shape.
- Decide on purpose whether every call goes through one entry point of your own
  ([AI gateways](/gradient_ascent/techniques/ai-gateways/)) and whether any of it runs on your
  own hardware ([running models locally](/gradient_ascent/techniques/local-inference/)). Both are
  operations decisions before they are engineering ones.

## Run it

**What to monitor.** Cost and latency per level or technique, cache hit rate, and how close traffic is
  running to a rate limit before it starts failing requests, not after.

**Cost at volume.** Caching and batching both trade a bit of complexity for a real discount on
  repeated or non-urgent work, per the maker figures cited above; routing by difficulty changes
  which model's price applies to how much of your traffic, which usually matters more than either
  discount.

**How it fails in production.** A cache's TTL expires between requests that used to land inside it, and the
  discount silently disappears with nothing failing outright: the bill goes up while every
  individual request still succeeds, which is why it needs its own metric, not just an error
  count.

**What to log.** Tokens in and out, wall time, model id, cache hit or miss, and which level or
  technique a call belongs to, on every call. These are the same fields the example's estimator reads back
  out of a trace file.

## Try it

1. **Use it.** Find an AI product you pay for by usage and check its documentation for whether it caches repeated content or batches non-urgent work. If it doesn't say, that silence is itself an answer worth noting.
2. **Build it.** Run python -m examples.ops --demo from the repo root and read the per-level report it prints. Then edit DEMO_PRICES in examples/ops/__main__.py to double the RAG model's output price and run it again. Which level's mean cost changes, and by how much?
3. **Either lane.** Pick one technique's Cost and latency strip elsewhere on this site and estimate what running it 10,000 times a day would cost, using the illustrative numbers on that page. Then note which lever here (caching, batching, routing) would cut that number the most.


## Sources

1. [Rate limits](https://platform.claude.com/docs/en/api/rate-limits) — Anthropic (Claude Platform Docs) (accessed 2026-09-19)
2. [Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) — Anthropic (Claude Platform Docs) (accessed 2026-09-19)
3. [Batch API](https://developers.openai.com/api/docs/guides/batch) — OpenAI (API documentation) (accessed 2026-09-19)
4. [What is OpenTelemetry?](https://opentelemetry.io/docs/what-is-opentelemetry/) — OpenTelemetry (accessed 2026-09-19)
5. [Ollama](https://ollama.com/) — Ollama (accessed 2026-09-19)


Last reviewed 2026-09-19.
