Topics at every level

Operations

Cost, speed, monitoring and running models on your own hardware.

Sourced

Concept at a glance

Run the system and respond to what it does.

Feedback loopConceptual illustration
Run the system and respond to what it does.Deploy a version leads to Serve requests. Serve requests leads to Measure. Measure leads to Adjust. Adjust leads to Serve requests as feedback. Operational decisions need measured quality, cost, latency, and reliability.Deploy a versionA defined configurationServe requestsRun the actual workloadMeasureQuality, latency, cost,errorsAdjustTune, roll back, or repairRun the system and respond to what it does.Deploy a version leads to Serve requests. Serve requests leads to Measure. Measure leads to Adjust. Adjust leads to Serve requests as feedback. Operational decisions need measured quality, cost, latency, and reliability.Deploy a versionA defined configurationServe requestsRun the actual workloadMeasureQuality, latency, cost,errorsAdjustTune, roll back, or repair

Ending or continuingMonitoring continues after release; each change can be rolled back.

Read the connections in words
  • Deploy a version → Serve requests: Run the actual workload.
  • Serve requests → Measure: Quality, latency, cost, errors.
  • Measure → Adjust: Tune, roll back, or repair.
  • Adjust → Serve requests: feedback informs another turn.
Key idea

Operational decisions need measured quality, cost, latency, and reliability.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

Operations: see it in practice.

Deploying and maintaining an AI system with monitoring, versioning, incident response, and recovery.

What you’ll walk through

Follow a change from a candidate release into monitored use and a possible rollback. Inspect how operational decisions depend on user outcomes, not just whether the service responds.

The task in this version

Roll out a support update gradually and recover from regression.

What you’ll learn to check

Release manifest, staged rollout, quality alert, rollback event, and incident record.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Roll out a support update gradually and recover from regression.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Release R2 serves a fictional 5% segment. Servers are healthy but reviewed unsupported answers increase.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Model, prompt, retrieval, and dependency changes can all alter behavior. Versions and rollback options must be known.

1 / 6
In this topic

4 pages under operations

Each one goes further into a part of this page than this page does.

Observability

Sourced

Recording what each run did, so a bad result can be traced to the step that caused it.

AI gateways

Sourced

One entry point in front of several model providers, for keys, routing, limits, fallback and logs.

Cost optimization

Sourced

Spending fewer tokens and less time for the same result: caching, batching, smaller models, shorter context.

Running models locally

Sourced

Running open-weight models on your own hardware: what fits, quantization, and what you give up.

Apply this to your project

Describe your task to your own model and use Operations as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

Every technique on this site has a cost strip because a working system has to keep working at volume, not just answer one question correctly. Ops is the set of levers for that: pay for less by caching what repeats and batching what can wait, keep latency down by choosing the smallest model that still passes, watch rate limits before they turn into a queue of failed requests, and know what a system is doing well enough to say why it costs what it costs. Running a model on your own hardware trades a per-call bill for hardware and upkeep you own; whether that trade pays depends on volume, on how sensitive the data is, and on whether a model small enough to run there does the job well enough.

None of this changes what a technique does; it changes what it costs to run enough times to matter. Every “Cost and latency” strip on this site measures one question on its own, which is exactly the part these levers do not settle.

This page is sourced, not measured: what each lever does comes from primary sources, and no production trace exists for this site, so what follows describes the levers and shows how to read what one already has.

Practical guidance

When a product you pay for by usage is slow, expensive, or flaky, three questions to the vendor find out why faster than guessing at settings on your own side. Send them together, in writing, rather than asking one at a time over a call.

Does it cache anything that repeats? Anthropic’s own documentation describes prompt caching as reusing “specific prefixes in your prompts” at a fraction of the price: a cache write costs “1.25 times” the ordinary rate for a five-minute cache, “2 times” for a one-hour one, and a cache read costs “0.1 times the base input tokens price,” with a footnote naming two current models whose reads are cheaper still[2]. A product built around a repeated system prompt or a shared document set that is not doing this is paying, and probably charging you, full price on every call.

Does it batch anything that does not need an instant answer? OpenAI’s Batch API documents “50% cost discount compared to synchronous APIs” for results that complete “within 24 hours (and often more quickly)” instead of immediately[3]. Ask whether any of what you pay for is eligible for that and simply is not using it.

When it slows down or goes quiet, ask what specifically was hit: “Which limit did we hit, requests, input tokens, or output tokens, and when does it clear?” Anthropic’s own API is rate-limited on requests, input tokens and output tokens per minute, with a “token bucket algorithm” so capacity “is continuously replenished up to your maximum limit, rather than being reset at fixed intervals”, and a 429 error naming which limit was exceeded[1]. A well-built product can answer that question specifically; one that just says the system is having issues has not looked, or is not telling you.

If the honest answer is that the product runs on hardware someone owns instead of paying per call, that is its own trade, not a strict upgrade: Ollama, one tool built for this, states “Nothing you run locally ever leaves your machine”[5], in exchange for buying and maintaining the machine it runs on, and a ceiling on how large a model that machine can actually hold.

Implementation details

The example is a small reporting tool: it reads trace files in the exact shape examples/common/trace.py writes, and a price table the caller supplies, and reports cost and latency per question and per level. No price is written into the code: a maker’s per-token price is specific to one model and changes without notice, so hard-coding one here would eventually be wrong and would read as this site’s own claim about a real price rather than the caller’s.

examples/ops/run.py · lines 74–87
def cost_for_trace(trace: dict, prices: PriceTable) -> QuestionCost:
    tokens_in = sum(step["tokens_in"] for step in trace["steps"])
    tokens_out = sum(step["tokens_out"] for step in trace["steps"])
    ms = sum(step["ms"] for step in trace["steps"])
    usd = _cost(trace["model_id"], tokens_in, tokens_out, prices)
    return QuestionCost(
        example=trace["example"],
        level=trace["level"],
        model_id=trace["model_id"],
        tokens_in=tokens_in,
        tokens_out=tokens_out,
        ms=ms,
        usd=usd,
    )

A trace’s model id might not be in the price table at all: a model retired since the table was built, a typo, a local model with no per-token price because nothing is billed for it. The estimator reports that as usd: None, not as free, and names every unpriced model id it found so the gap is visible instead of silently zeroed out:

examples/ops/run.py · lines 94–107
def summarize_by_level(costs: list[QuestionCost]) -> list[LevelSummary]:
    summaries = []
    for level in sorted({c.level for c in costs}):
        subset = [c for c in costs if c.level == level]
        priced = [c.usd for c in subset if c.usd is not None]
        summaries.append(
            LevelSummary(
                level=level,
                n=len(subset),
                mean_usd=_mean(priced) if priced else None,
                mean_ms=_mean([c.ms for c in subset]),
            )
        )
    return summaries

python -m examples.ops --demo needs no files: it writes two synthetic trace files with a real Tracer (one shaped like a level 1 chat call, one like the level 2 RAG run this site’s own rag page illustrates) and prices them against a small table that is explicitly made up for the demo, kept apart from examples.ops.run itself so nothing about the reusable code depends on it. A real report would point --traces at recorded trace.json files and --prices at a table built from a maker’s current, dated price page instead.

tests/test_example_ops.py uses only made-up prices (fake-small, fake-big), never a real vendor’s figures, and checks the arithmetic directly: 1,000 input and 1,000 output tokens against a $1/$2-per-1,000-token table comes to exactly $3.00, a trace with an unpriced model id reports None rather than $0, and summarize_by_level groups and averages correctly across several traces at the same level. That is the whole measurement for this example: it answers no question, so the site’s 60-question set does not score it, and docs/EVALS.md says so with the reason.

This site’s trace.json is its own small format, built for stepping through one recorded run on a page, not for production monitoring across every call a system makes. A real deployment generally reaches for a shared standard instead: OpenTelemetry describes itself as “vendor- and tool-agnostic”, an “observability framework and toolkit” for producing traces, metrics and logs that many different backends can read, and says that “The backend (storage) and the frontend (visualization) of telemetry data are intentionally left to other tools”[4]. These are the same tokens-in, tokens-out and wall-time fields this example reads out of a trace file, but emitted in a shape a tracing backend already knows how to store, query and alert on.

When you do not need this

Skip caching, batching and a routing layer before there is enough traffic for any of them to pay for themselves. A cache write costs more than a plain call and only earns that back once the same content is read again inside its window; a routing layer is a second system to build, test and keep in sync with whatever it is choosing between. Below some real volume, the simplest version (one model, called directly, priced as is) costs less in engineering time than any of these levers saves in tokens.

Skip building a tracing format of your own, too. This site’s own trace.json exists to step through one recorded run on a page; a real deployment should reach for a standard such as OpenTelemetry from the start rather than growing its own and migrating later.

Add a lever once traffic is real and sustained: caching once the same content is genuinely read again inside its cache window, batching once a real share of requests do not need an instant answer, and routing once questions arrive that a cheaper model would demonstrably have answered just as well as the expensive one.

Failure modes

A cache that quietly stops paying off

How to notice it
The bill creeps up over weeks with no single request failing or slowing down, because a cache's TTL started expiring between requests that used to land inside it.
How to test for it
Track cache hit rate as its own metric, not just total spend; a hit rate that drifts down with no code change is this failure, and a spend total alone will not show it until much later.

An unpriced model id goes unnoticed

How to notice it
A cost report understates the real bill because a model id (retired, mistyped, or newly added) has no entry in the price table and gets silently treated as free instead of flagged.
How to test for it
Check a report for a named list of unpriced model ids, the way this page's own example's summarize_by_level does, rather than trusting a total that a missing price can quietly shrink.

A retry storm looks like the product being slow

How to notice it
Requests line up and time out during a traffic spike, and from the outside it reads as the product being generally unreliable rather than a specific rate limit being hit.
How to test for it
Check whether a rate-limit error is logged with which limit it hit, separately from an ordinary timeout; if the two look the same in the logs, a slow period cannot be told apart from an unrelated one.

Local hardware sized for the wrong model

How to notice it
A self-hosted setup that ran a smaller model comfortably starts missing its latency target once a task needs a larger local model to pass the same eval, and nothing about the original sizing accounted for that trade.
How to test for it
Run the actual eval the product needs to pass on the smallest local model that could plausibly work before committing to hardware, not just on whichever model happened to be handy.

Routing sends the hard question to the cheap model

How to notice it
A router built to send easy requests to a cheaper model occasionally misjudges a hard one as easy, and the wrong-sized model answers it badly with no separate signal that routing, not the model itself, made the mistake.
How to test for it
Score accuracy broken out by which model actually answered, not only by question kind; a gap between the router's intended difficulty split and the model that actually handled a question is this failure.

At each level

  • Conventional software: there are no tokens and no per-call bill, the same zero row level 0’s own cost strip reports; cost is the compute you already pay for, and the ops questions below don’t really start yet.
  • Direct prompting: one question’s cost and latency is the floor every model-calling level’s cost strip on this site compares against, the baseline chat’s own page shows.
  • Added context: how much you put in the window can cost more than the question itself does; caching the part that repeats across questions is where the first real savings show up, which is the whole subject of context engineering.
  • Workflows: the same fixed step run thousands of times a day, with no need for an answer inside the next minute (one step of prompt chaining, run unconditionally on every question) is exactly the case batching was built for.
  • Tool use: whether a question needs a tool at all is the model’s own call, so cost swings between one call and two for the same kind of question (function calling’s own cost strip shows that split directly) and a cost estimate here has to budget for both, not just the typical case.
  • Agent loops: a loop with no fixed number of steps is the hardest thing on the ladder to put a ceiling on; a hard cap on steps or tokens matters here the way it matters to a single agent’s own max_steps and max_tokens, or to --budget-tokens on this site’s own eval runner.
  • Teams of Agents: several models working one task multiplies every per-call number by however many agents are in the team, the way a lead and its workers cost several times a single call for the same question. This is exactly where routing a smaller model to the easy seats and a stronger one to the hard seat earns its keep.
  • Always-on agents: a system that runs continuously has a cost that is a rate, not a number per question (an always-on assistant spends one call every tick whether or not anything gets proposed) so what’s worth alerting on shifts from “did this one thing cost too much” to “is this hour costing more than a normal hour does.”

Practices

  • Cache what repeats, batch what can wait, and route by difficulty: the three levers that move a bill, each with its own break-even. Cost optimization works through all three.
  • Measure cost and latency broken out by level or technique, not only as one aggregate number, so a change to one part of a system doesn’t hide in an average across all the others.
  • Log enough on every call to answer “why did this cost what it cost” by reading a record instead of reproducing the call, and alert on a rate rather than a running total for anything that runs continuously. Observability is what to record and in what shape.
  • Decide on purpose whether every call goes through one entry point of your own (AI gateways) and whether any of it runs on your own hardware (running models locally). Both are operations decisions before they are engineering ones.

Run it

What to monitor

Cost and latency per level or technique, cache hit rate, and how close traffic is running to a rate limit before it starts failing requests, not after.

Cost at volume

Caching and batching both trade a bit of complexity for a real discount on repeated or non-urgent work, per the maker figures cited above; routing by difficulty changes which model's price applies to how much of your traffic, which usually matters more than either discount.

How it fails in production

A cache's TTL expires between requests that used to land inside it, and the discount silently disappears with nothing failing outright: the bill goes up while every individual request still succeeds, which is why it needs its own metric, not just an error count.

What to log

Tokens in and out, wall time, model id, cache hit or miss, and which level or technique a call belongs to, on every call. These are the same fields the example's estimator reads back out of a trace file.

Try it

  1. Use it

    Find an AI product you pay for by usage and check its documentation for whether it caches repeated content or batches non-urgent work. If it doesn't say, that silence is itself an answer worth noting.

  2. Build it

    Run python -m examples.ops --demo from the repo root and read the per-level report it prints. Then edit DEMO_PRICES in examples/ops/__main__.py to double the RAG model's output price and run it again. Which level's mean cost changes, and by how much?

  3. Either lane

    Pick one technique's Cost and latency strip elsewhere on this site and estimate what running it 10,000 times a day would cost, using the illustrative numbers on that page. Then note which lever here (caching, batching, routing) would cut that number the most.

How it connects

Before, after and instead of this

Optional: products, tools, and models

13 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

Explore 7 more examples
  • MLX Tool or framework · Apple

    Training and inference on Apple hardware

    Checked 09/18/2026
  • Ollama Tool or framework · Ollama

    Runs models locally

    Checked 09/18/2026
  • OpenRouter Tool or framework · OpenRouter

    One API for many models

    Checked 09/18/2026
  • OpenTelemetry Tool or framework · open standard

    Tracing standard

    Checked 09/18/2026
  • Phoenix Tool or framework · Arize AI

    AI observability and evaluation

    Checked 09/19/2026
  • SGLang Tool or framework · SGLang community

    Model serving runtime

    Checked 09/19/2026
  • vLLM Tool or framework · open source

    Model server

    Checked 09/18/2026
In practice

Operate a support assistant

Track failures, response time, and spend, then investigate regressions and roll back changes when needed.

Out there

Named products, tools and models

Tools13
  • HeliconeHelicone · tracing and cost tracking
  • LangfuseClickHouse · tracing and cost tracking
  • LangSmithLangChain · eval and tracing platform
  • LiteLLMBerriAI · one API for many models
  • llama.cppopen source · runs models locally
  • LM StudioElement Labs · runs models locally
  • MLXApple · training and inference on Apple hardware
  • OllamaOllama · runs models locally
  • OpenRouterOpenRouter · one API for many models
  • OpenTelemetryopen standard · tracing standard
  • PhoenixArize AI · AI observability and evaluation
  • SGLangSGLang community · model serving runtime
  • vLLMopen source · model server

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. Rate limits · Anthropic (Claude Platform Docs) (accessed 09/19/2026)
  2. Prompt caching · Anthropic (Claude Platform Docs) (accessed 09/19/2026)
  3. Batch API · OpenAI (API documentation) (accessed 09/19/2026)
  4. What is OpenTelemetry? · OpenTelemetry (accessed 09/19/2026)
  5. Ollama · Ollama (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page