# AI gateways

_Topics at every level · sourced_

One entry point in front of several model providers, for keys, routing, limits, fallback and logs.


## Guided worked example · Engineering & technical work

Fictional scripted fixture, not a measured run. No model or external actions execute.

**Overview:** Follow a request through a shared access layer and a provider failure or fallback. Inspect which policies remain true when the destination or model changes.

**Assumptions:** Providers may differ in data handling, tool support, and output behavior. Successful routing does not guarantee equivalent results.

**Design choices:** Centralize policies that benefit from consistent enforcement. Permit fallback only to destinations compatible with the request's requirements.

**Request:** Handle a provider timeout without violating data policy.

**Starting evidence:** Primary unavailable. B is allowed for public data only. Request contains restricted customer data.

**Action and control:** Check routing, data policy, capability, and budget before fallback.

**Stage records (authored, not executed):**

### Input record

Primary unavailable. B is allowed for public data only. Request contains restricted customer data.

What changed: Establish the facts supplied for this version of the task.

### Design note

Centralize policies that benefit from consistent enforcement. Permit fallback only to destinations compatible with the request's requirements.

What changed: Choose an approach before treating a proposed result as accepted.

### Proposed work

Check routing, data policy, capability, and budget before fallback.

What changed: Turn the request and evidence into the next action or transformation.

### Result record · illustrative

No allowed fallback: return unavailable with the request identifier rather than leak data.

What changed: Inspect the result of the authored example; this is not an executed model run.

### Verification plan

Route decision, timeout, allowed fallback, budget refusal, request identifier, and resulting quality check.

If the result falls short:
When no compatible provider is available, return a useful failure or defer the work. Avoid changing privacy or capability assumptions merely to produce a response.

What changed: Separate what needs checking from what the illustration establishes.

### Adaptation handoff

Use a gateway when several applications or providers share policy needs. A single simple integration may not need an additional routing layer.

What changed: Decide which assumptions, tools, and controls should change for your own task.

**Sample result:** No allowed fallback: return unavailable with the request identifier rather than leak data.

**Change something — Request contains approved public data only:** B is permitted under this fixture policy. Record the route and check quality separately.

**Decision:** Should any available provider receive the failed request?

**Answer:** No; fallback must satisfy policy and capability constraints.

**Why:** Fallback can change behavior, privacy assumptions, and duplicate-request risk; compatibility is not guaranteed.

**Review criteria:** Route decision, timeout, allowed fallback, budget refusal, request identifier, and resulting quality check.

**Recovery:** When no compatible provider is available, return a useful failure or defer the work. Avoid changing privacy or capability assumptions merely to produce a response.

**Adapt it:** Use a gateway when several applications or providers share policy needs. A single simple integration may not need an additional routing layer.

An AI gateway is one entry point that every call to a model provider goes through instead of
calling each provider directly. It centralizes provider keys, routing and fallback between
providers, rate limits and budgets per caller, caching, a log of every call, and policy checks on
the way out and back. LiteLLM's own documentation describes its proxy mode as a "Self-hosted LLM
Gateway (Proxy) with virtual keys, cost tracking, and an admin UI," offering "Virtual keys with
per-key/team/user budgets" and "Centralized logging, guardrails, and caching"[2].

A gateway is not an agent framework: it decides nothing about what a model should do next and
holds no task state. It overlaps with an API aggregator without being the same thing: LiteLLM's
own docs describe a library that "gives you a single, unified interface to call 100+ LLMs" and a
separately named "Gateway (Proxy)" run on top of it, adding keys, budgets and policy[2].
Two costs come with the entry point either way: one component everything now depends on, and one
more party that sees your traffic. Neither is a reason not to use one; both are decisions to make
on purpose.

This topic is not a level on the ladder; it applies at every level, the way
[ops](/gradient_ascent/techniques/ops/) does. This page is sourced, not measured: what each
gateway does comes from its own documentation, and no gateway here has been put in front of real
traffic and scored, so no number below is one this site took.

## Practical guidance

This one is not yours, and no amount of looking will make it visible. A gateway is a piece of a
company's own plumbing: one entry point its code calls instead of calling each model provider
directly. There is no setting for you to change, nothing to type, and no product feature that
switches one on. Two things still land on you, though, and both are worth five minutes.

The first is who sees your text. A gateway of either kind sees every prompt sent and every
response returned, in full, because that is what routing and logging a call requires. A hosted one
is a third party seeing that traffic on top of whatever the model provider itself sees, under its
own logging and retention terms, not the model maker's. OpenRouter, a hosted service, calls itself
"The Unified Interface For Every Model" on its front page, and the feature it lists there for
outages is "Higher Availability", described as "Reliable AI models via our distributed
infrastructure. Fall back to other providers when one goes down"[1]. Cloudflare AI
Gateway is another, whose documented features include caching, rate limiting, logging and "model
fallbacks in case of an error"[3]. Those are the makers' words for what their products
do, not results this site has checked. So when you are buying an AI product for your team, put one
sentence in writing to the vendor: "Does our text pass through any service other than the model
provider you name, and whose retention terms cover it?" A vendor who cannot answer that in a
sentence has not thought about it. It is [safety, privacy and
governance](/gradient_ascent/techniques/safety/)'s question, asked of one more company on the route.

The second is what an outage looks like. A product that starts failing every request at once,
rather than getting slower or dropping one feature, is showing a single component going down, not
five providers failing together. Report it that way instead of spending an afternoon on your own
account settings.

If you came here because you are paying the bill, the page you want is
[ops](/gradient_ascent/techniques/ops/).

## Implementation details

The example is a small in-process gateway: one call per key routes to that key's primary
provider, falls back to its secondary on any error, and is refused outright once the key's token
budget is spent. Both "providers" are `Model`s from `examples.common.model`: here two
`StubModel`s standing in for what would really be an Ollama tag and a Claude model id behind the
same interface, so the gateway's own code never has to know which one actually answered:

`examples/ai_gateways/run.py` (lines 57-109)

```python
class Gateway:
    routes: dict[str, Route]
    budgets: dict[str, int]  # key -> tokens remaining
    redact: bool = True
    log: list[LogEntry] = field(default_factory=list)

    def complete(self, key: str, messages: list[Message], *, tracer: Tracer | None = None, **kwargs) -> Completion:
        """Route one call for `key`, falling back to the secondary provider on any error.

        Raises `BudgetExceeded` before calling anything if the key is spent. If the fallback
        fails too, that provider's exception propagates unchanged; see the module docstring.

        `tracer` is optional and unused by the demo CLI: when given, it records the refusal, the
        fallback (if one happened) and the call that answered as `code` steps -- routing and
        budget enforcement are fixed rules the gateway applies, never a choice a model made, so
        nothing here is ever `decided_by: "model"`.
        """
        if self.budgets.get(key, 0) <= 0:
            if tracer is not None:
                tracer.record(kind="code", decided_by="code", title=f"Refuse: {key} has no budget left", detail=key)
            raise BudgetExceeded(f"key {key!r} has no budget left")
        route = self.routes[key]
        fell_back = False
        try:
            completion = route.primary.complete(messages, **kwargs)
        except Exception:  # noqa: BLE001 - any provider failure triggers fallback, on purpose
            if tracer is not None:
                tracer.record(kind="code", decided_by="code", title=f"Fall back: {key}'s primary failed", detail=key)
            completion = route.fallback.complete(messages, **kwargs)
            fell_back = True
        self.budgets[key] -= completion.tokens_in + completion.tokens_out
        prompt = None if self.redact else "\n".join(content_text(m.content) for m in messages)
        self.log.append(
            LogEntry(
                key=key,
                model_id=completion.model_id,
                tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out,
                fell_back=fell_back,
                prompt=prompt,
            )
        )
        if tracer is not None:
            tracer.record(
                kind="model",
                decided_by="code",
                title=f"{key}: call answered" + (" via fallback" if fell_back else ""),
                detail=completion.text[:200],
                tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out,
                ms=completion.ms,
            )
        return completion
```

Budget is checked before the call and reconciled after it against the tokens the provider
actually reported, which is the right order: a key with nothing left is refused without a
provider ever being contacted. Be precise about what that buys, though, because it is easy to
oversell. The check asks whether anything is left, not whether enough is left, so one call of any
size may still start and finish below zero; the overshoot is bounded by one call, and the tests
pin it. And check, call, subtract is three steps, so two calls racing each other can both read
the same remaining budget and both go through: a real gateway reserves against the budget under
a lock or in a shared store, and this one runs in a single thread and says so.

Fallback is the mechanism that needs a warning rather than a caveat. `complete` falls back on any
exception from the primary, and an exception does not tell you whether the provider did the work
before it failed. For a plain completion that is harmless. For anything with a side effect (a
call that charges something, files something, or sends something), retrying on a different
provider can do it twice, so falling back on everything that raises is a choice to make per
route, not a default to leave on.

Redaction is a constructor flag, not an afterthought. `Gateway.redact` defaults to `True`, so a
call's prompt text never reaches `LogEntry.prompt`; a caller has to turn it off deliberately to
see it, the same shape of default OpenTelemetry uses for anything that might carry
[sensitive content](/gradient_ascent/techniques/observability/). The flag covers the gateway's
own log and nothing else. If both providers fail, the second one's exception propagates
unchanged, and a real provider client often quotes part of the request in that string: a second
way prompt text leaves a system whose logs are careful.

`examples/ai_gateways/README.md` (lines 12-12)

```text
python -m examples.ai_gateways --demo
```

`tests/test_example_ai_gateways.py` checks all three mechanisms and then attacks them: a failing
primary falls back and the log records `fell_back=True`; a key with an exhausted budget raises on
its next call while a different key's budget is untouched; a planted secret never appears in the
log's `repr` with redaction on, only when it is explicitly turned off; and three more tests pin
the limits above: the single-call overshoot, the two calls that both read the budget before
either subtracts, and the refused call that reaches no provider and names no prompt in its error.

What the example leaves out matters as much as what it includes. It has no policy check: a real
gateway's "guardrails" step, the same idea [guardrails](/gradient_ascent/techniques/guardrails/)
covers, would run here, on the request before it is sent or the answer before it is returned. It
also has no cache: a real gateway's cache is keyed on more than the prompt text alone, because
two different keys asking the identical question are not necessarily allowed to share an answer,
and a cache that ignores which key asked is a way for one tenant's data to reach another's
response.

## When you do not need this

Skip a gateway while a project calls one provider directly with one key and has no budget,
routing or fallback need of its own: a gateway adds a component that itself has to stay running,
and a single call to a single provider has nothing for a gateway to route between.

Move up once more than one provider is in play, once separate callers need separate budgets or
keys tracked centrally, or once the same policy or logging check needs to apply no matter which
provider actually answers. A gateway is also where the levers on
[cost optimization](/gradient_ascent/techniques/cost-optimization/) get enforced once for
everything rather than reimplemented in each application that calls a model.

## Failure modes

### The gateway becomes the single point of failure

- **How to notice it:** Every provider behind the gateway is healthy, but every request still fails, because the one thing in front of all of them is down.
- **How to test for it:** Take the gateway itself offline in a test environment and confirm the failure is visible and distinguishable from a provider outage in whatever you monitor, not lumped in with "the model is down."

### Fallback hides a real outage instead of surfacing it

- **How to notice it:** A primary provider is failing every request, but because the fallback quietly answers every time, nothing downstream notices until someone asks why costs or latency changed.
- **How to test for it:** Check whether a fallback event is logged and counted on its own, the way this example's LogEntry.fell_back is, not merged into a single "request succeeded" metric that looks identical either way.

### A cache serves one caller's answer to a different caller

- **How to notice it:** Two different keys ask a similar or identical question, and a cache keyed only on the prompt text returns one caller's cached answer to the other, which can leak content across tenants.
- **How to test for it:** Send the same prompt under two different keys and confirm the cache key includes which key asked, not the prompt text alone.

### A budget check runs after the call instead of before it

- **How to notice it:** A key goes over budget because the check that should have refused the call ran only after the provider had already answered and been billed.
- **How to test for it:** Confirm a key with zero budget remaining is refused before any provider is called, the way this example's Gateway.complete checks first, not billed once more and then flagged.

### Fallback retries a request that should not run twice

- **How to notice it:** A primary provider fails after doing the work rather than before, the gateway cannot tell the two apart, and the retry on the secondary repeats a side effect: something charged, filed or sent twice for one request.
- **How to test for it:** List what each route can actually cause to happen, and confirm fallback is enabled only on the routes where repeating the request is harmless; for the rest, the gateway should surface the error rather than retry it somewhere else.

### A budget is a floor, not a cap, and one call goes under it

- **How to notice it:** A key with a little budget left starts a very large call, because the check asked whether anything remained rather than whether enough remained, and the key finishes the call below zero.
- **How to test for it:** Send one deliberately oversized request against a nearly-spent key and read the remaining budget afterwards. If it is negative, the overshoot is real and worth bounding by request size, not only by what is left.

### The log redacts nothing, or redacts what a reader actually needed

- **How to notice it:** A log built for debugging keeps full prompt and answer text by default, which is useful right up until the log itself becomes the thing someone has to secure and explain in an audit, or the opposite: redaction is on for everything, including the one field a real incident needed to see.
- **How to test for it:** Check what a log actually contains after a real request, not what a redaction flag is named; this example's own tests assert the secret string is absent with redact=True and present with it off, which is the same check to run against a real deployment.

## At each level

- [Direct prompting](/gradient_ascent/levels/1/): a gateway is one extra hop in front of the single
  call [chat](/gradient_ascent/techniques/chat/) makes; whether it earns its place here depends
  entirely on whether more than one provider or key is already in the picture.
- [Workflows](/gradient_ascent/levels/3/): your own code already decided which step calls which
  model, so the gateway's routing is redundant with that decision: what it adds here is one
  budget and one log across every step, not new fallback logic your code didn't already have.
- [Tool use](/gradient_ascent/levels/4/): a tool definition passes through the gateway unchanged
  in either direction, so nothing about [function
  calling](/gradient_ascent/techniques/function-calling/) is this layer's concern: only the token accounting around the call is.
- [Agent loops](/gradient_ascent/levels/5/): the number of calls a run makes is not fixed in advance,
  so a gateway's per-key budget is a real backstop here, not just bookkeeping. This is the same ceiling
  [a single agent](/gradient_ascent/techniques/single-agent/)'s own `max_tokens` cap is meant to
  provide, enforced one layer further out in case the cap inside the loop fails to hold.
- [Teams of Agents](/gradient_ascent/levels/6/): several agents can share one gateway key or
  each hold their own, which is exactly the choice that decides whether
  [a lead and its workers](/gradient_ascent/techniques/orchestrator-workers/) show up as one
  line in a cost report or several.
- [Always-on agents](/gradient_ascent/levels/7/): traffic is continuous rather than one burst
  per question, so a per-key budget stops being an occasional backstop and becomes the thing that
  decides how much a runaway loop can spend before anyone is awake to notice.

## Practices

- Check a key's budget before calling a provider, not after, and reserve against it rather than
  subtracting afterwards, so two simultaneous calls cannot both spend the last of it.
- Turn fallback on per route, not globally. A request that can be repeated safely is a different
  thing from one that charges, files or sends, and only the first is a fallback candidate.
- Key a cache on more than the prompt text: who asked matters as much as what was asked, or one
  caller's cached answer can reach another's request.
- Log every fallback as its own event, not folded into a plain success count, so a primary
  provider quietly failing every request is visible before someone asks why costs changed.
- Redact prompt and answer content in logs by default, and then check the paths the flag does not
  cover: a provider's raw error string is the usual one, and
  [observability](/gradient_ascent/techniques/observability/) lists the rest.
- Decide up front whether the gateway itself is allowed to be a single point of failure for your
  system, and if not, plan for what happens when it, not a provider behind it, is the thing that
  is down.

## Run it

**What to monitor.** Fallback rate per key, separately from plain error rate; budget remaining per key;
  and cache hit rate if caching is on, the same figures the ops track already asks for but split
  out per key instead of aggregated.

**Cost at volume.** A gateway centralizes routing and budget decisions that would otherwise be
  duplicated in every calling application; what it costs instead is running and securing one more
  service that everything else now depends on.

**How it fails in production.** A fallback masks a primary provider's outage until someone asks why answers
  changed, or a cache keyed only on the prompt text returns one caller's answer to a different
  caller.

**What to log.** Which key called, which provider actually answered, whether it fell back, tokens in
  and out, and whether a policy check ran and what it decided, with prompt and answer content
  redacted unless explicitly captured.

## Try it

1. **Use it.** Find a product you use that calls more than one AI provider (check its status page or documentation for the providers it names). Look for signs of a gateway in front of them: one combined status indicator, or a single outage that took down access to every provider at once.
2. **Build it.** Run python -m examples.ai_gateways --demo from the repo root and read the printed log. Then change team-a's budget in _demo_gateway (examples/ai_gateways/__main__.py) from 10 to 10000 and run it again: does the second call for team-a still get refused?
3. **Either lane.** Pick one of the failure modes above and write down, for a real system you use or built, which of its two mitigations (the check itself, or the log that would reveal the check failed) you actually have today.


## Sources

1. [OpenRouter](https://openrouter.ai/) — OpenRouter (accessed 2026-09-19)
2. [LiteLLM documentation](https://docs.litellm.ai/docs/) — BerriAI (accessed 2026-09-19)
3. [AI Gateway](https://developers.cloudflare.com/ai-gateway/) — Cloudflare (accessed 2026-09-19)


Last reviewed 2026-09-19.
