Topics at every level

AI gateways

One entry point in front of several model providers, for keys, routing, limits, fallback and logs.

Sourced

Concept at a glance

Put one controlled entry point before providers.

Decision pathsConceptual illustration
Put one controlled entry point before providers.Gateway leads to Provider A. Gateway leads to Provider B. Gateway leads to Local model. Routing and fallback follow your policy; limits and logs live at the shared entry point.GatewayKeys, routing, budgets, logsProvider APrimary routeProvider BAlternate or fallbackLocal modelA route for suitable tasksPut one controlled entry point before providers.Gateway leads to Provider A. Gateway leads to Provider B. Gateway leads to Local model. Routing and fallback follow your policy; limits and logs live at the shared entry point.GatewayKeys, routing, budgets, logsProvider APrimary routeProvider BAlternate or fallbackLocal modelA route for suitable tasks
Read the connections in words
  • Gateway → Provider A: Primary route.
  • Gateway → Provider B: Alternate or fallback.
  • Gateway → Local model: A route for suitable tasks.
Key idea

Routing and fallback follow your policy; limits and logs live at the shared entry point.

A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.

GUIDED WORKED EXAMPLE Fictional fixtures · scripted outputs · no live model or external actions

AI gateways: see it in practice.

A shared layer managing access, routing, policy, budgets, and telemetry across model services.

What you’ll walk through

Follow a request through a shared access layer and a provider failure or fallback. Inspect which policies remain true when the destination or model changes.

The task in this version

Handle a provider timeout without violating data policy.

What you’ll learn to check

Route decision, timeout, allowed fallback, budget refusal, request identifier, and resulting quality check.

The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.

Engineering & technical workAn authored case with its own evidence, changed condition, and decision.
The task in this example

Handle a provider timeout without violating data policy.

Authored case. Select any record below; nothing is sent to a model.
FOLLOW THE EXAMPLE1 / 6
Interpret this honestlySample evidence, not your actual data.No real messages, tools, training, or hardware operations run.The sequence illustrates the concept; it is not a recorded agent trace.
THE VISIBLE WORKStarting evidence
Input record
AUTHORED TEACHING RECORD · NOT A LIVE RUN
Primary unavailable. B is allowed for public data only. Request contains restricted customer data.

What changed: Establish the facts supplied for this version of the task.

WHY THIS MATTERS

What this case assumes

Providers may differ in data handling, tool support, and output behavior. Successful routing does not guarantee equivalent results.

1 / 6

Apply this to your project

Describe your task to your own model and use AI gateways as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.

Go deeper: practical guidance, failure modes, and implementation

An AI gateway is one entry point that every call to a model provider goes through instead of calling each provider directly. It centralizes provider keys, routing and fallback between providers, rate limits and budgets per caller, caching, a log of every call, and policy checks on the way out and back. LiteLLM’s own documentation describes its proxy mode as a “Self-hosted LLM Gateway (Proxy) with virtual keys, cost tracking, and an admin UI,” offering “Virtual keys with per-key/team/user budgets” and “Centralized logging, guardrails, and caching”[2].

A gateway is not an agent framework: it decides nothing about what a model should do next and holds no task state. It overlaps with an API aggregator without being the same thing: LiteLLM’s own docs describe a library that “gives you a single, unified interface to call 100+ LLMs” and a separately named “Gateway (Proxy)” run on top of it, adding keys, budgets and policy[2]. Two costs come with the entry point either way: one component everything now depends on, and one more party that sees your traffic. Neither is a reason not to use one; both are decisions to make on purpose.

This topic is not a level on the ladder; it applies at every level, the way ops does. This page is sourced, not measured: what each gateway does comes from its own documentation, and no gateway here has been put in front of real traffic and scored, so no number below is one this site took.

Practical guidance

This one is not yours, and no amount of looking will make it visible. A gateway is a piece of a company’s own plumbing: one entry point its code calls instead of calling each model provider directly. There is no setting for you to change, nothing to type, and no product feature that switches one on. Two things still land on you, though, and both are worth five minutes.

The first is who sees your text. A gateway of either kind sees every prompt sent and every response returned, in full, because that is what routing and logging a call requires. A hosted one is a third party seeing that traffic on top of whatever the model provider itself sees, under its own logging and retention terms, not the model maker’s. OpenRouter, a hosted service, calls itself “The Unified Interface For Every Model” on its front page, and the feature it lists there for outages is “Higher Availability”, described as “Reliable AI models via our distributed infrastructure. Fall back to other providers when one goes down”[1]. Cloudflare AI Gateway is another, whose documented features include caching, rate limiting, logging and “model fallbacks in case of an error”[3]. Those are the makers’ words for what their products do, not results this site has checked. So when you are buying an AI product for your team, put one sentence in writing to the vendor: “Does our text pass through any service other than the model provider you name, and whose retention terms cover it?” A vendor who cannot answer that in a sentence has not thought about it. It is safety, privacy and governance’s question, asked of one more company on the route.

The second is what an outage looks like. A product that starts failing every request at once, rather than getting slower or dropping one feature, is showing a single component going down, not five providers failing together. Report it that way instead of spending an afternoon on your own account settings.

If you came here because you are paying the bill, the page you want is ops.

Implementation details

The example is a small in-process gateway: one call per key routes to that key’s primary provider, falls back to its secondary on any error, and is refused outright once the key’s token budget is spent. Both “providers” are Models from examples.common.model: here two StubModels standing in for what would really be an Ollama tag and a Claude model id behind the same interface, so the gateway’s own code never has to know which one actually answered:

examples/ai_gateways/run.py · lines 57–109
class Gateway:
    routes: dict[str, Route]
    budgets: dict[str, int]  # key -> tokens remaining
    redact: bool = True
    log: list[LogEntry] = field(default_factory=list)

    def complete(self, key: str, messages: list[Message], *, tracer: Tracer | None = None, **kwargs) -> Completion:
        """Route one call for `key`, falling back to the secondary provider on any error.

        Raises `BudgetExceeded` before calling anything if the key is spent. If the fallback
        fails too, that provider's exception propagates unchanged; see the module docstring.

        `tracer` is optional and unused by the demo CLI: when given, it records the refusal, the
        fallback (if one happened) and the call that answered as `code` steps -- routing and
        budget enforcement are fixed rules the gateway applies, never a choice a model made, so
        nothing here is ever `decided_by: "model"`.
        """
        if self.budgets.get(key, 0) <= 0:
            if tracer is not None:
                tracer.record(kind="code", decided_by="code", title=f"Refuse: {key} has no budget left", detail=key)
            raise BudgetExceeded(f"key {key!r} has no budget left")
        route = self.routes[key]
        fell_back = False
        try:
            completion = route.primary.complete(messages, **kwargs)
        except Exception:  # noqa: BLE001 - any provider failure triggers fallback, on purpose
            if tracer is not None:
                tracer.record(kind="code", decided_by="code", title=f"Fall back: {key}'s primary failed", detail=key)
            completion = route.fallback.complete(messages, **kwargs)
            fell_back = True
        self.budgets[key] -= completion.tokens_in + completion.tokens_out
        prompt = None if self.redact else "\n".join(content_text(m.content) for m in messages)
        self.log.append(
            LogEntry(
                key=key,
                model_id=completion.model_id,
                tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out,
                fell_back=fell_back,
                prompt=prompt,
            )
        )
        if tracer is not None:
            tracer.record(
                kind="model",
                decided_by="code",
                title=f"{key}: call answered" + (" via fallback" if fell_back else ""),
                detail=completion.text[:200],
                tokens_in=completion.tokens_in,
                tokens_out=completion.tokens_out,
                ms=completion.ms,
            )
        return completion

Budget is checked before the call and reconciled after it against the tokens the provider actually reported, which is the right order: a key with nothing left is refused without a provider ever being contacted. Be precise about what that buys, though, because it is easy to oversell. The check asks whether anything is left, not whether enough is left, so one call of any size may still start and finish below zero; the overshoot is bounded by one call, and the tests pin it. And check, call, subtract is three steps, so two calls racing each other can both read the same remaining budget and both go through: a real gateway reserves against the budget under a lock or in a shared store, and this one runs in a single thread and says so.

Fallback is the mechanism that needs a warning rather than a caveat. complete falls back on any exception from the primary, and an exception does not tell you whether the provider did the work before it failed. For a plain completion that is harmless. For anything with a side effect (a call that charges something, files something, or sends something), retrying on a different provider can do it twice, so falling back on everything that raises is a choice to make per route, not a default to leave on.

Redaction is a constructor flag, not an afterthought. Gateway.redact defaults to True, so a call’s prompt text never reaches LogEntry.prompt; a caller has to turn it off deliberately to see it, the same shape of default OpenTelemetry uses for anything that might carry sensitive content. The flag covers the gateway’s own log and nothing else. If both providers fail, the second one’s exception propagates unchanged, and a real provider client often quotes part of the request in that string: a second way prompt text leaves a system whose logs are careful.

examples/ai_gateways/README.md · lines 12–12
python -m examples.ai_gateways --demo

tests/test_example_ai_gateways.py checks all three mechanisms and then attacks them: a failing primary falls back and the log records fell_back=True; a key with an exhausted budget raises on its next call while a different key’s budget is untouched; a planted secret never appears in the log’s repr with redaction on, only when it is explicitly turned off; and three more tests pin the limits above: the single-call overshoot, the two calls that both read the budget before either subtracts, and the refused call that reaches no provider and names no prompt in its error.

What the example leaves out matters as much as what it includes. It has no policy check: a real gateway’s “guardrails” step, the same idea guardrails covers, would run here, on the request before it is sent or the answer before it is returned. It also has no cache: a real gateway’s cache is keyed on more than the prompt text alone, because two different keys asking the identical question are not necessarily allowed to share an answer, and a cache that ignores which key asked is a way for one tenant’s data to reach another’s response.

When you do not need this

Skip a gateway while a project calls one provider directly with one key and has no budget, routing or fallback need of its own: a gateway adds a component that itself has to stay running, and a single call to a single provider has nothing for a gateway to route between.

Move up once more than one provider is in play, once separate callers need separate budgets or keys tracked centrally, or once the same policy or logging check needs to apply no matter which provider actually answers. A gateway is also where the levers on cost optimization get enforced once for everything rather than reimplemented in each application that calls a model.

Failure modes

The gateway becomes the single point of failure

How to notice it
Every provider behind the gateway is healthy, but every request still fails, because the one thing in front of all of them is down.
How to test for it
Take the gateway itself offline in a test environment and confirm the failure is visible and distinguishable from a provider outage in whatever you monitor, not lumped in with "the model is down."

Fallback hides a real outage instead of surfacing it

How to notice it
A primary provider is failing every request, but because the fallback quietly answers every time, nothing downstream notices until someone asks why costs or latency changed.
How to test for it
Check whether a fallback event is logged and counted on its own, the way this example's LogEntry.fell_back is, not merged into a single "request succeeded" metric that looks identical either way.

A cache serves one caller's answer to a different caller

How to notice it
Two different keys ask a similar or identical question, and a cache keyed only on the prompt text returns one caller's cached answer to the other, which can leak content across tenants.
How to test for it
Send the same prompt under two different keys and confirm the cache key includes which key asked, not the prompt text alone.

A budget check runs after the call instead of before it

How to notice it
A key goes over budget because the check that should have refused the call ran only after the provider had already answered and been billed.
How to test for it
Confirm a key with zero budget remaining is refused before any provider is called, the way this example's Gateway.complete checks first, not billed once more and then flagged.

Fallback retries a request that should not run twice

How to notice it
A primary provider fails after doing the work rather than before, the gateway cannot tell the two apart, and the retry on the secondary repeats a side effect: something charged, filed or sent twice for one request.
How to test for it
List what each route can actually cause to happen, and confirm fallback is enabled only on the routes where repeating the request is harmless; for the rest, the gateway should surface the error rather than retry it somewhere else.

A budget is a floor, not a cap, and one call goes under it

How to notice it
A key with a little budget left starts a very large call, because the check asked whether anything remained rather than whether enough remained, and the key finishes the call below zero.
How to test for it
Send one deliberately oversized request against a nearly-spent key and read the remaining budget afterwards. If it is negative, the overshoot is real and worth bounding by request size, not only by what is left.

The log redacts nothing, or redacts what a reader actually needed

How to notice it
A log built for debugging keeps full prompt and answer text by default, which is useful right up until the log itself becomes the thing someone has to secure and explain in an audit, or the opposite: redaction is on for everything, including the one field a real incident needed to see.
How to test for it
Check what a log actually contains after a real request, not what a redaction flag is named; this example's own tests assert the secret string is absent with redact=True and present with it off, which is the same check to run against a real deployment.

At each level

  • Direct prompting: a gateway is one extra hop in front of the single call chat makes; whether it earns its place here depends entirely on whether more than one provider or key is already in the picture.
  • Workflows: your own code already decided which step calls which model, so the gateway’s routing is redundant with that decision: what it adds here is one budget and one log across every step, not new fallback logic your code didn’t already have.
  • Tool use: a tool definition passes through the gateway unchanged in either direction, so nothing about function calling is this layer’s concern: only the token accounting around the call is.
  • Agent loops: the number of calls a run makes is not fixed in advance, so a gateway’s per-key budget is a real backstop here, not just bookkeeping. This is the same ceiling a single agent’s own max_tokens cap is meant to provide, enforced one layer further out in case the cap inside the loop fails to hold.
  • Teams of Agents: several agents can share one gateway key or each hold their own, which is exactly the choice that decides whether a lead and its workers show up as one line in a cost report or several.
  • Always-on agents: traffic is continuous rather than one burst per question, so a per-key budget stops being an occasional backstop and becomes the thing that decides how much a runaway loop can spend before anyone is awake to notice.

Practices

  • Check a key’s budget before calling a provider, not after, and reserve against it rather than subtracting afterwards, so two simultaneous calls cannot both spend the last of it.
  • Turn fallback on per route, not globally. A request that can be repeated safely is a different thing from one that charges, files or sends, and only the first is a fallback candidate.
  • Key a cache on more than the prompt text: who asked matters as much as what was asked, or one caller’s cached answer can reach another’s request.
  • Log every fallback as its own event, not folded into a plain success count, so a primary provider quietly failing every request is visible before someone asks why costs changed.
  • Redact prompt and answer content in logs by default, and then check the paths the flag does not cover: a provider’s raw error string is the usual one, and observability lists the rest.
  • Decide up front whether the gateway itself is allowed to be a single point of failure for your system, and if not, plan for what happens when it, not a provider behind it, is the thing that is down.

Run it

What to monitor

Fallback rate per key, separately from plain error rate; budget remaining per key; and cache hit rate if caching is on, the same figures the ops track already asks for but split out per key instead of aggregated.

Cost at volume

A gateway centralizes routing and budget decisions that would otherwise be duplicated in every calling application; what it costs instead is running and securing one more service that everything else now depends on.

How it fails in production

A fallback masks a primary provider's outage until someone asks why answers changed, or a cache keyed only on the prompt text returns one caller's answer to a different caller.

What to log

Which key called, which provider actually answered, whether it fell back, tokens in and out, and whether a policy check ran and what it decided, with prompt and answer content redacted unless explicitly captured.

Try it

  1. Use it

    Find a product you use that calls more than one AI provider (check its status page or documentation for the providers it names). Look for signs of a gateway in front of them: one combined status indicator, or a single outage that took down access to every provider at once.

  2. Build it

    Run python -m examples.ai_gateways --demo from the repo root and read the printed log. Then change team-a's budget in _demo_gateway (examples/ai_gateways/__main__.py) from 10 to 10000 and run it again: does the second call for team-a still get refused?

  3. Either lane

    Pick one of the failure modes above and write down, for a real system you use or built, which of its two mitigations (the check itself, or the log that would reveal the check failed) you actually have today.

How it connects

Before, after and instead of this

Read first

Optional: products, tools, and models

3 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.

In practice

Apply one policy across providers

Route requests through a shared entry point that handles credentials, budgets, logging, and permitted fallbacks.

Out there

Named products, tools and models

Tools3
  • Cloudflare AI GatewayCloudflare · hosted gateway for model calls
  • LiteLLMBerriAI · one API for many models
  • OpenRouterOpenRouter · one API for many models

Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.

Where this comes from

Primary sources

  1. OpenRouter · OpenRouter (accessed 09/19/2026)
  2. LiteLLM documentation · BerriAI (accessed 09/19/2026)
  3. AI Gateway · Cloudflare (accessed 09/19/2026)

Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page