Primary sources
- OpenRouter · OpenRouter (accessed 09/19/2026)
- LiteLLM documentation · BerriAI (accessed 09/19/2026)
- AI Gateway · Cloudflare (accessed 09/19/2026)
One entry point in front of several model providers, for keys, routing, limits, fallback and logs.
Sourced
Concept at a glance
Routing and fallback follow your policy; limits and logs live at the shared entry point.
A focused engineering & technical work example. Additional perspectives appear where they provide a useful contrast.
A shared layer managing access, routing, policy, budgets, and telemetry across model services.
Follow a request through a shared access layer and a provider failure or fallback. Inspect which policies remain true when the destination or model changes.
Handle a provider timeout without violating data policy.
Route decision, timeout, allowed fallback, budget refusal, request identifier, and resulting quality check.
The setting makes the example concrete. Carry the underlying pattern into your own work; adapt the sources, tools, and level of oversight to your task.
Handle a provider timeout without violating data policy.
Authored case. Select any record below; nothing is sent to a model.What changed: Establish the facts supplied for this version of the task.
Providers may differ in data handling, tool support, and output behavior. Successful routing does not guarantee equivalent results.
Describe your task to your own model and use AI gateways as a reference. Ask whether it fits, which alternatives meet the same automation needs, and how you would implement and check the result.
An AI gateway is one entry point that every call to a model provider goes through instead of calling each provider directly. It centralizes provider keys, routing and fallback between providers, rate limits and budgets per caller, caching, a log of every call, and policy checks on the way out and back. LiteLLM’s own documentation describes its proxy mode as a “Self-hosted LLM Gateway (Proxy) with virtual keys, cost tracking, and an admin UI,” offering “Virtual keys with per-key/team/user budgets” and “Centralized logging, guardrails, and caching”[2].
A gateway is not an agent framework: it decides nothing about what a model should do next and holds no task state. It overlaps with an API aggregator without being the same thing: LiteLLM’s own docs describe a library that “gives you a single, unified interface to call 100+ LLMs” and a separately named “Gateway (Proxy)” run on top of it, adding keys, budgets and policy[2]. Two costs come with the entry point either way: one component everything now depends on, and one more party that sees your traffic. Neither is a reason not to use one; both are decisions to make on purpose.
This topic is not a level on the ladder; it applies at every level, the way ops does. This page is sourced, not measured: what each gateway does comes from its own documentation, and no gateway here has been put in front of real traffic and scored, so no number below is one this site took.
This one is not yours, and no amount of looking will make it visible. A gateway is a piece of a company’s own plumbing: one entry point its code calls instead of calling each model provider directly. There is no setting for you to change, nothing to type, and no product feature that switches one on. Two things still land on you, though, and both are worth five minutes.
The first is who sees your text. A gateway of either kind sees every prompt sent and every response returned, in full, because that is what routing and logging a call requires. A hosted one is a third party seeing that traffic on top of whatever the model provider itself sees, under its own logging and retention terms, not the model maker’s. OpenRouter, a hosted service, calls itself “The Unified Interface For Every Model” on its front page, and the feature it lists there for outages is “Higher Availability”, described as “Reliable AI models via our distributed infrastructure. Fall back to other providers when one goes down”[1]. Cloudflare AI Gateway is another, whose documented features include caching, rate limiting, logging and “model fallbacks in case of an error”[3]. Those are the makers’ words for what their products do, not results this site has checked. So when you are buying an AI product for your team, put one sentence in writing to the vendor: “Does our text pass through any service other than the model provider you name, and whose retention terms cover it?” A vendor who cannot answer that in a sentence has not thought about it. It is safety, privacy and governance’s question, asked of one more company on the route.
The second is what an outage looks like. A product that starts failing every request at once, rather than getting slower or dropping one feature, is showing a single component going down, not five providers failing together. Report it that way instead of spending an afternoon on your own account settings.
If you came here because you are paying the bill, the page you want is ops.
The example is a small in-process gateway: one call per key routes to that key’s primary
provider, falls back to its secondary on any error, and is refused outright once the key’s token
budget is spent. Both “providers” are Models from examples.common.model: here two
StubModels standing in for what would really be an Ollama tag and a Claude model id behind the
same interface, so the gateway’s own code never has to know which one actually answered:
class Gateway:
routes: dict[str, Route]
budgets: dict[str, int] # key -> tokens remaining
redact: bool = True
log: list[LogEntry] = field(default_factory=list)
def complete(self, key: str, messages: list[Message], *, tracer: Tracer | None = None, **kwargs) -> Completion:
"""Route one call for `key`, falling back to the secondary provider on any error.
Raises `BudgetExceeded` before calling anything if the key is spent. If the fallback
fails too, that provider's exception propagates unchanged; see the module docstring.
`tracer` is optional and unused by the demo CLI: when given, it records the refusal, the
fallback (if one happened) and the call that answered as `code` steps -- routing and
budget enforcement are fixed rules the gateway applies, never a choice a model made, so
nothing here is ever `decided_by: "model"`.
"""
if self.budgets.get(key, 0) <= 0:
if tracer is not None:
tracer.record(kind="code", decided_by="code", title=f"Refuse: {key} has no budget left", detail=key)
raise BudgetExceeded(f"key {key!r} has no budget left")
route = self.routes[key]
fell_back = False
try:
completion = route.primary.complete(messages, **kwargs)
except Exception: # noqa: BLE001 - any provider failure triggers fallback, on purpose
if tracer is not None:
tracer.record(kind="code", decided_by="code", title=f"Fall back: {key}'s primary failed", detail=key)
completion = route.fallback.complete(messages, **kwargs)
fell_back = True
self.budgets[key] -= completion.tokens_in + completion.tokens_out
prompt = None if self.redact else "\n".join(content_text(m.content) for m in messages)
self.log.append(
LogEntry(
key=key,
model_id=completion.model_id,
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
fell_back=fell_back,
prompt=prompt,
)
)
if tracer is not None:
tracer.record(
kind="model",
decided_by="code",
title=f"{key}: call answered" + (" via fallback" if fell_back else ""),
detail=completion.text[:200],
tokens_in=completion.tokens_in,
tokens_out=completion.tokens_out,
ms=completion.ms,
)
return completionBudget is checked before the call and reconciled after it against the tokens the provider actually reported, which is the right order: a key with nothing left is refused without a provider ever being contacted. Be precise about what that buys, though, because it is easy to oversell. The check asks whether anything is left, not whether enough is left, so one call of any size may still start and finish below zero; the overshoot is bounded by one call, and the tests pin it. And check, call, subtract is three steps, so two calls racing each other can both read the same remaining budget and both go through: a real gateway reserves against the budget under a lock or in a shared store, and this one runs in a single thread and says so.
Fallback is the mechanism that needs a warning rather than a caveat. complete falls back on any
exception from the primary, and an exception does not tell you whether the provider did the work
before it failed. For a plain completion that is harmless. For anything with a side effect (a
call that charges something, files something, or sends something), retrying on a different
provider can do it twice, so falling back on everything that raises is a choice to make per
route, not a default to leave on.
Redaction is a constructor flag, not an afterthought. Gateway.redact defaults to True, so a
call’s prompt text never reaches LogEntry.prompt; a caller has to turn it off deliberately to
see it, the same shape of default OpenTelemetry uses for anything that might carry
sensitive content. The flag covers the gateway’s
own log and nothing else. If both providers fail, the second one’s exception propagates
unchanged, and a real provider client often quotes part of the request in that string: a second
way prompt text leaves a system whose logs are careful.
python -m examples.ai_gateways --demotests/test_example_ai_gateways.py checks all three mechanisms and then attacks them: a failing
primary falls back and the log records fell_back=True; a key with an exhausted budget raises on
its next call while a different key’s budget is untouched; a planted secret never appears in the
log’s repr with redaction on, only when it is explicitly turned off; and three more tests pin
the limits above: the single-call overshoot, the two calls that both read the budget before
either subtracts, and the refused call that reaches no provider and names no prompt in its error.
What the example leaves out matters as much as what it includes. It has no policy check: a real gateway’s “guardrails” step, the same idea guardrails covers, would run here, on the request before it is sent or the answer before it is returned. It also has no cache: a real gateway’s cache is keyed on more than the prompt text alone, because two different keys asking the identical question are not necessarily allowed to share an answer, and a cache that ignores which key asked is a way for one tenant’s data to reach another’s response.
Skip a gateway while a project calls one provider directly with one key and has no budget, routing or fallback need of its own: a gateway adds a component that itself has to stay running, and a single call to a single provider has nothing for a gateway to route between.
Move up once more than one provider is in play, once separate callers need separate budgets or keys tracked centrally, or once the same policy or logging check needs to apply no matter which provider actually answers. A gateway is also where the levers on cost optimization get enforced once for everything rather than reimplemented in each application that calls a model.
max_tokens cap is meant to
provide, enforced one layer further out in case the cap inside the loop fails to hold.Fallback rate per key, separately from plain error rate; budget remaining per key; and cache hit rate if caching is on, the same figures the ops track already asks for but split out per key instead of aggregated.
A gateway centralizes routing and budget decisions that would otherwise be duplicated in every calling application; what it costs instead is running and securing one more service that everything else now depends on.
A fallback masks a primary provider's outage until someone asks why answers changed, or a cache keyed only on the prompt text returns one caller's answer to a different caller.
Which key called, which provider actually answered, whether it fell back, tokens in and out, and whether a policy check ran and what it decided, with prompt and answer content redacted unless explicitly captured.
Find a product you use that calls more than one AI provider (check its status page or documentation for the providers it names). Look for signs of a gateway in front of them: one combined status indicator, or a single outage that took down access to every provider at once.
Run python -m examples.ai_gateways --demo from the repo root and read the printed log. Then change team-a's budget in _demo_gateway (examples/ai_gateways/__main__.py) from 10 to 10000 and run it again: does the second call for team-a still get refused?
Pick one of the failure modes above and write down, for a real system you use or built, which of its two mitigations (the check itself, or the log that would reveal the check failed) you actually have today.
3 current examples · Products, tools, and models that demonstrate this concept. A selection, not a ranking.
Hosted gateway for model calls
Maker’s documentation Checked 09/18/2026One API for many models
Maker’s documentation Checked 09/18/2026One API for many models
Maker’s documentation Checked 09/18/2026Route requests through a shared entry point that handles credentials, budgets, logging, and permitted fallbacks.
Names listed 09/19/2026. 223 of 223 registry entries have been checked against the maker's own page; the registry marks the rest as unchecked.
Last reviewed 09/19/2026. Pages unreviewed for 90 days are flagged for another pass. Markdown version of this page