LLM Proxy vs LLM Gateway: How to Choose an Open Source Option
An LLM proxy forwards requests to model providers through one API, while an LLM gateway also authenticates, budgets, routes, and logs them. This guide explains the difference and how to choose an open source option, with Bifrost as the reference gateway.
TL;DR
- An LLM proxy is a pass-through service that gives applications one API for several model providers and forwards each request with stored credentials.
- An LLM gateway adds per-consumer access control, budgets, rate limits, routing with fallbacks, caching, and request-level observability on top of that unified API.
- A proxy is enough for a single team prototyping against one or two providers; production traffic shared by several teams needs a gateway.
- An open source license matters for gateways because the gateway sees every prompt, response, and provider key, so teams need to self-host and audit it.
- Bifrost is an Apache 2.0, Go-based open source LLM gateway that adds 11 microseconds of overhead per request at 5,000 RPS.
An LLM proxy is a service that sits between applications and model providers, exposes one API, and forwards each request to the right provider with the right credentials. Bifrost, the open-source AI gateway written in Go by Maxim AI, starts from that same unified API and adds the controls a production team needs: virtual keys, budgets, failover, caching, and observability. This guide explains where an LLM proxy stops and an LLM gateway begins, and how to choose the right open source option for your stage.
What Is an LLM Proxy?
An LLM proxy is a pass-through layer that accepts requests in one format, usually the OpenAI API schema, translates them for the target provider, and forwards them. Its job is connectivity: one endpoint, one SDK, and centrally stored provider keys, so application code does not change when a team adds a second model provider.
Most LLM proxies provide three capabilities:
- Schema translation between the OpenAI-style API and provider-specific APIs such as Anthropic or Google Gemini
- Credential storage so provider API keys live in one service instead of in every application
- Basic forwarding with optional retries when a provider returns a transient error
That scope is useful but narrow. A proxy treats every caller the same, so it cannot tell which team sent a request, cap what that team spends, or decide that a request should go to a different provider because the first one is failing. The same distinction exists in tool traffic, where the difference between an MCP proxy and an MCP gateway follows the same line between forwarding and governing.
What Is an LLM Gateway?
An LLM gateway is a control layer for model traffic that identifies each caller, applies policy, routes requests across providers, and records what happened. It includes everything an LLM proxy does, then adds authentication, budgets, rate limits, failover, caching, guardrails, and observability, which is why the terms "LLM gateway" and "AI gateway" describe the same category.

As Figure 1 shows, the gateway lane has decision points the proxy lane does not. Each request carries an identity, usually a gateway-issued key, and that identity determines which models it may call, how much it may spend, and which fallback chain applies. For a full breakdown of the category, see this complete guide to LLM gateways for enterprise AI.
A gateway also differs from a general API gateway. API gateways meter requests, while LLM gateways meter tokens and cost, stream responses, and understand provider-specific errors such as context-length failures; the comparison of AI gateways and API gateways for LLM traffic covers that boundary in detail.
LLM Proxy vs LLM Gateway: Key Differences
The core difference between an LLM proxy and an LLM gateway is control. A proxy answers "where should this request go?" while a gateway also answers "who sent it, is it allowed, what will it cost, and what happens if the provider fails?" The table below compares the two on the capabilities that matter once traffic reaches production.
| Capability | LLM proxy | LLM gateway |
|---|---|---|
| Unified OpenAI-compatible API | Yes | Yes |
| Central provider key storage | Yes | Yes, with weighted load balancing across keys |
| Caller identity | Shared credentials | Per-consumer keys for apps, teams, and users |
| Budgets and rate limits | Usually none | Per key, team, and customer |
| Provider failover | Basic retries, if any | Retries, key rotation, and fallback chains |
| Caching | Rare | Exact-match and semantic caching |
| Observability | Access logs | Per-request logs, token and cost metrics, traces |
| Guardrails and policy | None | Input and output checks before and after the provider call |
| Tool traffic (MCP) | Not covered | Governed alongside model traffic |
Two rows explain most migrations from proxy to gateway. Shared credentials mean a single runaway job can consume a provider's rate limit for every application, a problem made concrete by provider-enforced rate limits that are set at the organization level and vary by model. Missing failover means a provider incident becomes an application incident, which is why teams that handle LLM rate limits and outages with an AI gateway rarely go back to a plain proxy.
Why an Open Source AI Gateway Matters
An open source AI gateway lets teams run the layer that sees every prompt, completion, and provider key inside their own infrastructure, inspect its code, and avoid lock-in to a hosted vendor. Because the gateway is in the path of all model traffic, its license and deployment model are security decisions, not only procurement ones.
Three properties separate a usable open source LLM gateway from a source-available one:
- An OSI-approved license. A license that meets the Open Source Definition permits commercial use, modification, and redistribution. Bifrost is released under the Apache License 2.0, which also includes an explicit patent grant.
- Full self-hosting. The gateway should run with no calls to a vendor control plane, so prompts and keys never leave your network. Bifrost runs from a single binary or container and supports in-VPC deployments for regulated environments.
- A complete open core. Routing, virtual keys, budgets, caching, and observability should be in the open source edition rather than behind a license key, so the free version is production-capable.
Teams that want a ranked view of the options can compare the leading open source LLM gateways side by side, or narrow to open source LLM gateways built for self-hosted deployments.
How to Choose an Open Source LLM Gateway
Choosing an open source LLM gateway comes down to five questions: how much latency it adds, which providers it supports, how it governs access and spend, how it handles provider failures, and how it deploys in your environment. Score each candidate against the same workload, because published numbers rarely match production traffic patterns.
| Criterion | What to check | Why it matters |
|---|---|---|
| Overhead | Added latency per request at your target RPS | The gateway sits in every request path |
| Provider coverage | Native support for the providers and APIs you use | Gaps force teams back to direct SDK calls |
| Governance | Per-consumer keys, budgets, rate limits, RBAC | Required once more than one team shares providers |
| Reliability | Retries, key rotation, fallback chains, health-aware routing | Determines whether a provider outage reaches users |
| Deployment | Self-hosted binary or container, Kubernetes, clustering | Decides whether the gateway fits your platform |
| Extensibility | Plugin model and observability exports | Avoids forking the gateway to add custom logic |
| Security | Guardrails, secrets detection, audit trail | Gateway is the natural enforcement point for AI policy |
Security deserves its own pass. The OWASP Top 10 for LLM Applications lists risks such as prompt injection and sensitive information disclosure, and a gateway is the one place where controls for those risks can apply to every application at once. The existing guide on how to choose the right open source LLM gateway walks through each criterion with evaluation steps.
How Bifrost Works as an Open Source LLM Gateway
The Bifrost AI gateway is an open source LLM gateway written in Go and licensed under Apache 2.0. It exposes 25+ providers and 10,000+ models through one OpenAI-compatible API and adds 11 microseconds of overhead per request at 5,000 requests per second in sustained benchmarks, so the governance layer does not cost meaningful latency.

Figure 2 maps to specific Bifrost features:
- Drop-in setup. Bifrost starts with
npx -y @maximhq/bifrostordocker run -p 8080:8080 maximhq/bifrost, includes a web UI for configuration, and works as a drop-in replacement for OpenAI, Anthropic, and other SDKs by changing the base URL. - Virtual keys. Virtual keys are the primary governance entity, giving each application or team its own key with model and provider permissions.
- Budgets and rate limits. Hierarchical budgets and rate limits apply at the virtual key, team, and customer levels.
- Caching. Semantic caching serves exact-match repeats from a hash lookup and similar requests through embedding search, so the provider is never called on a hit.
- Failover. Automatic retries and fallbacks retry transient errors, rotate keys on rate-limit and auth failures, and move to the next provider in the chain.
Beyond the request path, Bifrost covers the rest of the gateway checklist. Prometheus metrics and OpenTelemetry traces export provider errors, latency, tokens, and cost. Bifrost also works as an MCP gateway, connecting to external tool servers and exposing them to clients under the same governance as model traffic.
For enterprise requirements, Bifrost Enterprise adds guardrails, role-based access control, and signed audit logs of administrative activity.
Clustering provides high availability for the gateway itself. Custom logic can be added through Go or WASM plugins without forking the gateway.
When an LLM Proxy Is Enough
An LLM proxy is enough when one team runs a prototype or internal tool against one or two providers, has no shared budget to protect, and can tolerate downtime when a provider fails. In that setting, a unified API and central key storage solve the immediate problem, and gateway features would add configuration without adding value.

The decision in Figure 3 usually changes for one of three reasons:
- A second team arrives. Shared provider keys make spend and rate limits impossible to attribute, which per-consumer virtual keys solve.
- A provider incident reaches users. Without fallback chains, the application inherits every provider outage, a pattern covered in this guide to LLM routers and how model routing works.
- Security or compliance review starts. Reviewers ask who can call which model and where prompts are logged, which requires access control and request logs at the gateway.
Starting with a gateway that also runs as a simple proxy avoids a migration later. Bifrost runs with zero configuration as a unified API on day one, and teams turn on virtual keys, budgets, and fallbacks when they need them, which the Bifrost governance guide describes step by step.
Frequently Asked Questions
What is an LLM proxy?
An LLM proxy is a service that gives applications a single API for multiple model providers. It translates requests into each provider's format, attaches stored credentials, and forwards them. An LLM proxy handles connectivity but usually lacks per-caller identity, budgets, and failover, which is the main difference from an open source LLM gateway such as Bifrost.
What is the difference between API and proxy?
An API is the interface a service exposes, while a proxy is an intermediary that forwards requests to that interface on a client's behalf. For LLM traffic, the provider's API defines how to call a model, and an LLM proxy sits in front of several provider APIs so applications call one endpoint. An LLM gateway is a proxy that also enforces policy.
Is the LLM gateway open source?
Some LLM gateways are open source and others are hosted-only or source-available. Bifrost is fully open source under the Apache 2.0 license, with routing, virtual keys, budgets, caching, and observability in the open source edition. Check for an OSI-approved license and full self-hosting before treating any gateway as open source.
Which open-source AI gateway is the best?
The best open source AI gateway depends on latency budget, provider coverage, and governance needs. For enterprises running mission-critical AI workloads, Bifrost is the strongest choice because it adds 11 microseconds of overhead at 5,000 RPS, supports 25+ providers, and includes virtual keys, budgets, failover, and MCP governance under Apache 2.0.
What is an AI gateway?
An AI gateway is a control layer between applications and AI model providers that unifies their APIs and applies authentication, budgets, rate limits, routing, caching, guardrails, and observability to every request. "AI gateway" and "LLM gateway" describe the same category; some AI gateways, including Bifrost, also govern MCP tool traffic alongside model calls.
Can an LLM proxy be upgraded to a gateway later?
Yes, but the migration touches every application that holds proxy credentials, because gateway governance depends on per-consumer keys. Teams avoid that work by starting with a gateway that runs as a simple unified API first. Bifrost runs with zero configuration and lets teams add virtual keys, budgets, and fallbacks later, with each application changing only the key it sends.
Start with Bifrost
The choice between an LLM proxy and an LLM gateway is a choice about control: a proxy unifies APIs, while a gateway governs who uses them, what they cost, and what happens when a provider fails. Bifrost gives teams both in one open source LLM gateway, from a zero-config unified API on day one to virtual keys, budgets, failover, and MCP governance in production. Book a demo to see how Bifrost fits your AI infrastructure.