Try Bifrost Enterprise free for 14 days. Request access

Best Self-Hosted AI Gateway in 2026

A self-hosted AI gateway routes, governs, and logs LLM traffic inside infrastructure you control, from one Docker container to an air-gapped cluster. This guide maps deployment models, data residency, secrets, and HA, then ranks Bifrost, LiteLLM, Kong AI Gateway, Apache APISIX, and Agent Router.

Best Self-Hosted AI Gateway in 2026

TL;DR

  • A self-hosted AI gateway runs inside infrastructure you control (a Docker host, a Kubernetes cluster, a private VPC, or an air-gapped data center), so prompts, logs, and provider keys stay inside your perimeter.
  • The deployment model sets most requirements: one Docker container with SQLite suits a pilot, while production needs Kubernetes, PostgreSQL, clustering, and a secret manager.
  • Bifrost adds 11 microseconds of overhead per request at 5,000 RPS, routes to 25+ providers and 10,000+ models, and documents Helm, in-VPC, on-premise, and air-gapped deployment paths.
  • LiteLLM, Kong AI Gateway, Apache APISIX, and Agent Router (formerly Envoy AI Gateway) can also be self-hosted, but each fits a narrower deployment profile.

Enterprise AI teams in 2026 are pulling LLM traffic back inside their own perimeter, and regulatory pressure, sensitive prompt data, and the unpredictable economics of agentic workflows have made the self-hosted AI gateway a foundational layer of modern AI infrastructure. Gartner's Predicts 2026: AI Sovereignty report frames this directly: achieving AI sovereignty requires decision-making authority across the entire AI stack, with on-premises and air-gapped deployments emerging as the architectural defaults for regulated workloads. A self-hosted AI gateway is what makes that practical, routing every LLM call through infrastructure the enterprise controls. This guide starts with the deployment decision (Docker, Kubernetes, in-VPC, or air-gapped) and then ranks the strongest self-hosted AI gateway options available today, led by Bifrost, the open-source AI gateway built by Maxim AI for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability.

What a Self-Hosted AI Gateway Actually Does

A self-hosted AI gateway is a deployable control plane that sits between applications and one or more LLM providers, running inside the enterprise's own infrastructure rather than as a managed SaaS. It centralizes authentication, routing, governance, observability, and cost control for every AI request without exposing prompt data to a third party. The same AI gateway architecture applies whether the upstream is a hosted API or a model you serve yourself, such as Ollama or vLLM on internal GPUs.

Apps and agents call Bifrost inside the perimeter, which uses PostgreSQL and a secret manager and routes to self-hosted or approved hosted models

Figure 1: The gateway is the only component that talks to model endpoints, so the perimeter boundary is enforced in one place.

The capabilities that separate a production-grade self-hosted AI gateway from a basic proxy include:

  • Unified API: a single OpenAI-compatible interface that abstracts away differences across providers
  • Provider routing and failover: automatic redistribution of traffic when an upstream provider degrades or returns errors
  • Governance primitives: virtual keys, per-team budgets, rate limits, and role-based access control
  • Semantic caching: response reuse based on semantic similarity rather than exact-match strings
  • MCP gateway capabilities: centralized tool access for agentic workflows built on the Model Context Protocol
  • Deployment flexibility: support for in-VPC, on-premises, air-gapped, and Kubernetes-native deployments
  • Observability hooks: native Prometheus metrics, OpenTelemetry traces, and audit-grade request logs

These are the dimensions that matter when an AI gateway will sit on the critical path of production inference traffic. The deployment dimension comes first, because it decides which of the others are even available.

Self-Hosted AI Deployment Models: Docker, Kubernetes, In-VPC, and Air-Gapped

A self-hosted AI gateway can run in four main deployment models: a single Docker container or VM, a Kubernetes cluster installed with Helm, an in-VPC deployment inside a private cloud account, and an on-premise or air-gapped deployment with no internet egress. Each step down that list adds isolation and operational work, so pick the lightest model that meets your compliance requirements.

Deployment model How Bifrost runs State and storage Fits
Docker or VM docker run or npx, one container SQLite in a persistent /app/data volume One team, pilots, internal tools
Kubernetes (Helm) Official Helm chart on any cluster, including EKS, GKE, and AKS PostgreSQL 16+ for config and logs; 3+ replicas with autoscaling Production platform teams
In-VPC Enterprise deployment inside your VPC on GCP, AWS, Azure, Cloudflare, or Vercel PostgreSQL inside the VPC; no external network dependencies Regulated teams on a public cloud
On-premise or air-gapped Enterprise image mirrored to an internal registry, run on Kubernetes or Docker PostgreSQL in your data center Defense, public sector, strict data residency

The gateway setup guide starts Bifrost with one command, and a single OSS instance handles roughly 3,000-5,000 RPS. The Helm deployment guide ships values files for SQLite-only, production HA, and external PostgreSQL, and the deployment overview adds ECS, Cloud Run, Fly.io, and Terraform.

Air-gapped AI deployments follow the on-premise deployment guide: the image moves across the boundary as a saved file, as Figure 2 shows. Teams comparing private-cloud options can also read the roundup of open-source AI gateway platforms for in-VPC teams.

Air-gapped AI image path: pull and save on a connected machine, transfer a tar file, push to an internal registry, run beside local models

Figure 2: The isolated network never reaches the public registry; only a saved image file crosses the boundary.

Data Residency and AI Sovereignty

Data residency in a self-hosted AI gateway means every artifact the gateway creates (prompts, responses, request logs, provider keys, and usage counters) is stored in a database and region you choose. The only traffic that leaves your network is the outbound call to the model providers you allow, and on-premise LLM deployments can remove that call entirely.

Bifrost keeps configuration, encrypted credentials, and request logs in a database you operate, and log exports offload request payloads to your own S3 or GCS bucket. An in-VPC deployment keeps all data processing inside your VPC with no external network dependencies. Under the EU General Data Protection Regulation, which restricts personal data transfers outside the EU, that boundary is what auditors check.

Multi-region deployments are practical because the database is not on the request path. Each Bifrost pod loads its configuration into memory at boot and writes logs asynchronously, so a cross-region deployment serves each region from local pods; strict residency rules also require checking where the PostgreSQL primary receiving those writes sits.

Secret Management for a Self-Hosted AI Gateway

Secret management decides where provider API keys and virtual key values live once the gateway runs on your infrastructure. A self-hosted AI gateway should resolve keys at runtime from an existing secret manager rather than store plaintext values in its database or container images.

Every Bifrost deployment can reference keys as env.<VAR> values populated from Kubernetes Secrets or container environment variables, and persisted credentials are encrypted with a stable encryption key. Bifrost Enterprise adds secret management with HashiCorp Vault, AWS Secrets Manager, and GCP Secret Manager: any secret field (provider keys, virtual key values, MCP auth headers) accepts a vault.<path> reference, and Bifrost stores only the reference.

The read_only mode resolves references and never writes to the vault, while read_and_write also pushes dashboard-saved values into it, which helps when migrating existing keys. Secret management requires a PostgreSQL config store and can use an IAM role (including IRSA on EKS) on AWS.

Clustering and High Availability

High availability for a self-hosted AI gateway means several gateway nodes share state, so one node can fail or be upgraded without dropping requests or losing budget counters, and the database stays off the request path so a database failure does not become an inference outage.

Bifrost clustering runs a peer-to-peer cluster with no single point of failure. Membership and liveness travel over a memberlist gossip layer, while configuration changes, virtual keys, routing rules, and governance counters replicate over a dedicated gRPC channel. Nodes find each other through six discovery methods (Kubernetes, Consul, etcd, DNS, UDP broadcast, and mDNS), and three or more nodes are recommended so the cluster tolerates one node failure.

A load balancer feeds three Bifrost cluster nodes that sync over gossip and gRPC, register with discovery, and read PostgreSQL at boot

Figure 3: Any node can fail or restart without losing budgets or routing rules, because every peer already holds the same state in memory.

The OSS multinode setup runs several nodes from one shared config.json and applies changes with a restart, while Enterprise clustering propagates UI and API changes to every node and supports zero-downtime rolling updates.

Key Criteria for Evaluating a Self-Hosted AI Gateway

Before selecting a self-hosted AI gateway, technical buyers should evaluate candidates against criteria that map to real production constraints. The criteria below act as a structured filter for the comparison that follows.

  • Performance overhead: latency added by the gateway at sustained, high-concurrency load, measured in microseconds, not theoretical RPS ceilings
  • Provider coverage: number of supported LLM providers and how quickly new models are integrated
  • MCP and agent support: native handling of Model Context Protocol, tool execution, and agentic routing
  • Governance depth: virtual keys, hierarchical budgets, RBAC, and per-consumer policy enforcement
  • Compliance posture: in-VPC and air-gapped deployment support, audit logs of administrative changes, and SOC 2, HIPAA, ISO 27001 readiness
  • Operational footprint: number of dependencies, configuration complexity, and the runtime needed to operate the gateway at scale
  • License and total cost: whether enterprise-grade features ship in the open source build or require a paid tier

For teams that want a more structured framework, the LLM Gateway Buyer's Guide maps each criterion to a capability matrix across the leading gateways. The companion ranking of open-source AI gateways for self-hosted LLM deployments applies the same filter with more weight on licensing.

The table compares the five gateways on deployment criteria; "Not published" means the project's documentation, read for this comparison, does not state it.

Gateway Runtime Kubernetes install Air-gapped path HA state sharing Secret manager integration
Bifrost Go; Docker image or npx Official Helm chart Documented image mirroring (Enterprise) P2P clustering, gossip + gRPC (Enterprise) Vault, AWS Secrets Manager, GCP Secret Manager (Enterprise)
LiteLLM Python Helm on EKS, GKE, or AKS Not published PostgreSQL + Redis once more than one instance runs Vault, AWS, Azure, Google, CyberArk (Enterprise)
Kong AI Gateway Kong Gateway plugins Kong Ingress Controller or Kubernetes Operator Not published Not published Not published
Apache APISIX APISIX plugins Helm charts Not published Not published Not published
Agent Router (Envoy AI Gateway) Envoy data plane Kubernetes-native CRDs Not published Not published Not published

Best Self-Hosted AI Gateways in 2026

The best self-hosted AI gateways in 2026 are Bifrost, LiteLLM, Kong AI Gateway, Apache APISIX, and Agent Router (formerly Envoy AI Gateway). Bifrost ranks first because it covers every deployment model above with AI-native governance and microsecond overhead; the other four fit narrower profiles.

1. Bifrost

The Bifrost AI gateway is a high-performance, open-source gateway that unifies access to 25+ LLM providers and 10,000+ models through a single OpenAI-compatible API. Bifrost is written in Go, distributed as a Docker image and an npx binary, and licensed under Apache 2.0. In sustained 5,000 requests-per-second benchmarks, Bifrost adds only 11 microseconds of overhead per request, with a 100% success rate. The gateway is fully self-hostable through npx, Docker, or Kubernetes, with zero-configuration startup.

Key capabilities:

  • Unified API across 25+ providers: OpenAI, Anthropic, AWS Bedrock, Google Vertex AI, Azure OpenAI, Mistral, Cohere, Groq, xAI, and more, accessible through a single endpoint
  • Automatic failover and load balancing: weighted distribution across API keys and providers, with zero-downtime fallback chains
  • Native MCP gateway: Bifrost acts as both an MCP client and server, with OAuth 2.0, tool filtering, and a Code Mode that cuts input tokens by up to 92.8% in large MCP deployments
  • Semantic caching: similarity-based response caching that reduces costs and latency for repeated query patterns
  • Governance via virtual keys: hierarchical budgets at the virtual key, team, and customer level, with rate limits and access control per virtual key
  • Enterprise deployment: in-VPC, on-premises, air-gapped, and Kubernetes-native deployments with clustering, RBAC, OIDC and SCIM provisioning (Okta, Entra, and more), and secret management through HashiCorp Vault, AWS Secrets Manager, or GCP Secret Manager
  • Observability: native Prometheus metrics, OpenTelemetry traces, and signed audit logs of administrative activity that support SOC 2, GDPR, HIPAA, and ISO 27001 reviews
  • Drop-in SDK compatibility: works as a drop-in replacement for OpenAI, Anthropic, Google GenAI, LiteLLM, and LangChain SDKs by changing only the base URL

Many teams start with one instance, such as a self-hosted AI gateway for Cursor routing to Claude or local Ollama models, and move to Kubernetes when the pilot becomes a platform service.

Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

2. LiteLLM

LiteLLM is a Python-based open-source LLM proxy that provides a unified OpenAI-compatible interface to more than 100 LLM providers. The core gateway is MIT-licensed, with an active contributor community and broad provider coverage. Teams self-host LiteLLM as a proxy server backed by PostgreSQL and Redis, typically deployed with Helm on EKS, GKE, or AKS alongside an external observability stack.

Key capabilities:

  • 100+ provider integrations through a unified OpenAI-format API
  • Virtual keys with personal, team, and team-member budgets
  • Latency-based, cost-based, and usage-based routing, plus Redis- and Qdrant-backed semantic caching and an MCP gateway
  • Logging integrations to S3, GCS, and external observability backends

Best for: Python-heavy engineering teams that prioritize maximum provider breadth and are comfortable operating PostgreSQL and Redis alongside the proxy.

Considerations: LiteLLM's Python architecture shapes the deployment. Production traffic runs on several proxy instances behind a load balancer, and Redis becomes a required dependency once more than one instance runs. Teams migrating from LiteLLM for performance or governance reasons can reference the LiteLLM alternative guide for a detailed capability comparison, or the enterprise comparison of Bifrost and LiteLLM.

3. Kong AI Gateway

Kong AI Gateway is an extension of Kong Gateway, the widely deployed API gateway built for hybrid and multi-cloud deployments. The AI Proxy plugin and related AI plugins add LLM-specific capabilities to Kong's existing API management infrastructure, which appeals to enterprises already running Kong for their broader API estate.

Key capabilities:

  • Multi-LLM routing through the AI Proxy plugin with OpenAI, Anthropic, Azure OpenAI, AWS Bedrock, Gemini, Vertex AI, Mistral, Ollama, and vLLM support
  • Semantic caching, prompt engineering, and request and response transformation plugins
  • MCP traffic governance with OAuth 2.1-based authorization, MCP traffic metrics, and AI MCP audit logs
  • Mature plugin ecosystem including OIDC, mTLS, rate limiting, and OpenTelemetry applicable to AI traffic

Best for: Teams already running Kong for non-AI API management that want to extend an existing gateway to handle LLM traffic without introducing a separate AI-native infrastructure layer.

Considerations: Several advanced AI capabilities, including the AI Semantic Cache plugin and the AI MCP OAuth2 plugin, are available only in the AI Gateway Enterprise offering. Self-hosting Kong for AI traffic means operating the full Kong Gateway data model, which pays off mainly for teams that already run it. The Kong alternatives for self-hosted AI gateways roundup covers teams without that investment.

4. Apache APISIX

Apache APISIX is a cloud-native API gateway hosted by the Apache Software Foundation. It has added a family of AI plugins (the ai-proxy series and related modules) that adapt common LLM providers and enable APISIX deployments to handle LLM traffic. As an Apache project, it benefits from strong open-source governance and a sizable contributor community.

Key capabilities:

  • The ai-proxy plugin family with adapters for OpenAI, Azure OpenAI, Anthropic, DeepSeek, Gemini, Vertex AI, Amazon Bedrock, and other OpenAI-compatible providers
  • Redis-backed exact and optional semantic caching through ai-cache, plus token rate limiting through ai-rate-limiting
  • Content review, access control, and rate limiting through the broader APISIX plugin ecosystem
  • Cloud-native architecture with strong Kubernetes support and Apache 2.0 licensing

Best for: Teams already operating APISIX for general API management that want to add LLM routing capabilities without standing up a separate AI gateway. It is also a reasonable choice where the AI gateway must coexist with a large estate of non-AI traffic on shared infrastructure.

Considerations: AI capabilities in APISIX are delivered as plugins rather than as a native AI-first architecture. APISIX has no dedicated MCP gateway (its mcp-bridge plugin, an SSE-to-stdio bridge, is deprecated), and AI-specific governance primitives such as hierarchical budgets are not published. Configuration complexity grows quickly for teams not already invested in the APISIX ecosystem.

5. Envoy AI Gateway (Now Agent Router)

Envoy AI Gateway, now named Agent Router and hosted by the Agentic AI Foundation, is built on Envoy Proxy, the data plane behind Istio and several other Kubernetes service meshes. It is an open-source project that extends Envoy Gateway with LLM-specific routing, token-based rate limiting, and provider fallback, and it reached its first generally available release (v1.0) in June 2026. The project's stated goal is resilient connectivity across LLM providers and self-hosted models with Kubernetes-native primitives.

Key capabilities:

  • Multi-provider routing with an OpenAI-compatible API surface
  • Token-based rate limiting and cost estimation
  • An MCP gateway that aggregates MCP servers behind one endpoint, with OAuth and tool-level authorization
  • InferencePool support for routing to self-hosted inference endpoints

Best for: Teams deeply invested in Kubernetes and the Envoy or Istio service mesh that want AI traffic managed through Kubernetes-native CRDs alongside their existing ingress and east-west traffic.

Considerations: Agent Router is new to general availability. Semantic caching and a virtual key budget hierarchy are not published in its documentation. The xDS configuration model has a steep learning curve for teams not already operating Envoy.

How to Pick a Self-Hosted AI Gateway

Pick a self-hosted AI gateway in two steps: first choose the deployment model your compliance and scale requirements demand (Figure 4), then choose the gateway that supports that model with the least extra infrastructure.

Decision flow asking about internet egress, cloud data residency, and high availability, each yes pointing to a self-hosted deployment model

Figure 4: Stricter isolation answers come first because they rule out every lighter option below them.

Your situation Deployment model Gateway fit
One team, pilot, or internal tool Docker or npx, single node with SQLite Bifrost OSS; LiteLLM with PostgreSQL
Production service on your own cluster Kubernetes with Helm, PostgreSQL, 3+ replicas Bifrost (OSS multinode or Enterprise clustering)
Regulated workload on a public cloud In-VPC on EKS, GKE, or AKS Bifrost Enterprise in-VPC deployment
No internet egress allowed On-premise or air-gapped, mirrored image Bifrost Enterprise with local models such as vLLM or Ollama

The right self-hosted AI gateway then depends on which constraints dominate the workload:

  • Performance and AI-native features matter most: Bifrost. Microsecond overhead, native MCP, semantic caching, and enterprise governance in a single Apache 2.0 gateway
  • Maximum provider breadth in Python-first teams: LiteLLM, with PostgreSQL and Redis to operate alongside the proxy
  • Existing Kong investment: Kong AI Gateway, accepting that semantic caching and MCP OAuth sit in the Enterprise offering
  • Existing APISIX investment: Apache APISIX, accepting the plugin-based AI feature set
  • Kubernetes-native, Envoy-first stack: Agent Router (formerly Envoy AI Gateway), accepting a newly GA project

For most enterprises building net-new AI infrastructure in 2026, the deciding factors converge on AI-native architecture, MCP readiness, and air-gapped deployment support. Bifrost is designed around all three, and the wider list of self-hosted LLM gateway options shows how the open-source field compares when licensing is the first filter. Teams replacing a hosted router can also review the self-hosted OpenRouter alternatives.

FAQ

What is an air-gapped AI deployment?

An air-gapped AI deployment runs with no connection to the public internet, so models, the gateway, and all data stay on an isolated network. Bifrost Enterprise reaches it as a saved image loaded into an internal registry, and requests route to local models such as vLLM or Ollama.

Is on-premise AI different from an in-VPC deployment?

On-premise AI runs on hardware in your own data center, while an in-VPC deployment runs in a private network inside a public cloud account such as AWS, GCP, or Azure. Both keep the gateway and its data under your control. In-VPC is usually faster to operate, and on-premise is required when data cannot touch a public cloud.

Can Bifrost run on Kubernetes?

Yes. Bifrost ships an official Helm chart for any compatible Kubernetes cluster, with guides for EKS, GKE, and AKS. The production example runs three replicas with PostgreSQL, autoscaling, and TLS ingress, and Bifrost Enterprise adds clustering with Kubernetes service discovery so replicas share governance state.

How does a self-hosted gateway protect provider API keys?

A self-hosted gateway keeps provider keys out of application code and resolves them at runtime. Bifrost reads keys from environment variables or Kubernetes Secrets in every edition and encrypts persisted credentials. Bifrost Enterprise also resolves vault.<path> references from HashiCorp Vault, AWS Secrets Manager, or GCP Secret Manager, so the database stores only references.

Do I need Bifrost Enterprise to self-host?

No. Open-source Bifrost self-hosts on Docker, npx, or Kubernetes under Apache 2.0, and one instance handles roughly 3,000-5,000 RPS. Bifrost Enterprise adds clustering with real-time state sync, in-VPC and air-gapped deployment support, secret manager integration, OIDC and SCIM, RBAC, and audit logs, and it offers a 14-day free trial.

Start Building with Bifrost

The category has matured to the point where running an AI gateway inside the enterprise's own perimeter is no longer a research project. Bifrost as a self-hosted AI gateway gives engineering teams microsecond-level performance, native MCP support, deep governance, and Apache 2.0 licensing in a gateway that runs end-to-end inside their infrastructure. For regulated industries, agentic workloads, and production AI traffic at scale, that combination is the new baseline.

To see how Bifrost fits as your self-hosted AI gateway of record, from a single Docker container to an air-gapped cluster, book a demo with the Bifrost team.