Top 5 AI Gateways for Air-Gapped AI and In-VPC LLM Deployments
Air-gapped AI keeps model traffic, prompts, and gateway control planes inside a network the organization controls. This guide compares Bifrost, LiteLLM, Kong AI Gateway, Envoy AI Gateway, and Gravitee for in-VPC and on-prem LLM deployments with no SaaS dependency.
TL;DR
- An air-gapped AI deployment needs a gateway whose data plane and control plane both run inside the network boundary, with no SaaS dependency at runtime.
- Bifrost runs as a self-hosted cluster in a VPC or on-prem, mirrors its container image into an internal registry, and keeps routing, guardrails, audit logs, and secrets inside the boundary.
- LiteLLM, Kong AI Gateway, Envoy AI Gateway, and Gravitee can all be self-hosted, but SSO, RBAC, admin audit logs, and secret managers are licensed tiers or not documented for every option.
- Hosted-only routers are a poor fit for regulated teams, because the request path and the configuration plane both leave the organization's network.
- The deciding questions are where the control plane runs, which features need outbound internet, and how the gateway stays highly available without it.
Air-gapped AI is the practice of running large language models and the infrastructure around them on networks with no automated connection to the public internet, and the AI gateway is the component most likely to break that rule. Bifrost, the open-source AI gateway written in Go and built by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability, including fully self-hosted, in-VPC, and air-gapped AI deployments. This guide compares five AI gateways for finance, healthcare, government, and defense teams that cannot accept a SaaS control plane or telemetry egress.
What Is Air-Gapped AI?
Air-gapped AI is an AI deployment in which models, gateways, data stores, and identity services run on a network that has no automated link to the internet. Data crosses the boundary only through manual, reviewed transfers. An in-VPC deployment is the cloud equivalent: everything runs inside a private network, with egress tightly restricted or removed.
The NIST glossary defines an air gap as an interface between two systems that are not physically connected, where any logical connection is not automated and data moves only under human control. For LLM workloads, that definition rules out more than public model APIs. It also rules out a gateway that pulls configuration from a vendor cloud, phones home with usage telemetry, or downloads a pricing catalog at startup.

Figure 1: In an air-gapped AI deployment every runtime dependency of the gateway sits inside the boundary; anything outside it is optional and policy-gated.
Figure 1 shows the minimum footprint. Applications call one gateway API, and the gateway reaches self-hosted models such as vLLM, a config and log store, an identity provider, and a secrets backend, all on the same side of the boundary. Our broader guide to air-gapped and on-prem AI gateways for regulated industries covers the compliance drivers behind this architecture.
Key Criteria for an AI Gateway in Air-Gapped and In-VPC Deployments
An AI gateway for air-gapped or in-VPC use must run its control plane locally, install without internet access, stay highly available without an external coordinator, and enforce identity, access, audit, and secret controls on infrastructure the organization owns. Routing features matter only after those deployment constraints are met.
A feature that needs outbound internet is a feature you cannot turn on, so deployment comes before routing. The criteria below reflect what security reviewers ask for in regulated AI governance programs.
| Criterion | What to verify | Why it matters offline |
|---|---|---|
| Control plane location | Configuration, dashboard, and policy engine run in your network | A SaaS control plane is a permanent egress path |
| Offline installation | Documented image mirroring to an internal registry | Clusters cannot pull from public registries |
| No runtime egress | Pricing data, telemetry, and guardrails work without internet | Hidden startup fetches fail or silently degrade |
| High availability | Multi-node clustering without a cloud coordinator | Gateway downtime stops every AI workload |
| Identity and RBAC | OIDC or SAML SSO to a self-hosted IdP, custom roles | Access must follow the corporate directory |
| Audit logs | Signed or tamper-evident admin trails, exportable to a SIEM | Auditors need who changed what, and when |
| Secrets management | Provider keys resolved from Vault or a cloud secret store | Plaintext keys in a database fail security review |
| Self-hosted models | Native support for vLLM, Ollama, and internal CAs | Offline deployments usually mean self-hosted inference |
Self-Hosted LLM Gateways Compared at a Glance
All five gateways in this list can run on self-hosted infrastructure, but they differ in where the control plane lives, how they cluster, and which security controls sit behind a license. The table records only what each vendor's documentation states; "Not published" means we found no documentation for it. The LLM gateway buyer's guide covers the wider evaluation.
| Gateway | Control plane | Documented air-gap install | High availability | SSO and RBAC | Admin audit logs | Secret managers |
|---|---|---|---|---|---|---|
| Bifrost | Fully self-hosted (in-VPC, on-prem) | Yes: mirror image to internal registry | Peer-to-peer cluster, gossip plus gRPC sync | OIDC, SCIM, custom roles | HMAC-signed, export to JSON, JSONL, Syslog | HashiCorp Vault, AWS, GCP |
| LiteLLM | Self-hosted | Offline cost-map flag; full guide not published | Stateless replicas with PostgreSQL and Redis | SSO free up to 5 users, then licensed | Enterprise license | Vault, AWS, Azure, Google, CyberArk (Enterprise) |
| Kong AI Gateway | Self-managed, or Konnect SaaS in hybrid mode | Not published | Database-backed node cluster behind a load balancer | Kong Manager OIDC and RBAC (licensed) | Admin API and database changes, RSA-signable | Vault, AWS, Azure, GCP, CyberArk (Enterprise) |
| Envoy AI Gateway | Self-hosted on Kubernetes | Not published | Kubernetes controller replicas | JWT, OIDC, mTLS via Envoy Gateway policies | Not published | Not published |
| Gravitee | Self-hosted; multi-environment needs Gravitee Cloud | Plugin install steps for air-gapped clusters | Not published for the LLM proxy | OIDC SSO, roles (Enterprise Edition) | Audit Trail (Enterprise Edition) | HashiCorp Vault, AWS secret plugins |
1. Bifrost
Bifrost, the self-hosted AI gateway, runs entirely inside the customer's network, with no SaaS control plane in the request or configuration path. The Bifrost Enterprise build adds clustering, SSO, RBAC, signed audit logs, guardrails, and secret management, all deployed in-VPC or on-prem.
Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.
Bifrost exposes 25+ providers and 10,000+ models through one OpenAI-compatible API, and adds 11 microseconds of overhead per request at 5,000 RPS with a 100% success rate in sustained gateway benchmarks. Inside an air gap, the relevant part of that catalog is the self-hosted side: Ollama, vLLM, and SGLang endpoints, plus custom providers with internal CA certificates for model servers behind a private PKI.
Deployment inside the boundary
- In-VPC deployment: In-VPC deployments run on GKE, EKS, or AKS inside your VPC, with network isolation and no external network dependencies for the gateway itself.
- On-prem and air-gapped installs: The on-premise deployment guide documents pulling the image on a connected host, saving it to a tar file, loading it on the air-gapped side, and pushing it to an internal registry.
- Kubernetes and Helm: The Bifrost Helm chart exposes
image.repositoryandimagePullSecrets, so the chart points at your internal registry instead of a public one.
High availability without a cloud coordinator
Bifrost clustering uses a peer-to-peer design in which every node is an equal participant. Membership and liveness run over memberlist gossip on port 10101 (TCP and UDP), while configuration, governance counters, routing rules, and RBAC state replicate over gRPC on port 10102. Peer discovery works through Kubernetes, Consul, etcd, DNS, UDP, or mDNS, so no external service is needed, and three or more nodes are recommended for fault tolerance.

Figure 2: Every Bifrost node is an equal peer, so losing one node removes capacity, not the control plane.
Identity, access, audit, and secrets
- SSO and provisioning: OIDC and SCIM user provisioning works with Okta and Entra, and with self-hosted identity providers such as Keycloak and Zitadel, which matters when the IdP must also live inside the air gap.
- RBAC and data access: Role-based access control ships Admin, Developer, and Viewer roles plus custom roles, and data access control limits which rows each operator can see.
- Audit logs: Bifrost audit logs record administrative activity, sign entries with an HMAC key, apply a retention window, and export as JSON, JSON Lines, or Syslog for an internal SIEM.
- Secrets: Secret management resolves provider keys and other credentials from HashiCorp Vault, AWS Secrets Manager, or GCP Secret Manager, so Bifrost never stores plaintext keys in its database.
Guardrails that work offline
Bifrost guardrails include Bifrost-managed profiles that need no external service. Secrets detection is Gitleaks-backed and has no external network dependency, and custom regex guardrails run in-process with a built-in PII detection template. Teams that need broader PII coverage can point the Presidio profile at a self-hosted analyzer, while cloud-hosted guardrail providers stay disabled inside the air gap.
2. LiteLLM
LiteLLM is an open-source, Python-based gateway that runs self-hosted on Kubernetes through Helm or on AWS and GCP through Terraform. Enterprise controls such as SSO beyond five users, SCIM, admin audit logs, RBAC, and secret managers require a license key on the same image.
LiteLLM documents two deployment modes: a monolithic image that serves LLM traffic, management APIs, and the UI, and a microservices layout with separate gateway, backend, and UI services. The proxy is stateless, but PostgreSQL is required for auth and spend tracking, and Redis is required once more than one instance runs. Both stores need to be provisioned inside the boundary before the gateway is useful.
Two details matter for air-gapped AI. By default, LiteLLM fetches its model cost map from GitHub at startup; setting LITELLM_LOCAL_MODEL_COST_MAP=True switches to the bundled backup map for offline operation. LiteLLM audit logs cover create, update, delete, and regenerate actions on keys, teams, users, and models under the Enterprise license. Teams comparing footprint can review the migration path from LiteLLM to Bifrost for self-hosted deployments.
Best for: Python-centric platform teams that want an open-source gateway in their own cloud and are prepared to license enterprise controls and operate PostgreSQL and Redis alongside it.
3. Kong AI Gateway
Kong AI Gateway is a set of AI plugins on Kong Gateway that supports both Konnect, Kong's SaaS control plane, and fully self-hosted deployments. For a disconnected deployment, only the self-managed traditional mode keeps configuration inside the network, because hybrid mode connects data planes to a Konnect control plane.
In traditional mode, Kong Gateway nodes share a database and form a cluster behind a load balancer, with each node caching configuration in memory. The AI Proxy plugin lists self-hosted targets including Ollama, vLLM, and Llama alongside hosted providers. Kong Gateway audit logs capture Admin API requests and database changes, can be signed with an RSA key, and are not available in Konnect; RBAC, Kong Manager OIDC, and vault backends such as HashiCorp Vault require an Enterprise license.
Kong's documentation does not publish an air-gapped installation path for the AI Gateway specifically. Our Kubernetes Helm deployment guide for Bifrost shows the equivalent self-managed path for a dedicated AI gateway.
Best for: Organizations that already operate self-managed Kong Gateway Enterprise and want LLM routing inside the same API platform and license.
4. Envoy AI Gateway
Envoy AI Gateway, now renamed Agent Router under the Agentic AI Foundation, is an open-source project built on Envoy Gateway that runs as a Kubernetes control plane with Envoy as the data plane. It requires Kubernetes 1.32 or later and Envoy Gateway 1.9.2 or later, and has no SaaS component.
The project's capabilities cover provider fallback, token-based quota policies, usage-based rate limiting, an MCP gateway, upstream authentication, and InferencePool routing for self-hosted inference endpoints. Access control comes from Envoy Gateway security policies: JWT validation, OIDC integration, mutual TLS, external authorization, and IP allowlists. Admin audit logs, a role-based admin console, and secret manager integrations are not published in the project documentation.
Envoy AI Gateway fits teams comfortable assembling governance from CRDs, policies, and their own logging pipeline. Teams that need packaged audit trails can compare it with the approaches in our audit log controls for LLM traffic guide.
Best for: Kubernetes-native platform teams that already run Envoy Gateway and want a vendor-neutral, open-source routing layer inside the cluster.
5. Gravitee
Gravitee is an API management platform whose Enterprise Edition includes an LLM proxy for OpenAI, Anthropic, Gemini, Bedrock, Vertex AI, and OpenAI-compatible providers. Gravitee documents a self-hosted architecture in which every control plane and data plane component runs on-prem or in a private cloud.
The self-hosted option installs through Docker, Kubernetes (EKS, AKS, GKE, and OpenShift), RPM, or ZIP packages, and Gravitee documents plugin deployment steps for air-gapped clusters. One constraint applies: a self-hosted control plane must connect to Gravitee Cloud to support multi-environment configuration. The LLM proxy handles text generation only, with semantic caching, token rate limiting, and a guard rails policy, while OIDC SSO and the Audit Trail are Enterprise Edition features.
Teams whose primary need is LLM traffic inside one VPC may find a dedicated gateway simpler, as covered in our comparison of open-source AI gateway platforms for in-VPC teams.
Best for: Enterprises standardizing on Gravitee for API and event management that want LLM and agent traffic under the same governance model.
Running an On-Premise LLM Gateway With No Internet Egress
Running an on-premise LLM gateway with no internet egress requires three things: a reviewed path for importing container images, local replacements for every file the gateway normally downloads, and a network policy that blocks anything not explicitly listed. Most failures in air-gapped AI rollouts come from the second item.

Figure 3: Nothing in an air-gapped AI cluster pulls from the internet; images and reference files cross the gap once, through a reviewed transfer.
Figure 3 shows the import flow used by the Bifrost on-premise guide: pull and save on a connected staging host, scan and transfer, then load, tag, and push into the internal registry referenced by the Helm values. The same discipline applies to reference data. The Bifrost gateway reads model pricing from the framework.pricing.pricing_url setting in the config.json schema, which can point at an internal file host instead of a public URL.
| Dependency | Default behavior to check | Air-gapped replacement |
|---|---|---|
| Container image | Pulled from a public or vendor registry | Mirror to an internal registry |
| Model pricing data | Fetched from a URL on a schedule | Host the file internally and set the URL |
| Model providers | Hosted APIs over the internet | vLLM, Ollama, or SGLang inside the boundary |
| Guardrail providers | Some are cloud APIs | In-process regex, secrets detection, self-hosted Presidio |
| Identity provider | Cloud IdP | Self-hosted Keycloak or Zitadel via OIDC |
| Metrics and traces | SaaS observability endpoints | Prometheus scrape of /metrics, internal OTel collector |
The Bifrost deployment requirements state that the gateway needs network access to every provider, database, MCP server, identity system, and observability endpoint configured for the deployment. Inside the boundary, that list becomes the allowlist. A Kubernetes NetworkPolicy that permits only those internal destinations, plus ports 10101 and 10102 between gateway pods, turns the configuration into an enforced boundary. Our Bifrost cluster mode guide covers sizing and failover for that cluster.
Choosing a Self-Hosted AI Gateway for Regulated Teams
Choosing a self-hosted AI gateway for a regulated team starts with where the control plane runs, then moves to identity, audit, and availability. A gateway that fails the first test is out of scope regardless of its feature list, which is why hosted-only routers rarely survive a security review for air-gapped AI.

Figure 4: The first filter is where the control plane runs; feature comparisons only matter for gateways that pass it.
Hybrid gateways with a cloud control plane suit teams whose policies allow configuration egress. Open-source gateways without built-in SSO, RBAC, or audit logs suit teams that will build those controls themselves. Teams that need all three, plus multi-node availability, need a self-hosted enterprise cluster.
The answer also varies by sector:
- Healthcare: PHI cannot reach third-party services without agreements in place, so healthcare and life sciences teams typically pair self-hosted models with in-process PII redaction, as described in our guide to PII redaction at the gateway layer.
- Financial services: Model-risk and audit teams expect signed, exportable admin trails and least-privilege roles mapped to the corporate directory.
- Government and defense: Fully disconnected networks require the image-mirroring path, a self-hosted IdP, and zero runtime egress, with risk documented against frameworks such as the NIST AI Risk Management Framework.
For a broader view of gateway controls in these environments, see the Bifrost governance resource page and our longer survey of on-prem AI gateway options.
Frequently Asked Questions
What is an air-gapped AI system?
An air-gapped AI system runs models, gateways, and supporting services on a network with no automated connection to the internet. Data enters or leaves only through manual, reviewed transfers. Regulated teams use air-gapped AI to keep prompts, outputs, and model weights inside a controlled boundary, which requires a self-hosted gateway, self-hosted models, and local identity and secrets services.
What does air-gapped mean?
Air-gapped means a system is isolated from other networks so that no automated data path exists between them. Any transfer happens manually under human control, for example by moving a scanned file on approved media. In cloud environments, teams often approximate an air gap with an isolated VPC that has no internet gateway and only private endpoints to approved services.
What is the purpose of an air gap?
The purpose of an air gap is to remove network paths that attackers or misconfigured software could use to exfiltrate data or receive instructions. For AI workloads, an air gap ensures that sensitive prompts, retrieved documents, and model outputs never reach external services, and that no vendor component can change behavior remotely.
Can an AI gateway run fully air-gapped with self-hosted LLMs?
Yes, if both its data plane and control plane run locally and it has no mandatory runtime downloads. The Bifrost AI gateway supports this with an image-mirroring install, routing to vLLM and Ollama, an internally hosted pricing file, and in-process guardrails. Gateways that depend on a SaaS control plane or startup fetches need those dependencies replaced first.
Is a hybrid AI gateway with a cloud control plane acceptable for regulated workloads?
A hybrid gateway can be acceptable when policy allows configuration and metadata to leave the network while request payloads stay local. Many finance, healthcare, and government teams do not allow that, because the control plane can change routing and policy remotely. For those teams, a fully self-hosted gateway deployed with in-VPC isolation is the safer default.
How does an in-VPC gateway support HIPAA-compliant LLM workloads?
An in-VPC gateway supports HIPAA-compliant LLM workloads by keeping PHI inside infrastructure the covered entity controls and enforcing safeguards at one point. That includes SSO and role-based access, signed audit logs of configuration changes, PII redaction before requests reach a model, and encrypted secrets. The gateway is one control within a HIPAA program, not a substitute for it.
Deploy Bifrost Inside Your Network
Air-gapped AI depends on a gateway that keeps routing, policy, identity, and audit inside the boundary without a SaaS control plane or hidden downloads. Bifrost runs as a clustered, self-hosted AI gateway in your VPC or data center, with signed audit logs, RBAC, guardrails, and secret management for regulated teams. To plan an in-VPC or air-gapped rollout, book a Bifrost demo with the Bifrost team.