Try Bifrost Enterprise free for 14 days. Request access

Top 5 AI Gateways for Kubernetes and Helm Deployments in 2026

AI gateways route, govern, and observe LLM traffic, and on Kubernetes they run as one more Helm-managed workload. This guide compares Bifrost, Agent Router, agentgateway, Kong AI Gateway, and LiteLLM on charts, scaling, probes, HA, and telemetry.

Top 5 AI Gateways for Kubernetes and Helm Deployments in 2026

TL;DR

  • Bifrost ships one official Helm chart (bifrost/bifrost) that deploys a stateless Deployment on PostgreSQL or a StatefulSet on SQLite, with HPA, ServiceMonitor, and /health probes in its values file.
  • Agent Router (formerly Envoy AI Gateway) and agentgateway install as Kubernetes controllers through OCI Helm charts, with CRDs installed first and Gateway API resources as the configuration surface.
  • Kong AI Gateway runs its data plane in your cluster through the kong/ingress chart, with AI Gateway configuration managed through Konnect.
  • LiteLLM publishes a monolithic chart and a componentized chart, and its Helm path requires PostgreSQL and Redis.
  • For multi-replica AI gateways, the deciding questions are where state lives, how replicas share counters, and how metrics reach Prometheus when pods sit behind a load balancer.

AI gateways are the layer that routes, governs, and observes LLM traffic between applications and model providers, and platform teams that run Kubernetes expect them to install, scale, and upgrade like any other Helm release. Bifrost, the open-source AI gateway written in Go and built by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability on their own clusters. This guide compares five AI gateways on how they behave as Kubernetes workloads: Helm charts, horizontal scaling, stateless versus stateful configuration, health probes, clustering, in-VPC deployment, GitOps-friendly config, and Prometheus or OpenTelemetry metrics.

What Kubernetes Teams Need From AI Gateways

AI gateways for Kubernetes need an official Helm chart with a documented values file, a clear answer on where state lives, HTTP health probes, horizontal scaling that does not break rate limits, and metrics that Prometheus or an OpenTelemetry collector can collect from every replica. Governance features matter only after these operational basics hold.

The common failure points on a cluster are predictable: a SQLite file pinned to one pod, per-replica rate-limit counters that drift apart under an HPA, long streaming responses cut off during a rolling update, and a /metrics endpoint that only shows whichever pod the load balancer picked. For a broader walkthrough of these issues, see this guide to deploying a high-throughput AI gateway on Kubernetes.

Applications call an Ingress and Service that front horizontally scaled AI gateway pods, which read a shared database, export metrics to Prometheus and OpenTelemetry, and route to LLM providers

Figure 1: The gateway is one more Kubernetes workload: an Ingress in front, an HPA beside it, a database underneath, and telemetry flowing out.

The evaluation criteria below map each requirement to the question a platform engineer should ask during a proof of concept. Readers new to the category can start with what an AI gateway does and how it is structured.

Criterion Question to ask Why it matters on Kubernetes
Helm chart Is there an official chart with a versioned values reference? Determines whether installs and upgrades fit existing Helm and GitOps pipelines
State model Does the chart deploy a Deployment or a StatefulSet, and why? Stateful pods with local volumes limit how freely the HPA can scale
Health probes Which endpoints back liveness and readiness? Wrong probes cause restart loops or route traffic to pods that are not ready
Horizontal scaling Do budgets and rate limits stay correct across replicas? Per-pod counters multiply effective limits as replicas scale out
Graceful shutdown Can in-flight streaming responses finish during scale-down? LLM streams often run longer than default termination windows
Observability Can metrics and traces reach Prometheus or OTel from all pods? Scraping through a load balancer can miss replicas
Secrets and config Can keys come from Kubernetes Secrets or a vault, and config from Git? Keeps credentials out of values files and makes config reviewable

AI Gateways for Kubernetes Compared at a Glance

The five AI gateways below all install with Helm, but they split into two models: a single gateway workload configured through values (Bifrost, LiteLLM) and a controller that reconciles Kubernetes custom resources (Agent Router, agentgateway, Kong's ingress controller). The table summarizes what each project publishes.

Cells marked "Not published" mean the capability was not described on the vendor pages reviewed, not that it is absent. For a wider buyer-oriented view across deployment models, see the buyer's guide to AI gateways for LLM workloads and the LLM gateway buyer's guide.

Gateway Helm install Kubernetes model Required dependencies Health probes Multi-replica state Metrics and tracing
Bifrost bifrost/bifrost from the Bifrost chart repo Deployment (PostgreSQL) or StatefulSet (SQLite) None for SQLite; PostgreSQL for HA GET /health for liveness and readiness Cluster mode with gossip and Kubernetes discovery (Enterprise) /metrics, ServiceMonitor, Push Gateway, OTLP traces and metrics
Agent Router OCI charts ai-gateway-crds-helm and ai-gateway-helm Controller plus CRDs on Envoy Gateway Envoy Gateway 1.9.2 or higher Not published Not published OpenTelemetry example stack documented
agentgateway OCI charts agentgateway-crds and agentgateway Gateway API control plane and proxy Kubernetes Gateway API CRDs Not published Not published Metrics and logs documented
Kong AI Gateway kong/ingress from the Kong chart repo Ingress controller with DB-less Kong Gateway Konnect, or self-hosted Kong Gateway with AI plugins Not published Not published Analytics and monitoring policies
LiteLLM litellm-helm (monolithic) or litellm (componentized) Deployment; gateway, backend, and UI can scale separately PostgreSQL and Redis /health/liveliness and /health/readiness Not published ServiceMonitor and dedicated metrics listener

1. Bifrost

The Bifrost AI gateway is open source, installs on Kubernetes from one official Helm chart and routes traffic to 25+ providers and 10,000+ models through one OpenAI-compatible API. Bifrost adds 11 microseconds of overhead per request at 5,000 RPS with a 100% success rate in sustained benchmarks.

Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

Helm chart and values

The Bifrost Helm chart quickstart starts with adding the chart repository:

helm repo add bifrost https://maximhq.github.io/bifrost/helm-charts
helm repo update

The chart will not start without image.tag, which keeps installs pinned to a known release. Every field in values.yaml maps directly to the config.json the chart generates, and the values reference documents image, replica, autoscaling, resource, ingress, probe, and secret-reference parameters. The chart also ships ready-made values files such as sqlite-only.yaml, production-ha.yaml, external-postgres.yaml, and secrets-from-k8s.yaml, and later -f files override earlier ones, so a base file plus an environment overlay works without templating.

Stateless scaling, probes, and graceful shutdown

The chart picks the workload type from the storage backend. With PostgreSQL only, Bifrost runs as a Deployment; when any store uses SQLite, it runs as a StatefulSet with a PVC, as described in the Helm storage options. Logs can also move to ClickHouse while the config store stays on PostgreSQL.

  • Autoscaling: autoscaling.enabled creates an HPA with CPU and memory targets, and the default scale-down stabilization window is 300 seconds to protect long-lived streams.
  • Health probes: liveness and readiness both call GET /health, with configurable initial delays and periods.
  • Graceful shutdown: the default preStop sleep and a 60-second terminationGracePeriodSeconds let in-flight SSE streams finish before a pod stops.
  • Scheduling: pod anti-affinity, node selectors, and tolerations are first-class values, so replicas can be spread across nodes.

Cluster mode for high availability

Multiple replicas on a shared PostgreSQL database are not enough on their own, because each Bifrost replica loads configuration into memory at startup. Cluster mode in the Helm chart closes that gap by syncing rate limits, budget counters, and governance data across pods through a gossip protocol, with Kubernetes API discovery by label selector as the recommended option. Cluster mode is an Enterprise capability and requires PostgreSQL.

Three Bifrost pods discovered through the Kubernetes API exchange state over gossip, share an external PostgreSQL store, and push metrics to a Prometheus Push Gateway or OTel collector

Figure 2: A shared database seeds each replica at startup; gossip keeps rate limits, budgets, and config changes consistent while pods run.

The broader Bifrost clustering architecture supports six discovery methods (Kubernetes, Consul, etcd, DNS, UDP, and mDNS) plus a broker mode for platforms without peer-to-peer networking. This walkthrough of Bifrost cluster mode for enterprise high availability covers failover behavior in more depth.

Observability, secrets, and in-VPC deployment

Bifrost exposes Prometheus metrics at /metrics, and the chart can create a ServiceMonitor for the Prometheus Operator. For multi-replica setups, a Push Gateway avoids scraping gaps behind a load balancer, and the OpenTelemetry plugin exports traces in GenAI semantic conventions and pushes OTLP metrics from every node.

  • Secrets: every sensitive value has an existingSecret alternative, and vault-backed secret management adds AWS Secrets Manager, GCP Secret Manager, and HashiCorp Vault on Enterprise.
  • Governance: virtual keys carry budgets, rate limits, and provider access, and the chart can make them mandatory on every request.
  • Private networks: in-VPC deployments run on GCP, AWS, Azure, Cloudflare, and Vercel, and Enterprise images come from a private registry pulled through imagePullSecrets.

Teams that prefer Terraform can use the Terraform and Kubernetes guide, whose module targets EKS, GKE, AKS, or a plain Kubernetes Deployment. The Bifrost enterprise deployment overview covers in-VPC, air-gapped, and multi-cloud patterns, and the Bifrost benchmarks document the overhead figures above.

For day-two operations, see these Helm best practices and common pitfalls for Bifrost.

2. Agent Router (formerly Envoy AI Gateway)

Agent Router is the new name for Envoy AI Gateway, now an Agentic AI Foundation project built on Envoy Gateway. It installs on Kubernetes with two OCI Helm charts, one for custom resource definitions and one for the AI Gateway controller, and it is configured through Kubernetes resources such as AIGatewayRoute.

The installation guide installs ai-gateway-crds-helm first and ai-gateway-helm second, both into the envoy-ai-gateway-system namespace. The prerequisites require Envoy Gateway 1.9.2 or higher and recommend a clean Envoy Gateway installation, because existing custom configuration can conflict. Rate limiting and InferencePool support are enabled by passing addon values files to the Envoy Gateway installation.

  • Configuration model: custom resources reconciled by a controller; the rename changed no CRD or API names.
  • Stated goals: failover across providers and self-hosted models, upstream authorization, rate limiting, and usage visibility.
  • Outside Kubernetes: the aigw run CLI starts a standalone OpenAI-compatible router for local use.

Best for: platform teams already operating Envoy Gateway who want AI routing expressed as Kubernetes custom resources. Teams comparing it against single-chart gateways can review these Envoy AI Gateway alternatives for LLM routing.

3. agentgateway

agentgateway is an open-source gateway for LLM, MCP, and agent traffic that runs on Kubernetes as a Gateway API control plane and proxy. Its Helm installation applies the Kubernetes Gateway API CRDs, then installs the agentgateway-crds and agentgateway charts from an OCI registry into the agentgateway-system namespace.

The Kubernetes documentation lists separate install paths for Helm, ArgoCD, and FluxCD, which suits teams that deliver cluster add-ons through GitOps controllers. Its documentation navigation covers virtual keys, budget and spend limits, guardrails, MCP authentication, CEL-based RBAC, inference routing, and metrics and logs. The project has joined the Agentic AI Foundation.

  • Configuration model: Gateway API resources plus agentgateway policies, applied with kubectl or a GitOps controller.
  • Upgrade path: CRDs are released as their own chart, so CRD and control plane upgrades are versioned separately.
  • Scope: LLM consumption, MCP connectivity, and agent connectivity in one proxy.

Best for: teams standardized on the Kubernetes Gateway API that want AI and MCP routing under the same resource model. Teams weighing Gateway API projects against a dedicated AI gateway can compare open-source AI gateway platforms for in-VPC teams.

4. Kong AI Gateway

Kong AI Gateway is Kong's connectivity and governance layer for LLM, MCP, and agent-to-agent traffic. On Kubernetes, data plane nodes run in your cluster and connect to Konnect for configuration and observability, while the Kong Ingress Controller installs from the kong/ingress Helm chart in Kong's chart repository.

The default kong/ingress values install the ingress controller in Gateway Discovery mode with a DB-less Kong Gateway, which Kong describes as the recommended topology. AI Gateway 2.x replaces the plugin-centric AI Proxy model with AI Model Provider and AI Model entities. Kong also documents running AI Gateway on self-hosted Kong Gateway with the Kong Gateway data model and AI plugins.

  • Configuration model: Konnect-managed control plane with in-cluster data planes, or self-hosted Kong Gateway with AI plugins.
  • Traffic types: LLM, MCP, and A2A through a single endpoint.
  • Operational fit: strongest where Kong already fronts internal APIs.

Best for: organizations already running Kong Gateway or Konnect that want AI traffic governed alongside existing APIs. Teams reconsidering that dependency can review Kong AI Gateway alternatives.

5. LiteLLM

LiteLLM is a Python-based LLM proxy that supports Kubernetes through two Helm charts: litellm-helm, a monolithic chart where one image serves LLM traffic, management APIs, and the UI, and a componentized litellm chart that deploys gateway, backend, and UI as separately scaled services.

Its deployment guide states that the Helm path needs PostgreSQL and Redis reachable from the cluster, and recommends managed services for both. The monolithic chart supports autoscaling through autoscaling.* or KEDA, PodDisruptionBudgets, a Prometheus ServiceMonitor, graceful drain on shutdown, and ArgoCD or Helm hooks for its database migration job. Example probes use /health/liveliness and /health/readiness.

  • Configuration model: Helm values plus a proxy config.yaml mounted into the pod.
  • Scaling model: the componentized chart scales the gateway independently of the management API and UI.
  • Dependencies: PostgreSQL and Redis must exist before install.

Best for: Python-centric platform teams that accept running Redis and PostgreSQL alongside the gateway. Teams moving off LiteLLM for performance or governance reasons can review the Bifrost LiteLLM alternative overview.

StatefulSet vs Deployment: Stateless and Stateful AI Gateways

An AI gateway should run as a stateless Deployment in production, with configuration, logs, and counters held in an external database or synced between replicas. A StatefulSet with a pod-local volume is reasonable for development or a single node, but it ties data to one pod and limits how freely the HPA can add or remove replicas.

Figure 3 shows how this works in the Bifrost chart: the storage backend decides the workload type. Any SQLite store produces a StatefulSet with a PVC; PostgreSQL alone produces a Deployment that an HPA can scale from three to fifteen replicas without moving volumes.

Decision flow showing that SQLite storage produces a StatefulSet with a persistent volume, while PostgreSQL-only storage produces a stateless Deployment that scales with the HPA

Figure 3: Moving state out of the pod into PostgreSQL is what turns the gateway into a stateless Deployment that the HPA can scale freely.

Stateless pods solve storage, but they do not solve shared in-memory state. Rate limits and budgets enforced per pod multiply as replicas scale, so a 100-requests-per-minute limit on three replicas can admit up to 300. Bifrost addresses this with cluster mode across replicas, and budgets and rate limits then hold at the cluster level. The Kubernetes Horizontal Pod Autoscaler documentation explains the scaling behavior the HPA settings control.

Concern StatefulSet with SQLite Deployment with PostgreSQL Deployment with PostgreSQL and cluster sync
Typical use Development, single node Production without strict shared limits Production with HPA and shared budgets
Config changes Local to the pod Read from the database at startup Propagated to running replicas
Rate limits and budgets Single pod only Counted per replica Synced across replicas

Choosing an AI Gateway for Helm and GitOps Workflows

The right AI gateway for a Kubernetes team depends on the platform it already runs. Teams that want governance, HA, and low overhead in one chart should start with Bifrost; teams standardized on Gateway API or Envoy should evaluate agentgateway or Agent Router; Kong estates should evaluate Kong AI Gateway; and Python-first teams may accept LiteLLM's extra dependencies.

Selection flow that maps team requirements such as built-in governance, Gateway API alignment, an existing Kong estate, or a Python-centric stack to a matching AI gateway type

Figure 4: Start from the platform you already operate, then check whether the gateway covers governance and HA without extra components.

GitOps adds one more requirement: configuration that lives in Git, renders the same way every time, and changes through pull requests. The OpenGitOps principles describe this as declarative, versioned, and continuously reconciled state. Bifrost fits that model in two ways: Helm values map one-to-one to the gateway config, and a declarative config.json for GitOps workflows can seed or fully define the gateway with a published JSON schema for validation in CI.

  • Pin versions: set the image tag explicitly and keep values files in the same repository as the rest of the cluster.
  • Review diffs: run helm diff upgrade before applying changes, as the Bifrost values reference shows.
  • Keep secrets out of Git: reference Kubernetes Secrets or a vault rather than plaintext keys.
  • Push metrics from every pod: use a Push Gateway or OTLP metrics push, then build dashboards on Prometheus metrics and dashboards for LLM observability and OpenTelemetry traces for LLM observability.

For a step-by-step build, see the Kubernetes AI gateway deployment guide.

Frequently Asked Questions

What is an AI gateway?

An AI gateway is a service that sits between applications and LLM providers, exposing one API while handling routing, failover, authentication, budgets, rate limits, and logging. On Kubernetes it runs as a Deployment or StatefulSet behind a Service and Ingress. Bifrost, for example, routes to 25+ supported providers through one OpenAI-compatible API.

Should an AI gateway run as a StatefulSet or a Deployment?

An AI gateway should run as a Deployment in production, with state in an external database such as PostgreSQL. A StatefulSet is appropriate when the gateway stores data in a pod-local volume, such as SQLite, which suits development or single-node setups. The Bifrost Helm chart chooses between the two automatically based on the configured storage backend.

Can the Kubernetes HPA scale an AI gateway?

Yes, the Kubernetes HPA can scale an AI gateway on CPU and memory targets, provided the gateway is stateless or shares state across replicas. Set a scale-down stabilization window so streaming responses are not cut off, and confirm that rate limits and budgets are synchronized, because per-pod counters let total throughput exceed the configured limits as replicas increase.

Which health probes should an AI gateway expose?

An AI gateway should expose an HTTP endpoint for liveness and readiness so Kubernetes can restart failed pods and withhold traffic from pods that are not ready. Bifrost uses GET /health for both probes, with configurable delays and periods in the chart. The Kubernetes probe documentation explains how each probe type affects pod lifecycle.

How do you manage AI gateway configuration with GitOps?

Manage AI gateway configuration with GitOps by storing Helm values or a declarative config file in Git, pinning image versions, and letting a controller such as Argo CD or Flux reconcile the cluster. Keep provider keys in Kubernetes Secrets or a vault. Bifrost Helm values map directly to its config.json, which carries a published schema in the Helm values reference for validation.

Is an AI gateway the same as the Kubernetes Gateway API?

No, an AI gateway and the Kubernetes Gateway API solve different problems. The Gateway API is a set of Kubernetes resources for describing HTTP and TCP routing, while an AI gateway understands LLM requests, tokens, providers, and model-level policy. Some AI gateways, such as agentgateway and Agent Router, use Gateway API resources as their configuration surface; Bifrost uses Helm values and governance features built into the gateway.

Try Bifrost on Your Cluster

AI gateways for Kubernetes should install from a versioned Helm chart, scale as stateless Deployments, keep limits consistent across replicas, and report metrics from every pod. Bifrost covers each of these in one chart, with cluster mode, in-VPC deployment, and vault-backed secrets available on Bifrost Enterprise. To see the Helm chart, cluster mode, and governance running against your own workloads, book a demo with the Bifrost team, or browse the Bifrost resources hub for deployment guides.