Best AI Gateway to Route Codex CLI to Any Model
An AI gateway routes Codex CLI to non-OpenAI models by exposing an OpenAI-compatible endpoint and translating each request. This comparison ranks Bifrost, LiteLLM, Kong AI Gateway, and OpenRouter on routing, governance, deployment, and Codex CLI setup.
TL;DR
- Codex CLI can reach non-OpenAI models only through an OpenAI-compatible endpoint, which is why teams put an AI gateway between Codex CLI and their providers.
- The routed model must support tool calling, because Codex CLI performs file edits, terminal commands, and code changes through function calls.
- Bifrost routes Codex CLI to 25+ providers and 10,000+ models with 11 microseconds of overhead per request at 5,000 requests per second.
- LiteLLM, Kong AI Gateway, and OpenRouter also route Codex CLI, trading off self-hosting, governance depth, and setup effort differently.
- The correct Bifrost setup uses a named
model_providersentry inconfig.tomlwith HTTPS mode, rather thanopenai_base_urlalone.
Codex usage grew from fewer than 1 million weekly active users in February 2026 to 5 million by early June, according to The New Stack, and much of that growth runs through the Codex CLI. By default, the CLI sends every request to OpenAI models. This article ranks the best AI gateways to route Codex CLI to any model in 2026, beginning with Bifrost, the open-source AI gateway built by Maxim AI that ships first-class Codex CLI integration and a one-command launcher. Bifrost is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability.
For teams that want a Codex router that sends Codex CLI to GPT-5.4 for hard reasoning, Claude Sonnet for explanations, Gemini Flash for cost-sensitive edits, or a Groq-hosted model for speed, the only clean answer is an AI gateway. The right Codex LLM gateway sits between Codex CLI and your LLM providers, translates the OpenAI-format request transparently, and handles routing, failover, governance, and observability behind one base URL.
Why Codex CLI Needs an AI Gateway
Codex CLI communicates with OpenAI over standard HTTP, controlled by openai_base_url or a named model_providers entry in ~/.codex/config.toml and an OPENAI_API_KEY value. Pointing Codex CLI at a gateway is the supported, OpenAI-documented path for routing Codex CLI through an LLM proxy or router.

Figure 1: The gateway turns model choice, failover, and spend control into configuration instead of per-developer setup.
Once the request hits the Codex gateway, it can be routed to any provider whose model supports tool calling, since Codex CLI relies heavily on function calls for file operations, terminal commands, and code edits. Without a gateway, every Codex CLI session is a direct call to OpenAI with no spend controls, no model access scoping, no failover, and no cross-team observability. With a gateway, those become infrastructure concerns instead of per-developer concerns.
Teams that also need spend caps can pair routing with the budget controls covered in the guide to managing Codex CLI token spend.
Key Criteria for Evaluating an AI Gateway for Codex CLI
The best AI gateway for Codex CLI is judged on eight criteria: Codex CLI integration, tool-use coverage, multi-provider routing, overhead, governance, observability, deployment model, and open-source posture. Before ranking, every option should be evaluated against the same baseline. The criteria that matter for Codex CLI specifically include:
- Codex CLI integration: a documented setup path with the correct provider endpoint (Codex CLI uses
/openai/v1paths and the Responses API) - Tool-use coverage: routing only to models that support tool calling reliably (Claude Sonnet, GPT-5.4, GPT-5.5, Gemini 2.5 Pro)
- Multi-provider routing: weighted distribution and explicit fallback chains across OpenAI, Anthropic, Google, and others
- Gateway overhead: latency added per Codex CLI request, especially under rapid tool-call sequences
- Governance: virtual keys, per-developer budgets, and rate limits with clear reset windows
- Observability: per-request token tracking, cost attribution, and model selection visibility
- Deployment model: self-hosted, managed, or hybrid (including in-VPC for regulated codebases)
- Open-source posture: license transparency and ability to inspect or extend the gateway
These criteria separate a basic OpenAI proxy from a production-grade Codex CLI gateway. Teams running side-by-side evaluations can use the LLM Gateway Buyer's Guide for a deeper capability matrix.

Figure 2: Weights spread load across providers, and the fallback chain keeps a session running when one provider returns errors.
1. Bifrost: The Best AI Gateway to Route Codex CLI to Any Model
Bifrost, the AI gateway, is a high-performance, open-source AI gateway written in Go. It ships dedicated Codex CLI integration and adds only 11 microseconds of overhead per request in sustained 5,000 RPS benchmarks, so the gateway is effectively invisible during Codex CLI's rapid tool-call sequences. For context, a network round trip to an LLM provider takes tens of milliseconds, three orders of magnitude larger than the 11 µs Bifrost adds.
How Bifrost routes Codex CLI to any model
Setup takes three steps. First, start the gateway with npx -y @maximhq/bifrost. Second, run /logout inside Codex CLI to clear any existing OAuth session (Codex CLI prefers an existing OAuth session over a custom API key). Third, export the Bifrost virtual key as OPENAI_API_KEY and add a Codex CLI custom provider entry to ~/.codex/config.toml, as the Bifrost Codex CLI setup guide shows:
model = "openai/gpt-5.4"
model_provider = "bifrost"
[model_providers.bifrost]
name = "Bifrost"
base_url = "http://localhost:8080/openai/v1"
env_key = "OPENAI_API_KEY"
wire_api = "responses"
supports_websockets = false
A named provider is preferred over openai_base_url, because with openai_base_url Codex keeps its built-in OpenAI provider and sends a client-side web tool that some providers, including Amazon Bedrock, reject. Setting supports_websockets = false keeps Codex CLI on HTTPS, which non-OpenAI models require.
From there, Codex CLI talks to Bifrost as if it were OpenAI, and Bifrost translates and routes requests to any of 25+ supported providers and 10,000+ models including Anthropic, AWS Bedrock, Google Vertex AI, Azure OpenAI, Mistral, and Groq. Developers can switch models mid-session using Codex CLI's /model command with the provider/model format (for example /model anthropic/claude-sonnet-4-5-20250929 or /model gemini/gemini-2.5-pro), and the gateway handles the provider translation transparently. The walkthrough on switching Codex CLI between Claude, Gemini, Llama, and more shows per-provider examples.

Figure 3: Codex CLI only ever speaks the OpenAI Responses API; the provider prefix in the model name decides where the gateway sends it.
Teams that want even faster setup can use the Bifrost CLI, an interactive launcher that installs Codex CLI if needed, sets the base URL, stores the virtual key in the OS keyring, and lists available models in a single npx -y @maximhq/bifrost-cli command. No environment variables, no manual config edits.
What sets Bifrost apart for Codex CLI
- First-class Codex CLI support: dedicated docs, the
/openai/v1endpoint path Codex CLI requires, and the Bifrost CLI for one-command launches - Weighted multi-provider routing: split traffic 70/30 between primary and secondary providers with provider routing, with automatic failover to the remaining providers
- Microsecond-level overhead: 11 µs per request at 5,000 RPS with a 100% success rate, verified through public benchmarks
- Hierarchical governance: virtual keys with per-developer, per-team, and per-customer budgets and rate limits
- MCP gateway: native Model Context Protocol support, so Codex CLI sessions can use centrally managed MCP tools alongside model routing
- Built-in observability: Prometheus metrics, OpenTelemetry traces, and an enterprise Datadog connector with zero custom instrumentation
- Enterprise-ready: in-VPC deployments, secret management with HashiCorp Vault, AWS, or GCP, OIDC, RBAC, and signed audit logs of administrative activity
For teams running Codex CLI across hundreds of developers, Bifrost generates structured telemetry on every request, including model used, provider routed to, input and output token counts, latency, virtual key identifier, and outcome. Platform teams can answer questions that are invisible in a direct-to-OpenAI setup, as the walkthrough on using Codex CLI with multiple model providers demonstrates: which team's Codex CLI sessions are generating the most tokens, which model is being used for which task type, and where latency spikes are occurring.
Best fit: engineering teams that want production-grade multi-provider routing for Codex CLI with hierarchical governance, observability, and an open-source core.
2. LiteLLM: Python-Native Codex CLI Routing
LiteLLM is an open-source Python proxy that exposes a unified OpenAI-compatible interface to 100+ LLM providers. A Codex CLI LiteLLM setup points Codex CLI's openai_base_url at a LiteLLM proxy, which is straightforward, and LiteLLM's broad provider coverage means almost any model with tool-use support is reachable.
LiteLLM's proxy includes virtual keys with per-key, team, and user budgets, load balancing with failover, and an MCP gateway. The trade-offs are runtime and supply-chain exposure. LiteLLM is written in Python and is distributed through PyPI, and a March 2026 PyPI supply-chain compromise affected published LiteLLM packages, which raised additional concerns for self-hosted deployments. Teams considering migration can review the LiteLLM alternatives comparison and the step-by-step guide to migrating from LiteLLM.
Best fit: Python-first teams that need maximum provider breadth and can own the patching cadence of a Python dependency chain.
3. Kong AI Gateway: API Management Extended to Codex CLI
Kong AI Gateway extends Kong's mature API management platform to LLM traffic, including Codex CLI, which Kong lists among the AI command-line tools it can proxy. In AI Gateway 2.x, the proxy is configured through the AI Model entity rather than the earlier AI Proxy plugins, with Codex CLI pointed at a Kong endpoint. Kong supports routing and load balancing across providers, failover, semantic caching, and token-based rate limiting through its AI Rate Limiting Advanced plugin.
Kong's strength is its plugin architecture and operational maturity. Organizations already running a Kong mesh can extend existing API governance policies to Codex CLI traffic without adopting a separate gateway. The trade-off is setup complexity. Kong's AI capabilities sit on top of a general API management platform, so configuring Codex CLI through Kong typically requires more declarative configuration than purpose-built AI gateways. Teams comparing both approaches can use the AI gateway buyer's guide feature matrix to map features.
Best fit: organizations already invested in the Kong ecosystem that want Codex CLI routing added to existing API infrastructure.
4. OpenRouter: Managed Routing Across a Large Model Catalog
OpenRouter aggregates hundreds of models behind a single OpenAI-compatible API. For Codex CLI, OpenRouter functions as a drop-in OpenAI-compatible endpoint that handles fallbacks automatically. For prototyping or solo developers, the breadth of model access and pay-as-you-go pricing is useful.
The constraints are governance and deployment. OpenRouter is a hosted service, so prompts and code leave the organization's network, and cost attribution by team or developer requires building an additional layer. For Codex CLI specifically, teams need to confirm that every model they expose supports tool calling, since Codex CLI fails on models without it. OpenRouter can also sit behind Bifrost as one upstream provider, which keeps Bifrost virtual keys and budgets in front of it; the supported providers list includes an openrouter prefix.
Best fit: solo developers and small teams that want the broadest model selection and are comfortable with a managed-only deployment.
How the Best AI Gateways for Codex CLI Compare
The four AI gateways for Codex CLI differ most on deployment model, governance depth, and runtime: Bifrost and LiteLLM are open-source and self-hosted, Kong extends an API management platform, and OpenRouter is hosted only.
| Capability | Bifrost | LiteLLM | Kong AI Gateway | OpenRouter |
|---|---|---|---|---|
| Documented Codex CLI integration | Yes | Yes | Yes | Via OpenAI-compatible endpoint |
| Gateway overhead | 11 µs at 5K RPS | Not published | Not published | Hosted, network-bound |
| Multi-provider weighted routing | Yes (per-VK weights) | Yes (load balancing) | Yes | Automatic fallbacks |
| Automatic failover | Native, configurable chains | Yes | Yes | Yes |
| Hierarchical governance | Yes (virtual keys, teams, customers) | Per-key, team, and user budgets | Token-based rate limiting | Not published |
| Native MCP gateway | Yes | Yes | Yes | No |
| Self-hosted | Yes (open source, Go) | Yes (open source, Python) | Yes | No, hosted only |
| One-command Codex CLI launch | Yes (Bifrost CLI) | No | No | No |
For a deeper feature-by-feature breakdown, see the complete LLM gateway buyer's guide, and for the broader cluster view, the ranking of the best AI gateway for Codex CLI covers governance and security beyond routing.
Choosing the Right Gateway to Route Codex CLI
The right gateway to route Codex CLI depends on team posture: deployment requirements, existing infrastructure, and how much governance the rollout needs. For Python-first teams with broad provider needs, LiteLLM offers reach with a Python dependency chain to maintain. For Kong-native API teams, Kong AI Gateway folds Codex CLI into existing infrastructure. For solo developers, OpenRouter delivers a large hosted model catalog. For engineering teams running Codex CLI at scale where multi-provider routing must combine microsecond-level overhead, hierarchical governance, native MCP support, and an open-source core, Bifrost is the most complete option.
| Team profile | Recommended gateway | Why |
|---|---|---|
| Engineering org rolling out Codex CLI to many developers | Bifrost | Virtual keys, hierarchical budgets, 11 µs overhead, self-hosted |
| Python-first platform team | LiteLLM | Python SDK and proxy in one ecosystem |
| Organization standardized on Kong | Kong AI Gateway | Reuses existing API management policies |
| Solo developer prototyping models | OpenRouter | Hosted catalog with no infrastructure to run |
Teams running more than one coding agent can apply the same routing layer to Gemini CLI; the guide to routing Gemini CLI and Codex through an AI gateway covers the shared setup.
Frequently Asked Questions
How can I use Codex with a multi-model LLM API gateway?
Point Codex CLI at the gateway's OpenAI-compatible endpoint with a named model_providers entry in ~/.codex/config.toml, set the gateway key in OPENAI_API_KEY, and run /logout to clear any OAuth session. With Bifrost, developers then pick any configured model with codex --model provider/model, and the gateway translates the request to that provider.
Can Codex CLI use LiteLLM?
Yes. Codex CLI can use LiteLLM by pointing its base URL at a LiteLLM proxy, which exposes an OpenAI-compatible interface to 100+ LLMs. Teams that want the same routing with a Go runtime, hierarchical budgets, and a one-command launcher use Bifrost instead, and the LiteLLM migration guide covers the switch.
What does openai_base_url do in Codex CLI?
openai_base_url changes the base URL of the built-in OpenAI provider so Codex CLI sends requests to a proxy or router instead of OpenAI. For Bifrost, a named model_providers entry is preferred, because it keeps Codex from sending a client-side web tool that providers such as Amazon Bedrock reject.
What custom model providers work with Codex CLI for coding tasks?
Any provider reachable through an OpenAI-compatible Responses endpoint works, provided the model supports tool calling. Through Bifrost, Codex CLI can use anthropic, gemini, vertex, bedrock, azure, mistral, groq, deepseek, xai, ollama, openrouter, and more with the provider/model format, in HTTPS mode rather than WebSocket mode.
Can Codex CLI use OpenRouter?
Yes. Codex CLI can call OpenRouter directly as an OpenAI-compatible endpoint, or route through Bifrost with OpenRouter configured as one upstream provider. The second option keeps virtual keys, budgets, and request logs in front of OpenRouter traffic, and lets teams fail over to another provider when an OpenRouter model is unavailable.
Do non-OpenAI models work with every Codex CLI feature?
Most coding workflows work when the model supports tool calling, but non-OpenAI models do not appear in the Codex CLI /models picker by default and receive fallback metadata. Bifrost documents a local model catalog file that adds them to the picker with correct context windows.
Try Bifrost as Your Codex CLI Gateway
Among the best AI gateways to route Codex CLI to any model in 2026, Bifrost is the only option that combines first-class Codex CLI integration, a one-command launcher, 11 µs overhead, weighted multi-provider routing, hierarchical governance, native MCP support, and a fully open-source core. Teams can start Bifrost with a single npx -y @maximhq/bifrost command, run /logout inside Codex CLI, add a Bifrost provider entry to config.toml, and route Codex CLI sessions through any tool-use-capable model on day one. To see Bifrost handling Codex CLI traffic at scale and discuss a deployment plan for your team, book a Bifrost demo.