Try Bifrost Enterprise free for 14 days. Request access

Policy-Based Governance at the Gateway: One Control Plane for Every AI Call

Policy-Based Governance at the Gateway: One Control Plane for Every AI Call

TL;DR

  • Gateway-enforced AI policy turns written rules for access, spend, routing, and content safety into checks that run on every model and tool request.
  • A Cloud Security Alliance survey published in April 2026 found that 82% of organizations had discovered previously unknown AI agents in their environments in the past year.
  • In Bifrost, policy attaches to virtual keys, budgets are checked at the customer, team, virtual key, and provider-config levels, and CEL expressions drive routing rules and guardrails.
  • The same control plane covers MCP tool access through per-key tool filtering and Virtual MCPs, and extends to endpoint AI with Bifrost Edge, currently in alpha.

Policy-based AI governance is the practice of declaring access, spend, routing, and safety rules once and enforcing them at runtime on every model request, instead of writing them down and trusting each application to comply.

Most enterprises now run several providers across production services, internal tools, IDE assistants, and agents, and the rules meant to control that traffic live in wiki pages rather than in the request path. Bifrost, the open-source AI gateway built in Go by Maxim AI, is built for enterprise teams that need one control plane for every AI call across providers, models, agents, and tools. This post covers how policy is modeled, evaluated, and enforced at the gateway, and what that architecture gives platform and security teams that application-level controls cannot. The broader path from written policy to runtime controls is covered in the complete guide to AI governance from policy to runtime enforcement.

What Policy-Based AI Governance Means at the Gateway

Policy-based AI governance at the gateway means access, spend, throughput, routing, and content rules are declared centrally and evaluated inside the request path. Every model call passes through the same enforcement point, so a rule written once applies to every application, provider, agent, and user without a change to application code.

A complete gateway policy answers six questions about a request before it reaches a provider:

  • Identity: which consumer, team, or customer is making this call?
  • Access: which providers, models, and provider API keys may they use?
  • Spend: which budgets does this call draw down, and what happens when one is exhausted?
  • Throughput: how many requests and tokens are permitted in the current window?
  • Content: which safety rules apply to the prompt and to the response?
  • Evidence: what record of the decision is retained, and for how long?

Because the Bifrost AI gateway sits in the data path between applications and providers, all six are resolved before the request leaves the network. Governance becomes part of the request pipeline rather than a reporting layer assembled after the fact, which is the core reason an AI gateway acts as the control plane for enterprise LLM traffic.

Why Policy at the Application Layer Does Not Hold

Application-level governance fails because policy is copied rather than centralized. Each service holds its own provider keys, its own approved-model list, and its own retry logic. Nothing evaluates those choices at runtime, and nothing produces a single record of who called which model with what data.

The failure modes are consistent across multi-provider environments:

  • Credential sprawl: raw provider keys sit in environment variables across dozens of services and cannot be scoped or revoked individually.
  • Untracked spend: cost lands on a provider invoice, attributed to an account rather than to a team, product, or customer.
  • Policy drift: an approved-model list in a document does not stop a new service from calling an unapproved model.
  • Blind spots: coding agents, IDE assistants, and desktop chat apps never pass through the platform team's code path.
  • Unusable evidence: logs are per-application, in different formats, with different retention windows.

The visibility gap is measurable. A Cloud Security Alliance survey published in April 2026 found that 82% of organizations had discovered previously unknown AI agents in their environments in the past year, and 65% reported an AI agent-related incident in the previous 12 months, even though 68% believed their visibility into agents was strong. Governance that depends on each team opting in produces exactly this result, which is why the governance model has to live at the infrastructure layer.

The Bifrost Policy Model: Virtual Keys, Hierarchy, and Access Profiles

In the open-source Bifrost gateway, policy attaches to a virtual key. A virtual key is the credential an application or user presents, and it carries the allowed providers and models, the budget, the rate limits, and the MCP tool access for that consumer. Real provider keys stay inside the gateway and are never handed to an application.

Virtual keys are the primary governance entity, and the model has a few properties worth knowing before designing policy:

  • Access is deny-by-default: a virtual key with an empty provider configuration blocks all providers unless an administrator turns on the explicit Allow all providers option, so permissions are granted deliberately rather than assumed.
  • Attachment is exclusive: a virtual key belongs to one team, or one customer, or neither, which keeps cost and access hierarchies unambiguous.
  • Keys can be pinned to specific provider API keys, so a key issued for a development environment cannot spend production quota.
  • Status is a switch: marking a virtual key inactive stops its access immediately, without a redeploy.
  • Authentication uses familiar headers, including the OpenAI, Anthropic, Gemini, and Azure styles, so existing clients present a virtual key the way they present a provider key.

Issuing keys by hand does not scale past a few teams. Access profiles solve that: a profile is a reusable policy template (provider list, model allow-list, budgets, rate limits, MCP tool access) that Bifrost copies per user and materializes as an auto-issued virtual key. Attach a profile to a role and users gaining that role are provisioned automatically, each with independent counters. Profile-managed keys are write-protected, so a user cannot edit around their own policy, and every profile change is recorded with a snapshot history.

The key-level mechanics are covered in more depth in enterprise AI governance with virtual keys.

Enforcing Spend and Rate Policy on Every AI Call

Bifrost enforces spend policy hierarchically. Budgets exist independently at the customer, team, virtual key, and provider-config levels, and every applicable budget is checked before a request proceeds. Any single budget without remaining balance blocks the call, so overspend is prevented at request time rather than discovered at invoice time.

The hierarchical budget structure supports the patterns platform teams actually need:

  • Customer level: cap total spend for an external tenant across all of their teams and keys.
  • Team level: give each internal team a monthly ceiling that its keys draw from.
  • Virtual key level: bound a single application, environment, or agent, with token and request rate limits alongside the budget.
  • Provider config level: cap spend and throughput per provider within one key, so an expensive frontier model cannot absorb the whole allocation.

Budget reset windows run from one day to one year (1d, 1w, 1M, 1Q, 1Y), while rate-limit windows run from one minute to one day. Calendar-aligned budgets reset at UTC day, week, month, quarter, or year boundaries instead of on a rolling window, which matches how finance teams close periods.

Cost is computed from token usage and model catalog pricing, which syncs every 24 hours by default, with prompt-cache and batch requests priced accordingly, so attribution reflects what was actually consumed.

Expression-Based Policy: CEL Rules for Routing and Guardrails

Static allow-lists cannot express rules like "send batch embedding traffic to a cheaper provider" or "inspect any prompt from the support application for credentials." Bifrost handles these with expression-based policy, using CEL, the open Common Expression Language maintained by Google, to evaluate conditions against the live request.

Routing rules run before governance provider selection and can override it. Rules are organized by scope with first-match-wins evaluation: virtual key scope first, then team, then customer, then global, with priority ordering inside each scope. A matched rule can chain, making its resolved provider and model the new context for another pass. If no rule matches, the incoming provider and model are used unchanged.

Guardrails use the same expression model with two objects:

  • Rules define when and what to evaluate, written in CEL, and can apply to inputs, outputs, or both.
  • Profiles define how content is evaluated, covering native Secrets Detection (Gitleaks-backed) and Custom Regex (including a PII template), plus external providers including Microsoft Presidio, Azure AI Language PII, AWS Bedrock Guardrails, Azure Content Safety, Google Model Armor, CrowdStrike AIDR, Gray Swan, Patronus AI, Check Point's AI Agent Security, and Repello Argus.

Profiles are reusable across rules, and a single rule can call several profiles for layered checks. The native secrets detection guardrail is a common first profile for developer traffic. Rules can inspect prompts before they reach a model, responses before they return to the caller, or both. Mapping written company policy onto these rules is covered in enterprise AI guardrails for company policy and compliance.

Compliance Policy: RBAC, Audit Logs, and Deployment Control

Evidence is a policy dimension of its own. Frameworks such as the NIST AI Risk Management Framework and the phased obligations under the EU AI Act emphasize ongoing risk management and record-keeping, which is simplest when the enforcement point and the record of enforcement are the same system.

Bifrost separates operator permissions from operator visibility:

  • Role-based access control governs which operations a user can perform, using the Admin, Developer, and Viewer system roles or custom roles, with assignment driven by OIDC groups and claims.
  • Data access control governs which rows a user can see, scoping virtual keys, prompts, and routing rules to the teams a role is entitled to view.
  • Audit logs record administrative activity with HMAC-signed entries, configurable retention, filtering by action, outcome, and date, export as JSON, JSON Lines, or Syslog, and continuous archival to S3 or GCS for long-term retention.

For regulated industries, the deployment target matters as much as the policy. Bifrost Enterprise runs in-VPC, and on-premise and air-gapped deployments are supported, so policy enforcement and the audit trail stay inside the same boundary as the data they describe.

One Control Plane for MCP Tools and Endpoint AI

Model calls are only part of the traffic. Agents reach tools through MCP servers, and those tool calls read files, query internal APIs, and take actions. Policy that stops at the model boundary leaves the tool boundary open, which is why tool scope belongs in the same control plane.

MCP tool filtering applies a deny-by-default tool allow-list to each virtual key, enforced at inference time and again at tool execution. Virtual MCPs, previously called MCP tool groups, are named bundles of tools drawn from one or more MCP servers, served at their own /mcp/<slug> endpoint and assigned to virtual keys, with Enterprise adding access-profile grants and Data Access Control scoping. Used as an MCP gateway, Bifrost applies the same virtual key identity and access policy to tool calls that it applies to model calls.

The remaining gap is traffic that never points at the gateway. The Bifrost AI gateway is the control plane and policy engine, and Bifrost Edge extends that same governance to the endpoint by running on company machines and routing AI traffic from desktop chat apps, browser AI, coding agents, and the MCP servers those tools connect to through the organization's Bifrost. The virtual keys, budgets, and guardrails already configured at the gateway are what apply on the laptop. Bifrost Edge is currently in alpha, and teams request access through the alpha sign-up on the Edge overview page. How this applies to developer machines is covered in how to govern AI coding agents at scale.

Where Each Policy Dimension Is Enforced

Each of the six policy questions maps to a specific Bifrost mechanism and a specific point in the request path. The table below summarizes that mapping for teams designing gateway policy.

Policy dimensionBifrost mechanismEnforcement point
IdentityVirtual keys, OIDC user provisioning, access profilesEvery request authenticates with a virtual key
AccessProvider configs, allowed models, key restrictions, routing rulesBefore provider selection
SpendHierarchical budgets at customer, team, virtual key, and provider configBefore the request proceeds
ThroughputToken and request rate limits at virtual key and provider configBefore the request proceeds
ContentCEL guardrail rules and profilesOn the prompt, the response, or both
EvidenceRequest logs, HMAC-signed audit logs, archivalEvery request and every administrative change

The layered view of policy, runtime, and observability tools is compared in AI governance tools by layer.

Common Questions About Policy-Based AI Governance

What is policy-based AI governance?

Policy-based AI governance is the practice of declaring rules for AI access, spend, routing, and content safety centrally and enforcing them on every model and tool request at runtime. In Bifrost, those rules attach to virtual keys and are evaluated in the gateway's request path, so they apply to every application and provider without code changes in each service.

How is this different from API gateway policy?

An API gateway authenticates callers and shapes HTTP traffic, but it is not built around tokens, model pricing, prompts, or tool calls. Policy-based AI governance adds model-aware dimensions: per-model access, token and cost budgets, prompt and response inspection, and MCP tool scope.

Does enforcing policy in the request path add latency?

Budget, usage, and provider configuration state is held in memory rather than looked up in a database per request. Bifrost adds 11 microseconds of overhead per request at 5,000 requests per second in sustained benchmarks, which keeps enforcement well inside the noise of a provider round trip.

Who defines policy, and who consumes it?

Platform and security teams define the governance rules once: profiles, budgets, guardrail rules, and routing rules. Developers consume them by presenting a virtual key against an OpenAI-compatible endpoint, which means adopting policy costs them a base URL change rather than a rewrite.

What is CEL, and why use it for AI policy?

CEL, the Common Expression Language, is an open expression language for evaluating conditions against structured data. Bifrost uses CEL for routing rules and guardrail rules because it can express conditions on the live request, such as the model, headers, or virtual key, without custom code. Routing rules evaluate first-match-wins across virtual key, team, customer, and global scopes.

Can policy-based governance cover AI tools that never call the gateway?

Yes, through endpoint enforcement. Bifrost Edge, currently in alpha, runs on company machines and routes traffic from supported desktop chat apps, browser AI, coding agents, and their MCP servers through the organization's Bifrost gateway. The virtual keys, budgets, guardrails, and logging configured at the gateway then apply to that endpoint traffic.

Getting Started with Policy-Based AI Governance

Policy-based AI governance holds only when the policy engine sits where the traffic is. Bifrost puts access, spend, routing, content safety, tool scope, and audit evidence in one control plane that evaluates every AI call, and extends that plane to MCP servers and endpoint AI so the rules cover the tools people actually use.

A practical sequence is to issue virtual keys and connect the identity provider first, add budgets and rate limits, layer guardrail and routing rules in CEL, then extend tool scope and endpoint coverage. For the organizational side of that rollout, see turning enterprise AI governance policy into gateway controls and the policy-to-runtime governance guide.

To see how policy-based AI governance maps to your providers, teams, and compliance requirements, book a demo with the Bifrost team.