Try Bifrost Enterprise free for 14 days. Request access

AI Governance with Virtual Keys for LLM and MCP Traffic

AI Governance with Virtual Keys for LLM and MCP Traffic

TL;DR

  • AI governance with virtual keys routes every LLM and MCP request through gateway-issued credentials that carry their own model access, budgets, rate limits, and tool permissions.
  • A Bifrost virtual key with no provider configuration blocks all providers, and a key with no MCP configuration exposes no tools except from clients marked Allow by Default.
  • Budgets are checked at every level of the virtual key, team, and customer hierarchy, and any single failing budget blocks the request.
  • Inactive or expired virtual keys are rejected at MCP tool execution time with a 403, cutting off model and tool access together.
  • Bifrost Enterprise adds access profiles, RBAC, OIDC and SCIM provisioning, Virtual MCPs, and audit logs for governance at scale.

IBM's 2025 Cost of a Data Breach Report found that 13% of organizations reported breaches of AI models or applications, and 97% of those lacked proper AI access controls. For enterprise teams running LLM and Model Context Protocol (MCP) traffic across multiple providers, AI governance with virtual keys replaces scattered provider API keys and informal spending limits with a single control point that engineering, finance, and security can all rely on. Bifrost, the open-source AI gateway built in Go by Maxim AI, makes virtual keys the primary governance entity: every request, budget, and tool permission is checked against the key that issued it. This post explains how virtual keys govern both LLM and MCP traffic, and how that governance scales across teams, budgets, and compliance requirements.

For the wider program these controls sit inside, see the complete guide to AI governance for enterprise LLM deployments.

What Is AI Governance with Virtual Keys?

AI governance with virtual keys is the practice of routing every LLM and MCP request through gateway-issued credentials that carry their own access permissions, budgets, and rate limits. Instead of sharing raw provider API keys, teams authenticate with virtual keys, so a central gateway enforces cost, access, and tool policies on every request.

A virtual key sits in front of your real provider API keys, and applications authenticate with it rather than a raw provider key. This lets Bifrost enforce policy on every request without exposing the underlying credentials to any team or service.

In Bifrost, virtual keys are the primary entity through which all governance is applied. Each key defines:

  • Access control: which providers and models the key is allowed to call, enforced on a deny-by-default basis.
  • Cost management: an independent budget with a configurable reset duration.
  • Rate limiting: token-based and request-based throttling over a defined period.
  • Key restrictions: an optional allow-list limiting the key to specific provider API keys.
  • Active or inactive status: the ability to enable or disable access instantly.

Keys authenticate through standard headers, including the OpenAI-style Authorization: Bearer, the Anthropic-style x-api-key, the Gemini-style x-goog-api-key, and the Bifrost-native x-bf-vk, so existing SDKs work without code changes beyond the base URL. This is what makes governance practical: policy lives at the gateway, not scattered across application code. The five ways to govern LLM access with virtual keys show these controls applied to common team setups.

Governance control Virtual key setting Default when unset
Provider and model access Provider configs with allowed_models All providers blocked
Provider API keys key_ids per provider config All keys denied unless ["*"]
Spend Budget with reset duration No budget cap
Throughput Request and token rate limits No rate limit
MCP tools MCP client configs with tool allow-lists No tools, except Allow by Default clients
Status Active or inactive flag Inactive keys are rejected

Why Enterprise Teams Need Centralized LLM Governance

Ungoverned AI usage is now a measurable risk rather than a hypothetical one. The OWASP Top 10 for LLM Applications ranks prompt injection and sensitive information disclosure among the most severe risks facing production LLM systems, and both are amplified when traffic flows through provider keys that no central system observes. The NIST AI Risk Management Framework sets the expectation that organizations can identify, monitor, and control every AI system they run, which is difficult when each team holds its own raw API keys.

Centralized governance through Bifrost addresses several enterprise problems at once:

  • Cost sprawl: without per-team budgets, a single misconfigured retry loop or runaway agent can produce a surprise invoice at the end of the month.
  • Access ambiguity: raw provider keys grant access to every model the account supports, with no way to restrict a team to approved models.
  • Audit gaps: when credentials are shared across services, there is no reliable record of which team or application generated a given request.
  • Tool exposure: agentic workflows connected to MCP servers can reach internal systems that were never meant to be available to every model.

Treating centralized AI governance as infrastructure, rather than a set of manual policies, is what lets enterprise teams expand AI adoption without losing control of spend, access, or data.

Governing LLM Traffic with Virtual Keys

For LLM traffic, Bifrost enforces governance at three layers on every request: access, budget, and rate limits. Because all three are attached to the virtual key, they apply uniformly whether the caller is a production service, a notebook, or a coding agent.

Access and routing. Through governance routing, a virtual key can be restricted to specific providers and models. With no provider configuration, the key blocks all traffic by default; once providers are added, the key is limited to exactly those provider and model combinations. The same configuration supports weighted load balancing across providers and automatic fallbacks, so teams can separate development, testing, and production environments while routing each to the models it is permitted to use.

Budgets and rate limits. Every key can carry an independent budget with a reset duration such as one day, week, month, quarter, or year, plus optional calendar-aligned resets that fire at the start of each UTC period. Bifrost calculates cost per request from token usage and model pricing that the Model Catalog synchronizes every 24 hours by default, then checks the request against the budget and rate limits before it proceeds. Rate limits operate in parallel as request-per-period and token-per-period thresholds, which protects providers from abuse and keeps a single team from exhausting shared capacity. Rate limits are checked at the provider-config and virtual key levels; when a provider config exceeds its limit, that provider is excluded from routing while other providers on the same key stay available.

Hierarchical cost control. Budgets are not limited to individual keys. A virtual key can belong to a team, and a team to a customer, with each level holding its own independent budget. When a request arrives, Bifrost checks every applicable budget in the hierarchy, and all of them must pass for the request to succeed. This lets a platform team set an organization-wide ceiling while each product team manages its own allocation underneath it. Teams comparing approaches to spend control can review the enterprise gateways for LLM cost tracking and budget controls.

Hierarchy level Budget Rate limits
Provider config on a virtual key Independent budget Request and token limits
Virtual key Independent budget Request and token limits
Team Independent budget None
Customer Independent budget None

Governing MCP Traffic with Virtual Keys

Agentic workflows introduce a second class of traffic that traditional API-key management ignores entirely: tool calls to MCP servers. When a model can create tickets, query databases, or trigger internal APIs through the MCP gateway, controlling which tools each key can reach becomes as important as controlling which models it can call.

MCP tool filtering applies the same virtual-key model to tool access. The behavior is deny-by-default: a key with no MCP configuration exposes no tools, except from MCP clients an admin marks as Allow by Default. Once you configure MCP clients on a key, Bifrost builds a strict allow-list and enforces it at two points, when the tool list is presented to the model and again when a tool is actually executed. For each connected client you can:

  • Permit specific named tools only, blocking everything else from that client.
  • Use a wildcard to allow all current and future tools from a client.
  • Leave a client's tool list empty, or omit the client entirely, to block it completely (an omitted client marked Allow by Default stays permitted).

Inactive or expired keys are rejected at tool execution time with a 403, so revoking a key immediately cuts off both model access and tool access. This closes a gap that pure network controls cannot reach, since the request is inspected at the gateway where both the model call and the tool call are visible. Teams that route agent traffic this way also benefit from Code Mode and other MCP optimizations covered in the MCP gateway deep dive on access control and token costs.

For allow-list design patterns, see MCP tool governance: filtering, allowlisting, and access control and tool-level permissions for production AI agents.

Scaling Governance Across Enterprise Teams

Individual virtual keys govern individual applications. At enterprise scale, the challenge shifts to issuing, updating, and auditing thousands of keys without manual effort. Bifrost Enterprise extends the same virtual-key foundation with capabilities designed for large organizations and regulated environments.

  • Access profiles. Rather than hand-writing keys, teams define an access profile once, as a reusable policy template carrying a provider list, model whitelist, budgets, rate limits, and MCP tool access. Assigning the profile to a user or role automatically issues a per-user virtual key with isolated budget and rate-limit counters, and those keys are write-protected so users cannot weaken their own policy.
  • Role-based access control. RBAC governs who can view or change gateway resources. Three system roles (Admin, Developer, and Viewer) cover common patterns, their permissions can be customized, and custom roles can be created for compliance, QA, or contractor access under the principle of least privilege.
  • Identity and team sync. Through OIDC and SCIM provisioning, users sign in with corporate credentials from identity providers such as Okta and Microsoft Entra, inherit roles from identity-provider groups, and are reconciled automatically as they join or leave teams, removing manual account creation.
  • Virtual MCPs. For agentic governance at scale, Virtual MCPs (previously called MCP tool groups) bundle curated tool subsets, serve each bundle at its own /mcp/<slug> path, and attach it to virtual keys. Enterprise adds access-profile grants and data access control on top.

For compliance, audit logs record administrative activity as events that can be HMAC-signed, retained, filtered in a dashboard, and exported or archived to object storage, with audit trails designed to be SOC 2, GDPR, HIPAA, and ISO 27001 friendly. Combined with in-VPC and on-prem deployment options, these capabilities make Bifrost Enterprise suitable for teams that need full control over data, access, and execution. Together, governance as infrastructure turns AI adoption from an audit liability into a defensible, documented control plane.

Frequently Asked Questions

What is a virtual key in an AI gateway?

A virtual key is a gateway-issued credential that applications use instead of raw provider API keys. In Bifrost, each virtual key carries its own allowed providers and models, budget, rate limits, MCP tool permissions, and active status. The gateway checks every request against the key, so the underlying provider credentials are never shared with teams or services.

How do virtual keys enforce AI budgets?

Bifrost calculates the cost of each request from token usage and synchronized model pricing, then checks it against every applicable budget: provider config, virtual key, team, and customer. Each budget resets on its own duration, from one minute to one year, and any single exhausted budget blocks the request.

Can virtual keys control which MCP tools an agent can use?

Yes. MCP tool filtering on a Bifrost virtual key is deny-by-default and builds a strict allow-list per MCP client, with specific tools, a wildcard, or nothing. The allow-list is applied when tools are presented to the model and enforced again at execution time, and inactive or expired keys are rejected with a 403.

Do existing SDKs work with Bifrost virtual keys?

Yes. Bifrost accepts virtual keys through the headers existing SDKs already send: Authorization: Bearer for OpenAI-style clients, x-api-key for Anthropic-style clients, and x-goog-api-key for Gemini-style clients, plus the native x-bf-vk header. Adoption usually means changing the base URL and swapping in a virtual key.

How does AI governance with virtual keys scale to thousands of users?

Bifrost Enterprise uses access profiles to issue per-user virtual keys automatically from reusable policy templates, with isolated budget and rate-limit counters. OIDC and SCIM provisioning assign roles and profiles from identity-provider groups, RBAC limits who can change gateway resources, and audit logs record administrative activity.

Getting Started with AI Governance in Bifrost

AI governance with virtual keys gives enterprise teams a single, enforceable control point for both LLM and MCP traffic: access is scoped per key, spend is capped per team and per key, tool permissions are deny-by-default, and administrative activity is logged. Teams building out the full program can pair this with policy-to-gateway controls for the enterprise AI governance lifecycle and the broader AI governance guide for enterprise LLM deployments.

Because Bifrost is a drop-in replacement for existing SDKs, adopting this model usually means changing a base URL and issuing keys, not rewriting application code.

To see how virtual keys can centralize AI governance across your models, budgets, and MCP servers, book a demo with the Bifrost team and map the policy model to your own stack.