Try Bifrost Enterprise free for 14 days. Request access

Why MCP Needs a Governance Layer: Access Control, Audit, and Cost

MCP adoption has outpaced the infrastructure needed to govern it. An MCP governance layer decides which tools each agent can call, who called them, what was returned, and what it cost. This guide covers tool-level access control, the two log types compliance needs, and where token cost actually goes

Why MCP Needs a Governance Layer: Access Control, Audit, and Cost

An MCP governance layer gives enterprise AI teams tool-level access control, audit trails, and cost visibility across every connected Model Context Protocol server.


TL;DR

  • An MCP governance layer is the control plane that decides which tools each agent can call, who called them, what was returned, and what it cost. Without one, each MCP server becomes its own island of policy, logging, and spend.
  • Ungoverned MCP maps onto three named OWASP LLM Top 10 risks at once: excessive agency, prompt injection through tool descriptions, and unbounded consumption.
  • Access control has to be enforced per tool, not per server. A single server routinely exposes a read tool and a delete tool side by side.
  • Tool executions belong in request logs with full arguments and latency; administrative changes belong in a separate signed audit log. Compliance programs need both, and they answer different questions.
  • Token cost and tool cost are one problem. Bifrost is open source under Apache 2.0 and reports both in the same view, with a Code Mode execution path that cut input tokens 92.8% at 508 tools.

Model Context Protocol (MCP) adoption has moved faster than the infrastructure needed to govern it. Teams connect an MCP server for file access, another for search, a third for internal APIs, then ten more, and within weeks an AI agent has a reach that no engineer would be handed on day one. An MCP governance layer is the missing control plane that decides which tools each agent can call, who is calling them, what each call returns, and what it costs. Bifrost, the open-source AI gateway by Maxim AI, provides this layer through a single MCP gateway that sits between your models and every connected MCP server.

The stakes are not hypothetical. The 2025 OWASP Top 10 for LLM Applications lists Excessive Agency as a top production risk, citing excessive functionality, excessive permissions, and excessive autonomy as the three root causes. MCP amplifies each of them.


What is an MCP Governance Layer

An MCP governance layer is the infrastructure between AI agents and the MCP servers they use. The governance layer enforces per-tool access control, records every tool invocation as a traceable event, and tracks both token and tool-level cost across all connected servers. It replaces scattered, per-server security with a single policy and observability plane, so teams can scale from one MCP server to dozens without losing control of any of it.

In practice the layer is an MCP gateway with policy attached. The gateway supplies the position in the network: it terminates the protocol on both sides, so it is the only component every tool call passes through. Governance is what that position is used for. A gateway without policy is a connection aggregator; policy without a gateway has nowhere to be enforced.

Three questions separate a governance layer from ordinary MCP plumbing:

QuestionWithout a governance layerWith one
Which tools can this agent call?Whatever its configured servers exposeAn explicit per-tool allow-list per consumer
What did the agent actually do?Scattered across per-server logs, if recordedOne correlated record per call, with attribution
What did the run cost?Model tokens onlyModel tokens plus per-tool spend, in one view

The distinction matters because the answers are what auditors, finance, and incident responders each ask for, and none of them can be reconstructed after the fact from per-server logs. What MCP governance means in practice covers the policy model in more depth.Why Ungoverned MCP Deployments Break at Scale

As soon as MCP moves from a local developer setup to a shared production environment, three structural problems surface.

  • Excessive agency. Agents are handed more tools than they need, often with broader permissions than the task requires. This is the exact failure mode OWASP catalogs under LLM06.
  • Tool poisoning and indirect prompt injection. Malicious or compromised MCP servers can embed hidden instructions inside tool descriptions, which the model reads and treats as authoritative. Microsoft's developer team has documented how tool poisoning works in MCP and why client-side validation alone is insufficient.
  • Unbounded token consumption. Every MCP tool definition from every connected server gets injected into the model's context on every single request. A 2025 Anthropic engineering post on code execution with MCP shows one Google Drive to Salesforce workflow dropping from 150,000 tokens to 2,000 tokens once tool definitions stopped being loaded on every turn.

Without a governance layer, none of these problems have a central place to be solved. Each MCP server becomes its own island of access policy, logging, and cost.


Why Ungoverned MCP Deployments Break at Scale

Ungoverned MCP deployments break at scale because the failure modes are structural rather than operational: each connected server multiplies permissions, context, and logging surface independently, and nothing in the protocol coordinates them. As soon as MCP moves from a local developer setup to a shared production environment, three structural problems surface.

  • Excessive agency. Agents are handed more tools than they need, often with broader permissions than the task requires. This is the exact failure mode OWASP catalogs under LLM06.
  • Tool poisoning and indirect prompt injection. Malicious or compromised MCP servers can embed hidden instructions inside tool descriptions, which the model reads and treats as authoritative. Microsoft's developer team has documented how tool poisoning works in MCP and why client-side validation alone is insufficient.
  • Unbounded token consumption. Every MCP tool definition from every connected server is injected into the model's context on each request. A 2025 Anthropic engineering post on code execution with MCP shows one Google Drive to Salesforce workflow dropping from 150,000 tokens to 2,000 tokens once tool definitions stopped being loaded on every turn.

Without a governance layer, none of these problems have a central place to be solved. Each MCP server becomes its own island of access policy, logging, and cost. The inventory problem this creates has a name: shadow MCP, where ungoverned servers run outside any register.


How MCP Governance Maps to the OWASP LLM Top 10

Ungoverned MCP is not one risk. It is three entries from the OWASP Top 10 for LLM Applications arriving together, which is why point fixes rarely hold. Mapping the controls to the named risks is also how most security reviews of an agent platform are actually structured.

OWASP LLM riskHow ungoverned MCP triggers itControl at the governance layer
LLM01 Prompt InjectionTool descriptions are read by the model as authoritative instructionsGuardrail evaluation of tool arguments and results, not only prompts
LLM02 Sensitive Information DisclosureTool results flow back into context with no redaction stepPII detection and secrets scanning between the tool and the model
LLM06 Excessive AgencyAgents hold more tools, and broader permissions, than any task requiresPer-tool allow-lists, deny-by-default, scoped credentials
LLM08 Vector and Embedding WeaknessesRetrieval tools reach corpora the caller should not seeTool-level authorization applied to retrieval servers
LLM10 Unbounded ConsumptionEvery tool definition is injected on every turn, with no spend ceilingBudgets and rate limits per key, plus a token-reduction execution path

The three OWASP calls out as root causes of excessive agency map directly onto what a gateway can enforce: excessive functionality is solved by the tool allow-list, excessive permissions by scoped credentials, and excessive autonomy by requiring explicit invocation rather than auto-execution. Bifrost does not auto-execute tool calls returned by a model; execution requires an explicit call, which keeps a human or an application in the loop for sensitive operations.

Guardrails are the enforcement half of this. Bifrost ships three managed checks, Prompt Guardrails, Custom Regex, and Secrets Detection, alongside eleven external providers including Presidio, Azure AI Language PII, AWS Bedrock Guardrails, Google Model Armor, and Patronus AI. Rules are written in CEL and evaluate MCP tool arguments and results, not only model prompts and responses, which is the part most guardrail deployments miss. The broader control set is enumerated in MCP security best practices.


Access Control: Scoping What Agents Can Actually Do

Access control for MCP has to operate at the tool level, not the server level. A single MCP server can expose filesystem_read alongside filesystem_write, or crm_lookup_customer alongside crm_delete_customer. Server-level allowlists treat these as a single unit, which defeats least privilege from the start.

Bifrost handles this through virtual keys and Virtual MCPs.

  • Virtual keys are scoped credentials issued to each consumer of the gateway: a user, a team, an internal application, or a customer integration. Each key carries an explicit list of the MCP tools it is allowed to call. The model attached to a key never sees definitions for tools outside that scope, so there is no prompt-level workaround.
  • Virtual MCPs are curated, addressable bundles of MCP tools served at their own /mcp/<slug> endpoint and attached to virtual keys. Creating them, selecting their tools, and assigning them to keys is part of the open-source build rather than an enterprise add-on.
  • Access profiles grant one or more Virtual MCPs to a role. Every user in that role inherits each granted bundle and all of its tools without a direct attachment, which is what makes the model administrable at team scale rather than per key.
  • Visibility scoping means an operator only sees a Virtual MCP if they hold an attached key or belong to an attached team or customer. Operators outside those relationships get a 404 rather than a permission error, so the existence of a bundle is itself scoped.

In a clustered deployment, definition changes reload from each node's own database copy, while attach and detach operations are carried with the affected virtual key so every node updates its in-memory assignment index.

This pattern aligns with where the broader MCP ecosystem is heading. The MCP specification has made authorization a first-class concern rather than an implementation detail, and identity providers are beginning to treat MCP as an authorization surface in its own right. Bifrost implements OAuth 2.0 for MCP servers with PKCE for public clients, Dynamic Client Registration under RFC 7591, and endpoint discovery under RFC 8414, alongside five other auth types covering shared headers, per-user credentials, and token exchange.

Per-user authentication is the part that decides whether user-scoped tools can be governed at all. A shared service credential makes every agent action attributable to the service rather than the person, which is exactly what an audit trail needs to avoid. The decision path across all six types is covered in MCP authentication patterns for OAuth, API keys, and token management, and the identity angle connects to the wider non-human identity problem, where machine principals now outnumber human ones in most enterprise estates. A governance layer is where those standards are enforced consistently, regardless of which upstream MCP servers support them natively.


Audit Logging: Making Agent Actions Traceable

The moment an AI agent can call production tools, every invocation has to be traceable. Two different records are needed, and conflating them is the most common gap in an MCP compliance story: one record of what the agent did, and one record of who changed what the agent was allowed to do.

Request logsAudit logs
RecordsLLM calls and MCP tool executionsAdministrative activity
FieldsInputs, outputs, tokens, cost, latency, tool namesInitiator, action, target resource, outcome
AnswersWhat did the agent doWho changed the configuration
Filter byLatency, token range, tool call name, virtual keyInitiator, target, IP, action, outcome, date
RetentionConfigurable, content logging separately disablableConfigurable, archival to S3 or GCS

An auditor asking whether an agent accessed customer records needs the first. An auditor asking who granted that agent access needs the second. Request logs capture each MCP tool execution with:

  • Tool name and the MCP server it came from
  • Arguments passed in and the result returned
  • Latency of the tool call
  • The virtual key that triggered the request
  • The parent LLM request that initiated the agent loop

Logging runs asynchronously and adds no latency to the request path. Audit logs are separate and record administrative activity as signed events: the initiator, the action taken, the target resource, and the outcome, with configurable retention and archival.

Teams can pull up any agent run and trace the exact sequence of tool calls, or filter by virtual key to audit what a specific team or customer has been running. Content logging can be disabled per environment when arguments or results carry sensitive data, while metadata (tool name, server, latency, status) is always captured.

This matters beyond debugging. Traceable records are required for SOC 2, HIPAA, GDPR, and ISO 27001 programs, and auditors increasingly expect them to cover AI tool invocations, not just API calls. Both record types support configurable retention and export to downstream SIEM and data lake tooling. Auditing every AI tool call at the gateway covers the review workflow in detail.


Cost Control: Token Bloat and Tool Call Economics

MCP cost has two components that a governance layer has to address together: the token cost of loading tool definitions, and the real-dollar cost of the tools themselves.

The token bloat problem

The default MCP execution model injects every tool definition from every connected server into the model's context on every single request. Five servers with thirty tools each means 150 tool definitions shipped before the user's prompt is even read. Industry research has converged on a fix: agents write code against the tool catalog instead of receiving the full catalog on every turn. Both Anthropic's engineering team and Cloudflare have published detailed analyses of the approach.

Bifrost implements this natively through Code Mode. Instead of dumping every tool definition into context, Code Mode exposes MCP servers as a virtual filesystem of lightweight Python stubs. The model reads only what it needs through four meta-tools (listToolFiles, readToolFile, getToolDocs, executeToolCode), and Bifrost executes the resulting script in a sandboxed Starlark interpreter. Bifrost's controlled MCP benchmarks show input token usage dropping 92.8% at 508 tools across 16 servers, with pass rate held at 100%.

The tool call economics problem

Not every MCP tool is free. Search APIs, enrichment vendors, code execution services, and paid data providers each carry a per-call price. Bifrost tracks cost at the tool level using a pricing configuration defined per MCP client, and surfaces those costs alongside LLM token costs in the same log view. Teams see the complete cost of an agent run, not just the model portion.


What an MCP Governance Layer Looks Like in Practice

A governance layer that scales has five non-negotiable properties:

  • A single endpoint. All connected MCP servers sit behind one /mcp URL that agents connect to. New servers appear without client-side reconfiguration.
  • Per-tool authorization. Access is scoped at the tool level, enforced by the gateway, and invisible to anything outside scope.
  • Unified audit. Every LLM call and every tool call land in one log model, correlated by request ID.
  • Cost visibility across both layers. Token spend and tool spend are reported together, broken down by virtual key, team, and provider.
  • Standards-based authentication. OAuth 2.0 with PKCE, identity-provider integration, and automatic token refresh, rather than static bearer tokens shared across services.

Bifrost centralizes all five inside its governance stack, which covers MCP traffic, model traffic, and the identity and budget primitives that unify them.


Where MCP Governance Sits Inside AI Agent Governance

AI agent governance is the wider programme: model selection, evaluation, human oversight, incident response, and the policy documents that describe them. MCP governance is one layer inside it, and specifically the enforcement layer, because it is the point where a written policy becomes something a system refuses to do.

The distinction is worth keeping because the two are often conflated in procurement. A governance platform documents that an agent should not reach customer records. A gateway is what returns an error when it tries. Most agentic AI governance programmes have more of the first than the second, which is how organizations end up with a complete policy binder and no ability to answer what an agent actually did last Tuesday.

LayerQuestion it answersWhere it lives
Policy and riskWhat should agents be allowed to doGovernance platform, risk register
EnforcementWhat can agents actually do right nowMCP gateway, virtual keys, guardrails
EvidenceWhat did agents do, and who authorized itRequest logs and audit logs
EvaluationAre agents doing it wellEvaluation and observability tooling

For teams building the enforcement layer, the practical starting point is a single gateway in front of every MCP server, with one virtual key per consumer and deny-by-default tool access. Everything else, including budgets, guardrails, and per-user auth, attaches to that structure afterwards rather than requiring it to be rebuilt. How MCP tools are discovered, invoked, and access-controlled covers what each enforcement point acts on, and the Bifrost governance stack documents how the primitives compose.


Getting Started with Bifrost MCP Gateway

MCP is now the default interface between AI agents and enterprise systems, and ungoverned deployments do not stay ungoverned quietly. Access drift, audit gaps, and runaway token bills all compound as the number of connected servers grows. An MCP governance layer is how teams move from early experimentation to production AI infrastructure without giving up control of what their agents can do, what those actions cost, or how those actions are recorded.

The adoption path is incremental rather than a migration. Point one agent surface at the gateway, issue it a virtual key with an explicit tool allow-list, confirm the request logs show what you expect, then move the remaining clients across one at a time. Nothing about the upstream MCP servers changes, which is what makes the first step cheap to reverse. Teams comparing options can start from the best MCP gateways for production AI systems.

To see how an MCP governance layer fits with your existing agents and servers, book a demo with the Bifrost team.


Frequently Asked Questions

What is an MCP governance layer?

An MCP governance layer is the infrastructure between AI agents and the MCP servers they call. It enforces per-tool access control, records each tool invocation with attribution, and tracks token and tool cost together. Without one, every MCP server carries its own access policy, logging, and spend, and none of it is coordinated.

Which OWASP LLM risks does an MCP governance layer address?

Primarily LLM06 Excessive Agency, by scoping tools per consumer, and LLM10 Unbounded Consumption, through budgets and a token-reduction execution path. It also mitigates LLM01 Prompt Injection and LLM02 Sensitive Information Disclosure when guardrails evaluate tool arguments and results rather than only model prompts, as set out in the enterprise MCP security checklist.

Should tool executions go in audit logs?

No. Tool executions belong in request logs, which capture arguments, results, latency, tokens, and cost. Audit logs should record administrative activity: who created a virtual key, who attached a tool bundle, who changed a budget. Gateway-level observability for AI tool calls walks through both record types. Compliance reviews need both, and each answers a question the other cannot.

How does an MCP governance layer control cost?

Through two mechanisms. Budgets and rate limits per virtual key cap spend before it happens. A token-reduction execution path addresses the larger cost driver, which is tool-schema injection: Code Mode cut input tokens 92.8% at 508 tools across 16 servers with pass rate held at 100%.