Try Bifrost Enterprise free for 14 days. Request access

How to Reduce MCP Token Costs for Claude Code at Scale

Claude Code defers MCP tool definitions by default, but not when ANTHROPIC_BASE_URL points to a gateway. This guide covers where MCP tokens go, what tool search misses, and how Bifrost Code Mode and per-key tool filtering cut input tokens by up to 92.8%.

How to Reduce MCP Token Costs for Claude Code at Scale

Reduce MCP token costs for Claude Code by up to 92% with Bifrost's MCP gateway, Code Mode execution, and centralized tool governance.

TL;DR

  • MCP token costs in Claude Code come from three places: tool definitions in context, tool results returned to the model, and the model round trips between multi-step tool calls.
  • Claude Code defers MCP tool definitions by default through tool search, but it loads them upfront when ANTHROPIC_BASE_URL points to a gateway or proxy, unless ENABLE_TOOL_SEARCH is set.
  • Measured without deferral, a typical four-server setup adds about 7,000 tokens of tool definitions per message, and heavy setups pass 50,000 tokens before the first prompt.
  • Bifrost's Code Mode cut input tokens by 58.2% at 96 tools, 84.5% at 251 tools, and 92.8% at 508 tools, and tool filtering per virtual key removes tools a developer should not see.
  • Claude Code connects to Bifrost with one claude mcp add command pointing at the /mcp endpoint, authenticated with a virtual key.

Connecting Claude Code to more than a handful of MCP servers almost always surfaces the same pattern: token usage climbs, response latency creeps up, and the API bill arrives larger than anyone expected. The root cause is not the tools themselves. It is how the Model Context Protocol (MCP) loads tool definitions into context on every request. To reduce MCP token costs for Claude Code without stripping out capability, teams need an infrastructure layer that governs tool exposure, caches what should be cached, and moves orchestration out of the prompt. Bifrost, the open-source AI gateway from Maxim AI, is built to do exactly that. This guide walks through where MCP token costs actually come from, what Claude Code's built-in optimizations can and cannot solve, and how Bifrost as an MCP gateway with Code Mode cut input tokens by up to 92.8% in benchmarks.


Why MCP Token Costs Explode in Claude Code

MCP token costs compound because tool schemas load into every single message, not once per session. Every MCP server Claude Code connects to injects its full tool definitions, names, descriptions, parameter schemas, expected outputs, into the model's context on every turn. Connect five servers with thirty tools each and the model is parsing 150 tool definitions before it ever sees the user's prompt.

Independent reporting has quantified the problem precisely. A recent analysis found that a typical four-server MCP setup in Claude Code adds around 7,000 tokens of overhead per message, with heavier setups crossing 50,000 tokens before a single prompt is typed. Another teardown showed multi-server configurations commonly adding 15,000 to 20,000 tokens of overhead per turn on usage-based billing.

Three dynamics make the problem worse at scale:

  • Per-message loading: Tool definitions reload on every turn, so a 50-message session pays the overhead 50 times.
  • Unused tools still cost: A Playwright server's 22 browser tools ride along even when the task is editing a Python file.
  • Verbose descriptions: Open-source MCP servers often ship with long, human-readable tool descriptions that inflate per-tool token cost.
  • Large tool results: Claude Code warns when a single MCP tool result exceeds 10,000 tokens and caps results at 25,000 tokens by default (MAX_MCP_OUTPUT_TOKENS), so one verbose tool call can cost more than every tool definition combined.

Token overhead is not just a line item. It crowds out the working context the model actually needs, which degrades output quality in long sessions and drives premature compaction. The same mechanics apply to any agent, as covered in code execution with MCP.


What Claude Code's Built-In Optimizations Cover

Anthropic has shipped several optimizations that address the easy cases. Understanding what they cover clarifies where an external layer is still needed.

Claude Code's official cost management guidance recommends a combination of tool search deferral, prompt caching, auto-compaction, model tiering, and custom hooks. Tool search is the most relevant for MCP: by default, Claude Code defers MCP tool definitions so only tool names and server instructions enter context until Claude uses a specific tool.

Claude Code leverWhat it reducesScope
Tool search (on by default)Tool definitions in contextOne developer's session
/mcp to disable unused serversTool definitions and connectionsOne developer's session
MAX_MCP_OUTPUT_TOKENSOversized tool results (25,000-token default cap)One developer's machine
Hooks that trim command outputTool and command resultsOne developer's configuration
/context and /usageNothing directly; shows what consumes context and usage per MCP serverLocal session history only

The gateway catch: tool search and ANTHROPIC_BASE_URL

Claude Code turns tool search off by default when ANTHROPIC_BASE_URL points to a non-first-party host, because most proxies do not forward the tool_reference blocks tool search depends on, according to the Claude Code MCP documentation. Teams that route Claude Code's model traffic through an LLM gateway therefore get every MCP tool definition loaded upfront again. Setting ENABLE_TOOL_SEARCH=true restores deferral only if the gateway forwards those blocks; otherwise requests fail. Using Bifrost only as the MCP endpoint, without changing ANTHROPIC_BASE_URL, leaves tool search on.

This is where reductions at the gateway matter most: tool filtering per virtual key and Code Mode shrink what reaches the model regardless of whether Claude Code defers definitions.

These built-in controls help, but they leave three gaps for teams running MCP in production:

  • No centralized governance: Tool deferral is a client-side optimization. It does not give a platform team control over which tools a given developer, team, or customer integration is allowed to call.
  • No orchestration layer: Even with deferral, every multi-step tool workflow pays for tool schemas it loads, intermediate tool results, and model round-trips on every step.
  • No cross-session visibility: Individual developers can run /context and /usage to audit their own sessions, but those figures come from local session history, so there is no organizational view of which MCP servers are burning tokens across the team.

For a single developer running Claude Code locally with two or three servers, the built-in optimizations are enough. For a platform team rolling Claude Code out to dozens or hundreds of engineers with shared MCP infrastructure, they are not, which is the gap an MCP gateway for production AI agents fills.


How Bifrost Reduces MCP Token Costs for Claude Code

Bifrost sits between Claude Code and the fleet of MCP servers your team depends on. Instead of Claude Code connecting directly to each server, it connects to one /mcp endpoint on Bifrost, the same pattern as connecting Claude Code to multiple MCP servers through one gateway. Bifrost handles discovery, tool governance, execution, and the orchestration pattern with the largest effect on token cost: Code Mode.

The result is documented in the Bifrost MCP gateway cost benchmark, which shows input tokens dropping by 58.2% with 96 tools connected, 84.5% with 251 tools, and 92.8% with 508 tools, all while pass rate held at 100%. The round-by-round data is in cutting MCP token costs by 92% at 500 tools.

Code Mode: orchestration that stops paying per-turn schema tax

Code Mode is the single largest driver of token reduction. Instead of injecting every MCP tool definition into context, Bifrost exposes connected MCP servers as a virtual filesystem of lightweight Python stub files. The model reads only what it needs, writes a short Python script to orchestrate the tools, and Bifrost executes that script in a sandboxed Starlark interpreter.

The model works with four meta-tools regardless of how many MCP servers are connected:

  • listToolFiles: Discover which servers and tools are available.
  • readToolFile: Load Python function signatures for a specific server or tool.
  • getToolDocs: Fetch detailed documentation for a specific tool before using it.
  • executeToolCode: Run the orchestration script against live tool bindings.

The pattern is conceptually similar to what Anthropic's engineering team described for code execution with MCP, where a Google Drive to Salesforce workflow dropped from 150,000 tokens to 2,000. Bifrost implements the same approach natively in the gateway, chooses Python over JavaScript for better LLM fluency, and adds the dedicated docs tool to further compress context. Cloudflare independently observed the same exponential savings pattern in their evaluation.

The savings compound as you add servers. Classic MCP pays for every tool definition on every request, so connecting more servers makes the tax worse. Code Mode's cost is bounded by what the model actually reads, not how many tools exist.

Virtual keys and Virtual MCPs: stop paying for access a consumer should not have

Every request through Bifrost carries a virtual key. Each key is scoped to a specific set of tools, and scoping works at the tool level, not just the server level. A key can be allowed to call filesystem_read without having access to filesystem_write from the same MCP server. The model only ever sees definitions for tools the key is allowed to call, so unauthorized tools cost zero tokens.

Filtering stacks at three levels: the client configuration, request headers that can narrow but never widen the tool list, and the virtual key, as documented in MCP tool filtering. At organizational scale, Virtual MCPs extend this further: a curated bundle of tools from several servers is served at its own /mcp/<slug> endpoint, reachable only through the virtual keys attached to it, so each team's Claude Code setup loads only its bundle.

Centralized gateway: one connection, one audit trail

Bifrost exposes all connected MCP servers through a single /mcp endpoint. Claude Code connects once and discovers every tool from every MCP server the virtual key allows. Add a new MCP server in Bifrost and it appears in Claude Code automatically with no client-side configuration change.

This matters for cost because it gives platform teams the visibility Claude Code's per-session tooling cannot. Tool executions appear as MCP entries in Bifrost's request logs next to the LLM calls, and header-based metadata such as a team or project ID can be attached to every LLM and MCP log entry, so usage can be grouped across every developer instead of one machine at a time.


Setting Up Bifrost as Your MCP Gateway for Claude Code

Setting up Bifrost as the MCP gateway for Claude Code takes four steps: register MCP servers in Bifrost, enable Code Mode on the heavy ones, scope virtual keys, and add Bifrost to Claude Code with one command. Bifrost runs as a drop-in replacement for existing SDKs, so no application code changes are required.

  1. Add MCP clients in Bifrost: Navigate to the MCP section in the Bifrost dashboard and register each MCP server you want to expose, with connection type (HTTP, SSE, or STDIO), endpoint, and any required headers.
  2. Enable Code Mode: Open the client settings and toggle Code Mode on. No schema changes, no redeployment. Token usage drops immediately as the four meta-tools replace full schema injection.
  3. Configure auto-execute and virtual keys: Under virtual keys, create scoped credentials for each consumer and select which tools each key is allowed to call. For autonomous agent loops, allowlist read-only tools for auto-execution while keeping write operations behind approval.
  4. Point Claude Code at Bifrost: Add Bifrost as an HTTP MCP server with the claude mcp add command below. Claude Code discovers every tool the virtual key allows through a single connection.
claude mcp add --transport http bifrost http://localhost:8080/mcp \
  --header "Authorization: Bearer your-virtual-key"

Bifrost also accepts the virtual key as X-Api-Key or x-bf-vk if the Authorization header conflicts with another tool. To route Claude Code's model traffic through Bifrost as well, set ANTHROPIC_BASE_URL to http://localhost:8080/anthropic and ANTHROPIC_AUTH_TOKEN to the virtual key, following the Claude Code integration docs; review the tool search note above before doing so. The broader walkthrough is in how to add and govern MCP servers in Claude Code.

From that point on, Claude Code sees a governed, token-efficient view of your MCP ecosystem, and every tool execution is recorded in the gateway's logs.


Measuring the Impact on Your Team

Reducing MCP token costs for Claude Code is only valuable if you can measure it. Bifrost's observability surfaces the data that matters for cost decisions:

  • Token cost per virtual key, per tool, and per MCP server over time.
  • Full trace of every agent run: which tools were called, in what order, with what arguments, and at what latency.
  • Spend breakdown combining LLM token costs and tool costs side by side, so you see the complete cost of every agent workflow.
  • Native Prometheus metrics and OpenTelemetry (OTLP) integration for Grafana, New Relic, Honeycomb, and Datadog.

A practical test is to record input tokens per session for a week with classic MCP, enable Code Mode on the two or three largest servers, and compare. Attributing spend to individual tools is covered in tracking per-tool costs across MCP servers.

Teams evaluating the cost impact at their own scale can cross-reference Bifrost's published performance benchmarks, which show 11 microseconds of overhead at 5,000 requests per second, and the LLM Gateway Buyer's Guide for a full capability comparison.


Beyond Token Costs: The Production MCP Stack

MCP without governance and cost control becomes unsustainable as soon as you move past a single developer's local setup. Bifrost's MCP gateway addresses the full set of production concerns in one layer:

  • Scoped access via virtual keys and per-tool filtering.
  • Organizational governance with MCP Tool Groups.
  • Complete audit trails for every tool call, suitable for SOC 2, GDPR, HIPAA, and ISO 27001.
  • Per-tool cost visibility alongside LLM token usage.
  • Code Mode to cut context cost without cutting capability.
  • The same gateway that governs MCP traffic also handles LLM provider routing, automatic failover, load balancing, semantic caching, and unified key management across 20+ AI providers.

When LLM calls and tool calls flow through the same gateway, model tokens and tool costs sit in one audit log under one access control model. That is the infrastructure pattern production AI systems actually need. Teams already using Claude Code with Bifrost can review the Claude Code integration guide for implementation details specific to that workflow.


Start Reducing MCP Token Costs for Claude Code

Reducing MCP token costs for Claude Code is not about trimming tools or accepting smaller capability. It is about moving tool governance and orchestration into the infrastructure layer where they belong. Bifrost's MCP gateway and Code Mode deliver token reductions of up to 92% on large tool catalogs while tightening access control and giving platform teams the cost visibility they need to operate Claude Code at scale.

To see how Bifrost can cut your team's Claude Code token bill and give you production-grade MCP governance, book a demo with the Bifrost team.


Frequently Asked Questions

Why does Claude Code use so many tokens with MCP servers?

Claude Code uses more tokens with MCP servers because tool definitions, tool results, and extra model round trips all enter the context window. Tool search defers definitions by default, but large tool results still load in full, up to a 25,000-token cap per call, and multi-step workflows resend the growing conversation on every step.

Does Claude Code tool search solve MCP token costs?

Tool search solves most of the tool-definition cost for one developer, because only tool names load until a tool is used. It does not reduce tool results or round trips, and Claude Code turns it off by default when ANTHROPIC_BASE_URL points to a gateway or proxy, loading every tool definition upfront again.

How do I check MCP token usage in Claude Code?

Run /context to see what fills the context window, including MCP tools, and /usage to see session token counts. On Pro, Max, Team, and Enterprise plans, /usage also attributes recent usage to individual MCP servers, based on local session history. Team-wide totals require a gateway that logs every developer's traffic, one of the roles covered in what an MCP gateway is.

How much does Code Mode reduce Claude Code token costs?

In Bifrost's benchmark, Code Mode reduced input tokens by 58.2% with 96 tools, 84.5% with 251 tools, and 92.8% with 508 tools, keeping a 100% pass rate. Savings grow with the number of connected tools, so teams with a few small servers see less benefit than teams with large shared tool catalogs.

How do I connect Claude Code to Bifrost's MCP endpoint?

Run claude mcp add --transport http bifrost http://localhost:8080/mcp with an Authorization: Bearer header carrying a Bifrost virtual key. Claude Code then sees every tool that key allows through one connection. The same virtual key can route model traffic through Bifrost via ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN.

Should I use CLI tools instead of MCP servers to save tokens?

For tools with a mature command-line interface, such as gh or cloud CLIs, calling the CLI through Bash avoids MCP tool definitions entirely. MCP servers remain the better choice for services without a CLI, for per-user authentication, and when a team needs central governance over which tools each developer can use.