Try Bifrost Enterprise free for 14 days. Request access

Top 5 AI Gateways with Prompt Caching Support in 2026

Compare the top 5 AI gateways with prompt caching support in 2026 on cache breakpoint injection, session affinity, cached-token cost tracking, and deployment.

Top 5 AI Gateways with Prompt Caching Support in 2026

TL;DR

  • Prompt caching is a provider feature: the provider reuses a cached request prefix and bills it at a reduced input rate, while semantic caching is a gateway feature that replays a stored response without calling the provider.
  • An AI gateway affects prompt caching in four ways: whether it forwards cache markers, whether it can inject them for clients that send none, whether it keeps a session on the provider key that holds the cache, and whether it prices cached tokens correctly.
  • Bifrost auto-injects cache breakpoints per provider, translates them to each provider's wire format, pins sessions to the provider and key that served them, and prices cache reads and writes separately.
  • LiteLLM and OpenRouter support provider prompt caching with sticky routing; Cloudflare AI Gateway adds cache-token cost rates; Kong documents semantic response caching and native-format passthrough only.

Prompt caching is a provider-side optimization that stores the processed prefix of an LLM request so later requests repeating that exact prefix are billed at a reduced input rate and return faster. Bifrost, the open-source AI gateway built in Go by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability, and it handles the prompt caching workflow end to end, from breakpoint injection to cached-token cost tracking. This guide compares five AI gateways on how much of that workflow each one handles.

What Is Prompt Caching?

Prompt caching is a feature of LLM providers that reuses the computed state of a request prefix (system prompt, tool definitions, earlier turns) when a later request repeats it byte for byte. The provider still runs the model and bills the call, but cached input tokens cost much less than fresh input, so agent loops that replay the same prefix every turn benefit most.

The mechanics differ by provider. Anthropic's prompt caching documentation prices cache reads at 0.1x the base input price on most Claude models and allows up to four explicit breakpoints per request. OpenAI's prompt caching guide describes automatic caching with reused tokens discounted by up to 95%.

Provider How caching is triggered Marker on the request Published pricing signal
Anthropic (Claude) Explicit breakpoints, or a single top-level cache_control field cache_control: {"type": "ephemeral"}, at most 4 breakpoints Reads 0.1x input on most models; writes 1.25x (5 min) or 2x (1 hour)
OpenAI Automatic (implicit); explicit mode on GPT-5.6 and later, with a 1,024-token minimum prefix prompt_cache_breakpoint plus prompt_cache_options.mode: "explicit" GPT-5.6+: writes 1.25x, reads 0.1x on most models
AWS Bedrock Explicit cache points for Claude and Amazon Nova on the Converse API cachePoint block Varies by model
Google Gemini Implicit prefix caching, or a server-side cached content resource No inline marker Varies by model
Turn 1 misses the provider prompt cache and writes the prefix at 1.25x input price; turn 2 repeats the prefix and reads it at 0.1x

Figure 1: The first turn pays a premium to write the prefix; every later turn that repeats it byte for byte pays a fraction of the input price.

As Figure 1 shows, prompt caching pays off only when a prefix is read more often than it is written, and one changed byte before the breakpoint (a timestamp, reordered tools) turns every turn into a write. Both problems sit in the request path, where Bifrost reads cache counters from supported providers that report cache usage through one usage block.

Prompt Caching vs Semantic Caching

Prompt caching and semantic caching solve different problems at different layers. Semantic caching runs inside the gateway and replays a stored response for an identical or similar request, so the provider is never called. Prompt caching runs at the provider, which still generates a fresh completion but bills the repeated prefix at a lower rate. The two are independent and can run together.

Attribute Provider prompt caching Gateway semantic caching
Where it runs LLM provider AI gateway
Provider called on a hit Yes, a new completion is generated No, a stored response is returned
What is reused Processed prefix of the prompt The full previous response
Match rule Exact prefix up to a breakpoint Exact hash or embedding similarity
Best fit Agent loops, long system prompts, RAG context Repeated FAQ-style questions
Billing on a hit Reduced input rate, normal output rate No provider charge
Top lane: the gateway response cache returns a stored answer without calling the provider. Bottom lane: the provider reuses a cached prompt prefix

Figure 2: A semantic cache hit removes the provider call entirely, while a prompt cache hit still generates a new completion at a lower input rate.

Many gateways advertise "caching" and mean response caching only. Bifrost ships the two as separate features: Auto Prompt Caching for provider prefixes and semantic caching with direct and similarity lookup for full responses.

For the response side, see this technical deep dive on semantic caching and the roundup of AI gateways with semantic caching for LLM cost reduction.

What an AI Gateway Needs for Prompt Caching

An AI gateway supports prompt caching well when it forwards provider cache markers intact, can add markers for clients that send none, routes each session back to the provider key that holds its cache, and reports cache reads and writes as separate, correctly priced token counts. Response caching is a separate, complementary capability.

Criterion Why it matters
Marker passthrough and translation Normalizing requests to one schema can strip cache_control or fail to convert it per provider
Automatic breakpoint injection Coding agents and many SDK callers send no markers, so nothing is cached on providers that require them
Cache-aware routing (session affinity) Provider caches are scoped to a provider and usually an API key, so load balancing each turn independently causes misses
Cached-token cost tracking Cache reads, writes, and fresh input carry different prices
Response caching Removes provider calls for repeated questions; independent of prompt caching
Turn 1 writes the cache on key A; without session affinity turn 2 lands on key B and misses, with affinity it returns to key A

Figure 3: Provider caches are scoped to a provider and usually an API key, so a gateway that load balances each turn independently pays the write rate repeatedly.

Figure 3 shows why weighted load balancing across two keys can double cache writes. These criteria sit alongside the usual gateway concerns covered in this production-ready comparison of the top LLM gateways and the LLM Gateway Buyer's Guide.

Top AI Gateways for Prompt Caching Compared

The five gateways below range from full prompt caching workflows to response caching only. Bifrost and LiteLLM inject breakpoints and pin sessions, OpenRouter translates markers and uses sticky routing, Cloudflare AI Gateway prices cache tokens, and Kong documents semantic response caching.

Gateway Provider cache markers Automatic breakpoint injection Cache-aware routing Cached-token cost tracking Response caching Deployment
Bifrost Translated per provider (cache_control, cachePoint, prompt_cache_breakpoint) Yes, per provider, with injection points and a per-request override Session affinity at provider and key level Separate cache-read and cache-write rates Exact-match and semantic Self-hosted, in-VPC, on-prem, air-gapped
LiteLLM Passthrough, plus prompt_cache_breakpoint for GPT-5.6+ Yes, injection points and a one-flag Claude mode Session affinity pre-call check Cost function handles cache pricing Exact-match and semantic Self-hosted proxy
OpenRouter cache_control and prompt_cache_breakpoint, converted between providers Not published Provider sticky routing Cached and cache-write tokens, cache_discount Exact-match Hosted service
Cloudflare AI Gateway Not published Not published Session affinity in Auto Router Custom cache-read and cache-write rates Exact-match Managed service
Kong AI Gateway Native-format passthrough (v3.10+) Not published Not published Not published Semantic (Enterprise tier) Self-hosted or Konnect

"Not published" means the capability did not appear in the vendor's documentation at the time of writing. Teams comparing access control as well can review the Bifrost governance model.

1. Bifrost

Bifrost, an open-source AI gateway, unifies 25+ providers and 10,000+ models behind one OpenAI-compatible API, and it treats provider prompt caching as a first-class routing concern. Bifrost injects cache breakpoints for clients that send none, keeps sessions on the provider key that holds their cache, and prices cached tokens separately in its cost tracking.

Coding agents and SDK apps send requests to Bifrost, which injects cache breakpoints, applies session affinity, routes to providers, and logs cached token costs

Figure 4: Breakpoint injection makes a request cacheable, session affinity sends it back to the cache it warmed, and cost tracking shows whether it worked.

Automatic breakpoint injection. Agentic clients such as Codex send no cache_control markers, so on Claude models nothing is cached. With automatic prompt cache breakpoints enabled through prompt_cache.auto_inject, Bifrost marks the first cacheable content block of each request that arrives without markers, so turn 1 writes the cache and later turns read it.

  • Off by default: a cache write costs more than fresh input, so injection is opt-in per provider.
  • Caller markers win: a request that already carries cache_control or prompt_cache_breakpoint is forwarded unchanged.
  • Capability gated per model: markers are only injected for models that accept them; implicit-caching providers such as Gemini, DeepSeek, Groq, and xAI receive nothing.
  • Four-marker ceiling: injection stops at four markers, matching Anthropic's limit.
  • Injection points: cache_control_injection_points target messages by role, index, or both.
  • Per-request override: the x-bf-prompt-cache-auto-inject header flips injection for one request, as listed in the request options reference.
  • Optional 1 hour TTL: "ttl": "1h" requests the longer cache lifetime.

One marker, translated per provider. Bifrost injects one internal marker and each provider translates it: cache_control for Claude on Anthropic, Vertex AI, and OpenRouter; a cachePoint block for Claude and Amazon Nova on AWS Bedrock; and prompt_cache_breakpoint with explicit mode for the OpenAI and Azure OpenAI gpt-5.6 family on the Responses API. Fallbacks re-evaluate injection against the new provider's configuration, so a failover from Anthropic to Bedrock receives a cachePoint when Bedrock has injection enabled. Tokens and cost for every request land in Bifrost request logs.

Session affinity. Session affinity keeps a session on the provider and key that last served it, after routing rules, virtual key load balancing, and the model catalog have built the candidate chain. Sessions come from the x-bf-session-id header, and Claude Code and Codex CLI get affinity with no configuration (Bifrost v2.0.0 and later) because Bifrost adopts the session headers they already send. Each decision is recorded in the routing trail; in Bifrost Enterprise, bindings replicate across a cluster and are dropped when the adaptive load balancer marks a provider as failed.

Cached-token cost tracking. Bifrost surfaces provider cache counters as cached_read_tokens and cached_write_tokens under usage.prompt_tokens_details. The model catalog applies separate rates to cache-read and cache-creation tokens, which keeps virtual key budgets accurate when most input is cached.

Semantic caching alongside it. Bifrost also runs a response cache with direct hash and embedding-based matching; a semantic hit skips the provider, and a miss still benefits from the provider's prefix cache.

Performance and deployment. Bifrost adds 11 microseconds of overhead per request at 5,000 RPS in sustained benchmarks. For regulated environments, Bifrost runs inside a private VPC or on-prem, with Bifrost Enterprise adding clustering and health-aware session invalidation.

Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

2. LiteLLM

LiteLLM is an open-source Python SDK and proxy that supports provider prompt caching for OpenAI, Anthropic, Google AI Studio, Vertex AI, Bedrock, DeepSeek, and xAI. It forwards cache_control markers, can auto-inject cache checkpoints, and offers session affinity as a router pre-call check.

  • Marker passthrough: cache_control blocks reach Anthropic, Bedrock, and Gemini, and prompt_cache_breakpoint is passed through for OpenAI GPT-5.6 and newer.
  • Auto-injection: cache_control_injection_points places checkpoints from the deployment config; an enable_anthropic_prompt_caching flag (v1.94.0+) adds a checkpoint on the system prompt and the trailing turn for Claude on Anthropic and Bedrock, respects the four-block limit, and never double-injects.
  • Session affinity: adding session_affinity to optional_pre_call_checks pins a conversation to the deployment that served its first request, reading a session header such as x-claude-code-session-id.
  • Cost tracking: responses follow the OpenAI usage format with cached_tokens and, for Anthropic, cache_creation_input_tokens; LiteLLM's cost function handles prompt caching rates.
  • Response caching: exact-match and semantic caches are separate; LiteLLM's docs warn that semantic caches behave poorly on agentic traffic.

Best for: Python-centric teams that want a self-hosted proxy with configurable cache injection. Teams evaluating throughput and governance at scale can compare LiteLLM alternatives.

3. OpenRouter

OpenRouter is a hosted model router that supports prompt caching on providers that offer it, with automatic caching on OpenAI, Grok, Moonshot AI, Groq, and DeepSeek models and cache_control breakpoints for Anthropic and Alibaba. Its main cache-related routing feature is provider sticky routing, which sends follow-up requests to the provider endpoint that already holds the cache.

  • Markers and translation: Anthropic breakpoints, top-level automatic caching, and OpenAI's prompt_cache_breakpoint are supported, and a block marked in one style is converted to the other across providers (TTLs are not translated).
  • Sticky routing: activates after a cache hit, or immediately when a session_id is sent; sessions expire after 10 minutes of inactivity.
  • Cost visibility: usage reports cached_tokens and cache_write_tokens, and a cache_discount field shows the saving per generation.
  • Response caching: a separate, opt-in cache returns identical requests with all billable usage reported as zero.
  • Breakpoint injection: gateway-side injection is not published; callers rely on provider automatic caching or set markers themselves.

Best for: teams that want a hosted, pay-as-you-go endpoint across many models and do not need self-hosting. Teams that need the gateway inside their own network can review this production comparison of OpenRouter alternatives.

4. Cloudflare AI Gateway

Cloudflare AI Gateway is a managed gateway whose built-in caching is response caching for identical requests. Its prompt caching relevance is narrower: custom cost rates for cache-read and cache-write tokens, and session affinity in its Auto Router so multi-turn conversations can reuse a provider prompt cache.

  • Response caching: exact match only, keyed on a SHA-256 hash of provider, endpoint, model, auth header, and full request body; semantic caching is listed as planned.
  • Cache controls: cf-aig-skip-cache, cf-aig-cache-ttl (60 seconds to one month), and cf-aig-cache-key headers.
  • Cache-token costs: the cf-aig-custom-cost header accepts per_cache_read_token and per_cache_write_token, and the gateway accounts for whether a provider includes cache tokens in its input count, so they are not double-counted.
  • Session affinity: a cf-aig-session-id header keeps Auto Router on the same model for a turn so requests can take advantage of prompt caching.
  • Breakpoint handling: marker translation or injection is not published.

Best for: teams already on Cloudflare that want edge-level exact-match caching and analytics with minimal setup. For self-hosted control of caching and routing, see this review of Cloudflare AI Gateway alternatives.

5. Kong AI Gateway

Kong AI Gateway extends the Kong API gateway with AI plugins, and its documented caching is semantic response caching through the AI Semantic Cache plugin in the AI Gateway Enterprise tier. Kong publishes no provider prompt caching features; its relevance is that native-format mode forwards provider requests without transformation.

  • Semantic cache: returns a cached response for semantically similar prompts from a vector database; Kong's cookbook notes that cache hits are model-agnostic and skip routing.
  • Native-format passthrough: setting config.llm_format to a provider format such as anthropic (v3.10+) passes requests upstream without payload conversion while keeping analytics, logging, and cost calculation, so markers a client sets reach the provider unchanged.
  • Prompt caching specifics: breakpoint injection, cache-aware routing, and cached-token pricing are not published.

Best for: organizations already standardized on Kong for API management that want semantic response caching in the same plugin model. Teams that need provider prompt caching handled at the gateway can compare Kong AI Gateway alternatives or the open-source LLM gateways for self-hosted deployments.

How to Choose a Prompt Caching AI Gateway

Choose based on where your cache misses come from. Clients that send no markers need breakpoint injection, and load balancing across keys or providers needs session affinity. Spend reconciliation needs cached-token pricing. If repeated questions dominate, add response caching.

  • Coding agents and agent loops: pick a gateway with injection and session affinity. Claude Code and Codex CLI replay long prefixes every turn, and routing Claude Code through Bifrost gives those sessions affinity from their existing session headers.
  • Multi-key or multi-provider load balancing: confirm that affinity applies at the key level, not just the provider level, since caches are usually scoped per API key.
  • One-shot traffic: leave injection off. A one-shot request never reads what it wrote, so a marker adds 25% to 100% to the marked prefix on Anthropic for no return.
  • Prefix hygiene and verification: keep timestamps and reordered tool lists out of the prefix, and compare cached_read_tokens with cached_write_tokens per session; writes with no reads mean the prefix is changing.

This guide on reducing Claude Code token costs pairs prompt caching with routing and budgets, and the LLM gateway buyer's checklist covers the remaining criteria.

Frequently Asked Questions

What is prompt caching and how does it work?

Prompt caching lets an LLM provider store the processed prefix of a request and reuse it when a later request repeats that prefix exactly. The provider hashes the prompt up to a cache breakpoint; on a match, it bills those tokens at a reduced cache-read rate. The model still generates a new response. Bifrost maps each provider's counters into one usage block, as shown in the Claude cache token mapping.

Does Anthropic support prompt caching?

Yes. Anthropic supports prompt caching on Claude models through cache_control markers on up to four content blocks, or one top-level field for automatic caching. Cache reads cost 0.1x the base input price on most Claude models, and writes cost 1.25x (5 minutes) or 2x (1 hour). Clients that send no markers get no caching unless a gateway injects them.

How long does Anthropic cache prompts?

Anthropic's default cache lifetime is 5 minutes, and a 1 hour lifetime is available by adding "ttl": "1h" to the marker at a higher write price. Each time the cached content is used, Anthropic refreshes the lifetime at no additional cost. The 1 hour option suits workflows where turns are minutes apart.

Should I use prompt caching?

Use prompt caching when the same long prefix is reused many times within the cache lifetime, such as agent loops, large system prompts, or RAG context. Avoid it for one-shot requests, because a cache write costs more than fresh input on providers like Anthropic. Configuring it per provider with a per-request opt-out lets one gateway serve both traffic patterns.

Does OpenAI cache prompts automatically?

Yes. OpenAI caches eligible prompt prefixes automatically, with no markers required; on GPT-5.6 and later the minimum cacheable prefix is 1,024 tokens. Those models also support explicit breakpoints through prompt_cache_breakpoint, and those models bill cache writes at 1.25x input. Bifrost translates a Claude-style marker into OpenAI's explicit format for the gpt-5.6 family on the Responses API.

Can prompt caching and semantic caching run together?

Yes. They operate at different layers. A semantic cache in the gateway answers repeated or similar questions without calling the provider, and requests that miss it still reach the provider, where prompt caching discounts the repeated prefix. Bifrost runs both independently, with semantic response caching keyed by a cache key header and prompt caching configured per provider.

Get Started with Prompt Caching on Bifrost

Prompt caching only lowers spend when requests carry cache markers, return to the provider key that holds the cache, and are priced with separate read and write rates. The Bifrost AI gateway handles all three at the gateway layer, alongside semantic caching and enterprise governance, with 11 microseconds of overhead. Explore the Bifrost resources hub, or book a demo to see prompt caching, session affinity, and cost tracking running on your own traffic.