Top 5 AI Gateways with Prompt Caching Support in 2026
Compare the top 5 AI gateways with prompt caching support in 2026 on cache breakpoint injection, session affinity, cached-token cost tracking, and deployment.
TL;DR
- Prompt caching is a provider feature: the provider reuses a cached request prefix and bills it at a reduced input rate, while semantic caching is a gateway feature that replays a stored response without calling the provider.
- An AI gateway affects prompt caching in four ways: whether it forwards cache markers, whether it can inject them for clients that send none, whether it keeps a session on the provider key that holds the cache, and whether it prices cached tokens correctly.
- Bifrost auto-injects cache breakpoints per provider, translates them to each provider's wire format, pins sessions to the provider and key that served them, and prices cache reads and writes separately.
- LiteLLM and OpenRouter support provider prompt caching with sticky routing; Cloudflare AI Gateway adds cache-token cost rates; Kong documents semantic response caching and native-format passthrough only.
Prompt caching is a provider-side optimization that stores the processed prefix of an LLM request so later requests repeating that exact prefix are billed at a reduced input rate and return faster. Bifrost, the open-source AI gateway built in Go by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability, and it handles the prompt caching workflow end to end, from breakpoint injection to cached-token cost tracking. This guide compares five AI gateways on how much of that workflow each one handles.
What Is Prompt Caching?
Prompt caching is a feature of LLM providers that reuses the computed state of a request prefix (system prompt, tool definitions, earlier turns) when a later request repeats it byte for byte. The provider still runs the model and bills the call, but cached input tokens cost much less than fresh input, so agent loops that replay the same prefix every turn benefit most.
The mechanics differ by provider. Anthropic's prompt caching documentation prices cache reads at 0.1x the base input price on most Claude models and allows up to four explicit breakpoints per request. OpenAI's prompt caching guide describes automatic caching with reused tokens discounted by up to 95%.
| Provider | How caching is triggered | Marker on the request | Published pricing signal |
|---|---|---|---|
| Anthropic (Claude) | Explicit breakpoints, or a single top-level cache_control field |
cache_control: {"type": "ephemeral"}, at most 4 breakpoints |
Reads 0.1x input on most models; writes 1.25x (5 min) or 2x (1 hour) |
| OpenAI | Automatic (implicit); explicit mode on GPT-5.6 and later, with a 1,024-token minimum prefix | prompt_cache_breakpoint plus prompt_cache_options.mode: "explicit" |
GPT-5.6+: writes 1.25x, reads 0.1x on most models |
| AWS Bedrock | Explicit cache points for Claude and Amazon Nova on the Converse API | cachePoint block |
Varies by model |
| Google Gemini | Implicit prefix caching, or a server-side cached content resource | No inline marker | Varies by model |

Figure 1: The first turn pays a premium to write the prefix; every later turn that repeats it byte for byte pays a fraction of the input price.
As Figure 1 shows, prompt caching pays off only when a prefix is read more often than it is written, and one changed byte before the breakpoint (a timestamp, reordered tools) turns every turn into a write. Both problems sit in the request path, where Bifrost reads cache counters from supported providers that report cache usage through one usage block.
Prompt Caching vs Semantic Caching
Prompt caching and semantic caching solve different problems at different layers. Semantic caching runs inside the gateway and replays a stored response for an identical or similar request, so the provider is never called. Prompt caching runs at the provider, which still generates a fresh completion but bills the repeated prefix at a lower rate. The two are independent and can run together.
| Attribute | Provider prompt caching | Gateway semantic caching |
|---|---|---|
| Where it runs | LLM provider | AI gateway |
| Provider called on a hit | Yes, a new completion is generated | No, a stored response is returned |
| What is reused | Processed prefix of the prompt | The full previous response |
| Match rule | Exact prefix up to a breakpoint | Exact hash or embedding similarity |
| Best fit | Agent loops, long system prompts, RAG context | Repeated FAQ-style questions |
| Billing on a hit | Reduced input rate, normal output rate | No provider charge |

Figure 2: A semantic cache hit removes the provider call entirely, while a prompt cache hit still generates a new completion at a lower input rate.
Many gateways advertise "caching" and mean response caching only. Bifrost ships the two as separate features: Auto Prompt Caching for provider prefixes and semantic caching with direct and similarity lookup for full responses.
For the response side, see this technical deep dive on semantic caching and the roundup of AI gateways with semantic caching for LLM cost reduction.
What an AI Gateway Needs for Prompt Caching
An AI gateway supports prompt caching well when it forwards provider cache markers intact, can add markers for clients that send none, routes each session back to the provider key that holds its cache, and reports cache reads and writes as separate, correctly priced token counts. Response caching is a separate, complementary capability.
| Criterion | Why it matters |
|---|---|
| Marker passthrough and translation | Normalizing requests to one schema can strip cache_control or fail to convert it per provider |
| Automatic breakpoint injection | Coding agents and many SDK callers send no markers, so nothing is cached on providers that require them |
| Cache-aware routing (session affinity) | Provider caches are scoped to a provider and usually an API key, so load balancing each turn independently causes misses |
| Cached-token cost tracking | Cache reads, writes, and fresh input carry different prices |
| Response caching | Removes provider calls for repeated questions; independent of prompt caching |

Figure 3: Provider caches are scoped to a provider and usually an API key, so a gateway that load balances each turn independently pays the write rate repeatedly.
Figure 3 shows why weighted load balancing across two keys can double cache writes. These criteria sit alongside the usual gateway concerns covered in this production-ready comparison of the top LLM gateways and the LLM Gateway Buyer's Guide.
Top AI Gateways for Prompt Caching Compared
The five gateways below range from full prompt caching workflows to response caching only. Bifrost and LiteLLM inject breakpoints and pin sessions, OpenRouter translates markers and uses sticky routing, Cloudflare AI Gateway prices cache tokens, and Kong documents semantic response caching.
| Gateway | Provider cache markers | Automatic breakpoint injection | Cache-aware routing | Cached-token cost tracking | Response caching | Deployment |
|---|---|---|---|---|---|---|
| Bifrost | Translated per provider (cache_control, cachePoint, prompt_cache_breakpoint) |
Yes, per provider, with injection points and a per-request override | Session affinity at provider and key level | Separate cache-read and cache-write rates | Exact-match and semantic | Self-hosted, in-VPC, on-prem, air-gapped |
| LiteLLM | Passthrough, plus prompt_cache_breakpoint for GPT-5.6+ |
Yes, injection points and a one-flag Claude mode | Session affinity pre-call check | Cost function handles cache pricing | Exact-match and semantic | Self-hosted proxy |
| OpenRouter | cache_control and prompt_cache_breakpoint, converted between providers |
Not published | Provider sticky routing | Cached and cache-write tokens, cache_discount |
Exact-match | Hosted service |
| Cloudflare AI Gateway | Not published | Not published | Session affinity in Auto Router | Custom cache-read and cache-write rates | Exact-match | Managed service |
| Kong AI Gateway | Native-format passthrough (v3.10+) | Not published | Not published | Not published | Semantic (Enterprise tier) | Self-hosted or Konnect |
"Not published" means the capability did not appear in the vendor's documentation at the time of writing. Teams comparing access control as well can review the Bifrost governance model.
1. Bifrost
Bifrost, an open-source AI gateway, unifies 25+ providers and 10,000+ models behind one OpenAI-compatible API, and it treats provider prompt caching as a first-class routing concern. Bifrost injects cache breakpoints for clients that send none, keeps sessions on the provider key that holds their cache, and prices cached tokens separately in its cost tracking.

Figure 4: Breakpoint injection makes a request cacheable, session affinity sends it back to the cache it warmed, and cost tracking shows whether it worked.
Automatic breakpoint injection. Agentic clients such as Codex send no cache_control markers, so on Claude models nothing is cached. With automatic prompt cache breakpoints enabled through prompt_cache.auto_inject, Bifrost marks the first cacheable content block of each request that arrives without markers, so turn 1 writes the cache and later turns read it.
- Off by default: a cache write costs more than fresh input, so injection is opt-in per provider.
- Caller markers win: a request that already carries
cache_controlorprompt_cache_breakpointis forwarded unchanged. - Capability gated per model: markers are only injected for models that accept them; implicit-caching providers such as Gemini, DeepSeek, Groq, and xAI receive nothing.
- Four-marker ceiling: injection stops at four markers, matching Anthropic's limit.
- Injection points:
cache_control_injection_pointstarget messages by role, index, or both. - Per-request override: the
x-bf-prompt-cache-auto-injectheader flips injection for one request, as listed in the request options reference. - Optional 1 hour TTL:
"ttl": "1h"requests the longer cache lifetime.
One marker, translated per provider. Bifrost injects one internal marker and each provider translates it: cache_control for Claude on Anthropic, Vertex AI, and OpenRouter; a cachePoint block for Claude and Amazon Nova on AWS Bedrock; and prompt_cache_breakpoint with explicit mode for the OpenAI and Azure OpenAI gpt-5.6 family on the Responses API. Fallbacks re-evaluate injection against the new provider's configuration, so a failover from Anthropic to Bedrock receives a cachePoint when Bedrock has injection enabled. Tokens and cost for every request land in Bifrost request logs.
Session affinity. Session affinity keeps a session on the provider and key that last served it, after routing rules, virtual key load balancing, and the model catalog have built the candidate chain. Sessions come from the x-bf-session-id header, and Claude Code and Codex CLI get affinity with no configuration (Bifrost v2.0.0 and later) because Bifrost adopts the session headers they already send. Each decision is recorded in the routing trail; in Bifrost Enterprise, bindings replicate across a cluster and are dropped when the adaptive load balancer marks a provider as failed.
Cached-token cost tracking. Bifrost surfaces provider cache counters as cached_read_tokens and cached_write_tokens under usage.prompt_tokens_details. The model catalog applies separate rates to cache-read and cache-creation tokens, which keeps virtual key budgets accurate when most input is cached.
Semantic caching alongside it. Bifrost also runs a response cache with direct hash and embedding-based matching; a semantic hit skips the provider, and a miss still benefits from the provider's prefix cache.
Performance and deployment. Bifrost adds 11 microseconds of overhead per request at 5,000 RPS in sustained benchmarks. For regulated environments, Bifrost runs inside a private VPC or on-prem, with Bifrost Enterprise adding clustering and health-aware session invalidation.
Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.
2. LiteLLM
LiteLLM is an open-source Python SDK and proxy that supports provider prompt caching for OpenAI, Anthropic, Google AI Studio, Vertex AI, Bedrock, DeepSeek, and xAI. It forwards cache_control markers, can auto-inject cache checkpoints, and offers session affinity as a router pre-call check.
- Marker passthrough:
cache_controlblocks reach Anthropic, Bedrock, and Gemini, andprompt_cache_breakpointis passed through for OpenAI GPT-5.6 and newer. - Auto-injection:
cache_control_injection_pointsplaces checkpoints from the deployment config; anenable_anthropic_prompt_cachingflag (v1.94.0+) adds a checkpoint on the system prompt and the trailing turn for Claude on Anthropic and Bedrock, respects the four-block limit, and never double-injects. - Session affinity: adding
session_affinitytooptional_pre_call_checkspins a conversation to the deployment that served its first request, reading a session header such asx-claude-code-session-id. - Cost tracking: responses follow the OpenAI usage format with
cached_tokensand, for Anthropic,cache_creation_input_tokens; LiteLLM's cost function handles prompt caching rates. - Response caching: exact-match and semantic caches are separate; LiteLLM's docs warn that semantic caches behave poorly on agentic traffic.
Best for: Python-centric teams that want a self-hosted proxy with configurable cache injection. Teams evaluating throughput and governance at scale can compare LiteLLM alternatives.
3. OpenRouter
OpenRouter is a hosted model router that supports prompt caching on providers that offer it, with automatic caching on OpenAI, Grok, Moonshot AI, Groq, and DeepSeek models and cache_control breakpoints for Anthropic and Alibaba. Its main cache-related routing feature is provider sticky routing, which sends follow-up requests to the provider endpoint that already holds the cache.
- Markers and translation: Anthropic breakpoints, top-level automatic caching, and OpenAI's
prompt_cache_breakpointare supported, and a block marked in one style is converted to the other across providers (TTLs are not translated). - Sticky routing: activates after a cache hit, or immediately when a
session_idis sent; sessions expire after 10 minutes of inactivity. - Cost visibility: usage reports
cached_tokensandcache_write_tokens, and acache_discountfield shows the saving per generation. - Response caching: a separate, opt-in cache returns identical requests with all billable usage reported as zero.
- Breakpoint injection: gateway-side injection is not published; callers rely on provider automatic caching or set markers themselves.
Best for: teams that want a hosted, pay-as-you-go endpoint across many models and do not need self-hosting. Teams that need the gateway inside their own network can review this production comparison of OpenRouter alternatives.
4. Cloudflare AI Gateway
Cloudflare AI Gateway is a managed gateway whose built-in caching is response caching for identical requests. Its prompt caching relevance is narrower: custom cost rates for cache-read and cache-write tokens, and session affinity in its Auto Router so multi-turn conversations can reuse a provider prompt cache.
- Response caching: exact match only, keyed on a SHA-256 hash of provider, endpoint, model, auth header, and full request body; semantic caching is listed as planned.
- Cache controls:
cf-aig-skip-cache,cf-aig-cache-ttl(60 seconds to one month), andcf-aig-cache-keyheaders. - Cache-token costs: the
cf-aig-custom-costheader acceptsper_cache_read_tokenandper_cache_write_token, and the gateway accounts for whether a provider includes cache tokens in its input count, so they are not double-counted. - Session affinity: a
cf-aig-session-idheader keeps Auto Router on the same model for a turn so requests can take advantage of prompt caching. - Breakpoint handling: marker translation or injection is not published.
Best for: teams already on Cloudflare that want edge-level exact-match caching and analytics with minimal setup. For self-hosted control of caching and routing, see this review of Cloudflare AI Gateway alternatives.
5. Kong AI Gateway
Kong AI Gateway extends the Kong API gateway with AI plugins, and its documented caching is semantic response caching through the AI Semantic Cache plugin in the AI Gateway Enterprise tier. Kong publishes no provider prompt caching features; its relevance is that native-format mode forwards provider requests without transformation.
- Semantic cache: returns a cached response for semantically similar prompts from a vector database; Kong's cookbook notes that cache hits are model-agnostic and skip routing.
- Native-format passthrough: setting
config.llm_formatto a provider format such asanthropic(v3.10+) passes requests upstream without payload conversion while keeping analytics, logging, and cost calculation, so markers a client sets reach the provider unchanged. - Prompt caching specifics: breakpoint injection, cache-aware routing, and cached-token pricing are not published.
Best for: organizations already standardized on Kong for API management that want semantic response caching in the same plugin model. Teams that need provider prompt caching handled at the gateway can compare Kong AI Gateway alternatives or the open-source LLM gateways for self-hosted deployments.
How to Choose a Prompt Caching AI Gateway
Choose based on where your cache misses come from. Clients that send no markers need breakpoint injection, and load balancing across keys or providers needs session affinity. Spend reconciliation needs cached-token pricing. If repeated questions dominate, add response caching.
- Coding agents and agent loops: pick a gateway with injection and session affinity. Claude Code and Codex CLI replay long prefixes every turn, and routing Claude Code through Bifrost gives those sessions affinity from their existing session headers.
- Multi-key or multi-provider load balancing: confirm that affinity applies at the key level, not just the provider level, since caches are usually scoped per API key.
- One-shot traffic: leave injection off. A one-shot request never reads what it wrote, so a marker adds 25% to 100% to the marked prefix on Anthropic for no return.
- Prefix hygiene and verification: keep timestamps and reordered tool lists out of the prefix, and compare
cached_read_tokenswithcached_write_tokensper session; writes with no reads mean the prefix is changing.
This guide on reducing Claude Code token costs pairs prompt caching with routing and budgets, and the LLM gateway buyer's checklist covers the remaining criteria.
Frequently Asked Questions
What is prompt caching and how does it work?
Prompt caching lets an LLM provider store the processed prefix of a request and reuse it when a later request repeats that prefix exactly. The provider hashes the prompt up to a cache breakpoint; on a match, it bills those tokens at a reduced cache-read rate. The model still generates a new response. Bifrost maps each provider's counters into one usage block, as shown in the Claude cache token mapping.
Does Anthropic support prompt caching?
Yes. Anthropic supports prompt caching on Claude models through cache_control markers on up to four content blocks, or one top-level field for automatic caching. Cache reads cost 0.1x the base input price on most Claude models, and writes cost 1.25x (5 minutes) or 2x (1 hour). Clients that send no markers get no caching unless a gateway injects them.
How long does Anthropic cache prompts?
Anthropic's default cache lifetime is 5 minutes, and a 1 hour lifetime is available by adding "ttl": "1h" to the marker at a higher write price. Each time the cached content is used, Anthropic refreshes the lifetime at no additional cost. The 1 hour option suits workflows where turns are minutes apart.
Should I use prompt caching?
Use prompt caching when the same long prefix is reused many times within the cache lifetime, such as agent loops, large system prompts, or RAG context. Avoid it for one-shot requests, because a cache write costs more than fresh input on providers like Anthropic. Configuring it per provider with a per-request opt-out lets one gateway serve both traffic patterns.
Does OpenAI cache prompts automatically?
Yes. OpenAI caches eligible prompt prefixes automatically, with no markers required; on GPT-5.6 and later the minimum cacheable prefix is 1,024 tokens. Those models also support explicit breakpoints through prompt_cache_breakpoint, and those models bill cache writes at 1.25x input. Bifrost translates a Claude-style marker into OpenAI's explicit format for the gpt-5.6 family on the Responses API.
Can prompt caching and semantic caching run together?
Yes. They operate at different layers. A semantic cache in the gateway answers repeated or similar questions without calling the provider, and requests that miss it still reach the provider, where prompt caching discounts the repeated prefix. Bifrost runs both independently, with semantic response caching keyed by a cache key header and prompt caching configured per provider.
Get Started with Prompt Caching on Bifrost
Prompt caching only lowers spend when requests carry cache markers, return to the provider key that holds the cache, and are priced with separate read and write rates. The Bifrost AI gateway handles all three at the gateway layer, alongside semantic caching and enterprise governance, with 11 microseconds of overhead. Explore the Bifrost resources hub, or book a demo to see prompt caching, session affinity, and cost tracking running on your own traffic.