Semantic Caching: The Top 5 AI Gateways in 2026
Semantic caching serves a stored LLM response when a new prompt means the same thing as an earlier one. This guide compares Bifrost, Kong AI Gateway, Azure API Management, and Cloudflare AI Gateway on match modes, vector stores, thresholds, TTLs, and cache scoping.
TL;DR
- Semantic caching returns a stored LLM response when a new request's embedding is close enough to a cached one, so paraphrased prompts skip the provider call.
- Bifrost runs exact-match (direct) caching first and a similarity lookup on a miss, with Redis/Valkey, Weaviate, Qdrant, or Pinecone as the vector store.
- Bifrost caches chat completions, text completions, the Responses API, embeddings, transcriptions, speech, and image generation, including streaming variants.
- Kong AI Gateway, Azure API Management, and Google Apigee offer semantic cache policies; Cloudflare AI Gateway caching is exact-match only.
- A semantic miss pays an embedding call on top of the full LLM call, so thresholds, TTLs, and cache key scope decide whether caching saves money.
Teams running LLM traffic at scale pay repeatedly for answers they have already generated, because users ask the same questions in different words and exact-match caches only catch identical requests. Matching on meaning closes that gap, and Bifrost, the open-source AI gateway written in Go, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. This post compares five AI gateways on how their caches actually work: match modes, vector store backends, similarity thresholds, TTLs, cache key scoping, covered request types, and streaming. Competitor details come from each vendor's own documentation; undocumented details are marked "Not published."
What Semantic Caching Means at the Gateway Layer
Semantic caching is a response cache that embeds each incoming prompt, searches a vector store for previously answered prompts with similar meaning, and replays the stored response when the similarity score clears a threshold. Running this cache in the AI gateway means one cache serves every application, provider, and model behind that gateway.
An application-level cache has to be rebuilt in every service and every language. A gateway-level cache sits on the request path once, so a support assistant, an internal search tool, and a batch job can share a cache policy while keeping their entries isolated. Our technical deep dive on embedding-based caching covers the embedding math; this post focuses on how gateways implement it.

Figure 1: The exact-match path is cheap and runs first; the semantic path adds an embedding call and only runs on an exact miss.
The lookup order in Figure 1 matters for cost: an exact-match lookup is one storage round trip, while a semantic lookup must call an embedding model before it can search. Academic results support the approach: the GPT Semantic Cache paper by Regmi and Pun reported cache hit rates of 61.6% to 68.8% and positive-hit accuracy above 97% across its query categories.
Semantic Caching vs Exact-Match and Prompt Caching
Exact-match caching replays a response only for byte-identical (or normalized-identical) requests. A semantic cache replays a response for requests with similar meaning. Prompt caching is different in kind: the provider reuses a cached prompt prefix, so the call still happens and is still billed, only at a lower input-token rate.
Bifrost supports automatic prompt caching as a separate feature that injects cache breakpoints for clients that send none. Provider pricing shows why the distinction matters: Anthropic's prompt caching documentation prices 5-minute cache writes at 1.25x base input and cache reads at 0.1x base input on most models, so the provider is still paid on every turn.
| Caching type | What is reused | Is the provider called? | Typical hit condition |
|---|---|---|---|
| Exact-match (direct) caching | Full response | No, on a hit | Identical normalized request |
| Semantic caching | Full response | No, on a hit | Embedding similarity above a threshold |
| Prompt caching | Prompt prefix on the provider side | Yes, billed at a reduced input rate | Repeated prefix up to a cache breakpoint |
For a fuller treatment of the provider-side mechanism, see our comparison of AI gateways with prompt caching support.
How We Compared AI Gateways for Semantic Caching
We compared each AI gateway on seven caching mechanics that decide hit rate, correctness, and cost: match modes, vector store backends, threshold controls, TTL controls, cache key scoping, covered request types, and streaming. Savings percentages were excluded because they depend on workload. The table summarizes what each gateway publishes.
| Capability | Bifrost | Kong AI Gateway | Azure API Management | Google Apigee | Cloudflare AI Gateway |
|---|---|---|---|---|---|
| Match modes | Exact (direct) and semantic, together or direct-only | Exact caching and semantic caching | Semantic (LLM policies) | Semantic | Exact-match only |
| Vector store | Redis/Valkey, Weaviate, Qdrant, Pinecone | Redis, Redis Cloud, Valkey, managed Redis, PostgreSQL with pgvector | Azure Managed Redis with RediSearch | Vertex AI Vector Search | Not applicable |
| Threshold | Cosine similarity, default 0.8, per-request override (stricter only) | vectordb.threshold, cosine or Euclidean |
score-threshold, 0.0 to 1.0, lower is stricter |
Threshold, default 0.9 for dot product |
Not applicable |
| TTL | Default 5 minutes, per-request override | Cache-Control max-age / s-maxage |
duration on the store policy |
TTLInSeconds, default 60 |
60 seconds to one month, per-request override |
| Cache key scoping | Required cache key, auto-scoped per virtual key, plus model and provider | Not published | vary-by expressions |
Not published | Hash of provider, endpoint, model, auth header, and body, or a custom key |
| Request types | Chat, text completions, Responses API, embeddings, transcription, speech, image generation | Chat requests (full list not published) | OpenAI Chat Completions or Responses, Anthropic Messages, Vertex AI | Not published | Text and image responses |
| Streaming | Cached and replayed chunk by chunk | Cached response can be streamed back | Not published | Not published | Not published |
| Availability | Open source, Enterprise tier | AI Gateway Enterprise only | All API Management tiers | Apigee policy | Cloudflare AI Gateway |
Teams building a shortlist beyond caching can use the LLM gateway buyer's guide for routing, governance, and deployment criteria.
1. Bifrost
Bifrost is an open-source AI gateway that runs exact-match and semantic caching as one plugin, scoped per virtual key, across 25+ providers and 10,000+ models through one OpenAI-compatible API. Bifrost adds 11 microseconds of overhead per request at 5,000 RPS, so the cache lookup, not the gateway, dominates added latency.
Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.
The Bifrost AI gateway implements semantic caching as the semantic_cache plugin, labeled Local Cache in the web UI. It offers two lookup paths that can run together or separately:
- Direct (hash) matching: the request is normalized and hashed from its input, parameters, and stream flag; identical requests replay instantly with no embedding call.
- Semantic matching: on a direct miss, Bifrost embeds the request with a configured embedding provider and serves a cached answer when cosine similarity meets the threshold (default 0.8).
- Direct-only mode: set
dimension: 1and omit the embedding provider to get exact-match caching with no embedding cost at all.

Figure 2: Every entry is partitioned by virtual key, cache key, model, and provider, so tenants never share cached answers by accident.
Vector stores. Bifrost stores cache entries in its vector store layer, which supports Redis/Valkey-compatible stores, Weaviate, Qdrant, and Pinecone. Direct-only mode requires Redis or Valkey, because the other three backends require a vector for every entry. Entries persist across restarts.
Cache keys and scoping. Caching engages only when a request carries an x-bf-cache-key header or the deployment sets a default_cache_key, which makes caching an explicit opt-in. When a request resolves to one of Bifrost's virtual keys, the cache partition is automatically scoped to that virtual key, so two teams can never share an entry even with the same cache key. Model and provider names are part of the key by default (cache_by_model, cache_by_provider).
Request types and streaming. Bifrost caches chat completions, text completions, the Responses API (including WebSocket), embeddings, transcriptions, speech, and image generation, including their streaming variants. Streamed responses are cached and replayed chunk by chunk, and only the final chunk carries the cache_debug payload.
Per-request controls. Every plugin default can be overridden per request:
| Header | Effect |
|---|---|
x-bf-cache-key |
Scopes the request to a cache partition; required unless default_cache_key is set |
x-bf-cache-ttl |
Overrides the TTL (default 5 minutes) for this request |
x-bf-cache-threshold |
Raises the similarity threshold; values below the configured threshold are floored, so callers can only make matching stricter |
x-bf-cache-type |
Limits lookup to direct or semantic |
x-bf-cache-no-store |
Serves cached hits but skips writing the response |
Observability and invalidation. Each response carries cache_debug metadata with cache_hit, hit_type, cache_id, and, on semantic hits, the similarity score and embedding tokens used. The built-in request logs tag hits with Direct Cache or Semantic Cache badges and support filtering by hit type. Entries can be cleared by cache_id or by cache key, and conversations longer than conversation_history_threshold messages (default 3) skip caching.
Because Bifrost is a drop-in replacement for OpenAI, Anthropic, and other SDKs, enabling the cache requires a base URL change and a header, not a code rewrite. The performance benchmarks cover gateway overhead under sustained load.
2. Kong AI Gateway
Kong AI Gateway offers meaning-based caching through its AI Semantic Cache plugin, which supports both an exact caching mode and embedding-based semantic matching. Kong documents the plugin as available only in its AI Gateway Enterprise offering, with Redis-family stores or PostgreSQL with pgvector as the vector database.
Best for: Organizations already standardized on Kong Gateway Enterprise that want a semantic cache for chat traffic inside the same plugin framework.
Kong's plugin page lists Redis with Redis Vector Search, Redis Cloud, Valkey, managed Redis on AWS ElastiCache, Azure Managed Redis, and Google Cloud Memorystore, plus PostgreSQL with pgvector. Kong's similarity documentation describes a vectordb.threshold setting passed directly to the vector engine and a distance_metric choice of cosine or Euclidean. Cache expiry follows Cache-Control directives: max-age and s-maxage set how long the vector database keeps a cached response.
What Kong's plugin page does not publish in the material we reviewed: a full list of cacheable request types beyond LLM chat requests, and how entries are partitioned by consumer or route. Teams that need an open-source similarity cache, or caching for embeddings and audio endpoints, can compare the Bifrost alternatives pages.
3. Azure API Management
Azure API Management provides a semantic cache through two policies, llm-semantic-cache-lookup and llm-semantic-cache-store, backed by Azure Managed Redis with the RediSearch module and an Azure-hosted embeddings deployment. Microsoft documents the policies as available across all API Management tiers.
Best for: Azure-centric teams that already publish model APIs through API Management and want similarity caching as a policy on those APIs.
The lookup policy takes a score-threshold from 0.0 to 1.0, where lower values require higher similarity, so lower is stricter. Microsoft recommends starting near 0.05 and warns that values above 0.2 may cause mismatches. The vary-by element partitions the cache by a runtime expression such as a subscription or user ID, ignore-system-messages removes system prompts before similarity scoring, and max-message-count skips caching for longer dialogs.
Supported APIs are OpenAI Chat Completions or Responses, the Anthropic Messages API (on v2 tiers), and the Google Vertex AI API. One operational constraint stands out: the RediSearch module can only be enabled when the Azure Managed Redis cache is created, not added later. Streaming behavior for cached responses is not published on the policy pages. Teams running Azure OpenAI alongside other providers can keep it behind Bifrost under one cache policy.
4. Google Apigee
Google Apigee supports similarity-based caching with the SemanticCacheLookup and SemanticCachePopulate policies, which use the Vertex AI text embeddings API to embed prompts and Vertex AI Vector Search to find similar ones. The setup requires a Vertex AI project, a Vector Search index, and a deployed index endpoint.
Best for: Google Cloud teams already managing APIs in Apigee that want a cache tied to Vertex AI embeddings and Vector Search.
The lookup policy's Threshold defaults to 0.9, which suits the default DOT_PRODUCT_DISTANCE measure; the configured distance measure decides whether the comparison is "greater than" or "less than." Only the highest matching data point is used. The populate policy sets entry lifetime with TTLInSeconds, which defaults to 60 seconds.
Apigee's policy reference does not publish an exact-match LLM mode, a per-tenant cache partitioning element, the list of request types covered, or streaming behavior. Apigee fits teams whose APIs already run on Apigee and whose models run on Vertex AI; the cache depends on Google Cloud services end to end.
5. Cloudflare AI Gateway
Cloudflare AI Gateway caches responses by exact match of the entire request and does not offer semantic caching; Cloudflare states that semantic search for caching is planned. The default cache key is a SHA-256 hash of the provider, endpoint, model, provider auth header, and full request body.
Best for: Teams on Cloudflare that need simple exact-match deduplication for identical requests and do not need paraphrase matching.
Cloudflare documents a TTL range of 60 seconds to one month. Per-request headers control behavior: cf-aig-skip-cache bypasses the cache, cf-aig-cache-ttl sets the TTL, and cf-aig-cache-key replaces the default key so requests can share a cached response. Caching covers text and image responses. Cloudflare also notes that its cache is volatile: two identical requests sent at the same time may both miss.
Exact-match caching suits retries and templated prompts, the same job Bifrost's direct mode does, but it produces no hits for paraphrased queries. Our guide to cutting LLM costs with semantic caching at the gateway explains which workloads repeat exactly and which only repeat in meaning.
Tuning Similarity Thresholds, TTLs, and Cache Keys in Production
A semantic cache pays off only when hits are frequent and correct. The three settings that decide that are the similarity threshold, the TTL, and the cache key scope, and each gateway exposes them differently.

Figure 3: A semantic miss is slower than no cache at all, so the threshold and cache key scope decide whether semantic caching pays off.
As Figure 3 shows, a Bifrost direct lookup costs one vector store round trip, sub-millisecond to a few milliseconds on a local Redis or Valkey. A semantic lookup adds one embedding API call (typically tens to a few hundred milliseconds) plus a vector search, paid on every direct miss. Cache writes are asynchronous, so they add no response latency. Practical guidance:
- Start with a strict threshold. Bifrost defaults to 0.8 cosine similarity; raise it per request with
x-bf-cache-thresholdfor answers where a near-miss would be wrong, such as account or pricing questions. - Match TTL to data volatility. A 5-minute default suits conversational answers; reference content can live for hours, while anything tied to live data should use a short per-request TTL or
x-bf-cache-no-store. - Scope keys to the sharing boundary. A coarse key (a feature name) maximizes hit rate; a per-user or per-session key keeps caches private. Bifrost adds virtual key scoping automatically on top.
- Use direct-only mode when prompts are templated. It removes embedding cost and latency entirely.
- Watch the similarity score on misses. Bifrost returns
similarityincache_debug, which shows how close a near-miss came before you lower a threshold.
Our walkthrough on optimizing LLM cost and latency with a gateway cache covers conversation-aware settings in more depth, and per-team spend limits that complement caching are covered on the Bifrost governance page.
Choosing an LLM Caching Strategy for Your Team
The right LLM caching strategy depends on two questions: whether your traffic repeats in meaning or only verbatim, and whether your stack is tied to one cloud API platform. Paraphrase matching across multiple providers and deployment environments points to a provider-neutral gateway; verbatim repetition can be served by any exact-match cache.

Figure 4: Paraphrase matching across providers and environments points to Bifrost; single-cloud and exact-match needs have narrower options.
Cloud platform policies (Azure API Management, Apigee) work when every model runs on that cloud, but the cache inherits that cloud's embedding service and deployment boundary. Bifrost serves teams that route across 25+ supported providers, need the same cache policy in every environment, and want to run the gateway in their own infrastructure.
For regulated workloads, the cache holds full prompts and responses, so where it runs matters. Bifrost runs inside your network with in-VPC deployments, scales horizontally with clustering, and keeps cache entries in a vector store you operate. Bifrost Enterprise adds governance and compliance controls.
For a broader view, see how semantic caching works and the tools that do it, or the fundamentals of meaning-based response caching for threshold theory.
Frequently Asked Questions
What is semantic caching?
Semantic caching is a technique that stores LLM responses alongside embeddings of the prompts that produced them, then serves a stored response when a new prompt's embedding is similar enough. Unlike exact-match caching, it returns hits for paraphrased questions. In Bifrost, semantic matching runs after an exact hash lookup misses, using a configurable cosine similarity threshold.
How to implement semantic caching?
A gateway-level implementation takes three parts: a vector store, an embedding model, and a cache policy. In Bifrost, enable a vector store such as Redis or Qdrant in config.json, turn on the cache plugin with an embedding provider, model, and dimension, then send requests with an x-bf-cache-key header. The gateway setup guide covers the first deployment.
What is the difference between semantic caching and prompt caching?
Semantic caching replays a full response from the gateway, so the provider is never called on a hit. Prompt caching happens at the provider: a repeated prompt prefix is reused at a lower input-token rate, but the call still runs and is still billed. The two are independent, and Bifrost supports both, including automatic prompt cache breakpoint injection for clients that send none.
What similarity threshold should a semantic cache use?
Start strict and loosen only with evidence. Bifrost defaults to 0.8 cosine similarity and returns the actual similarity score on each semantic lookup, so teams can see how close misses come before lowering the threshold. Threshold values are not portable across gateways, because Azure API Management scores lower-is-stricter. The semantic cache field reference lists every threshold option.
Does semantic caching work with streaming responses?
It depends on the gateway. Bifrost caches streamed responses and replays them chunk by chunk, for chat, Responses API, and other streaming request types, with cache metadata on the final chunk. Kong documents that a cached response can be streamed back. Azure API Management, Apigee, and Cloudflare do not publish streaming behavior for cached responses. Bifrost's streaming guide covers the request format.
Which vector database is best for a semantic cache?
Redis or Valkey is the safest default for Bifrost because it handles both exact-match and semantic entries. Bifrost also supports Weaviate, Qdrant, and Pinecone for semantic mode, but direct-only mode requires Redis or Valkey because the other stores require a vector for every entry. Pick the store your team already operates; the vector store reference lists per-store setup.
Try Bifrost for Semantic Caching
A gateway cache cuts repeat LLM spend only when the cache matches on meaning, scopes entries to the right tenant, expires them on time, and covers the request types your applications actually send. Bifrost does this with exact and semantic matching in one plugin, four vector store options, automatic virtual key scoping, and caching for chat, embeddings, audio, and image generation, including streams. Explore more on the Bifrost resources hub, or book a demo to see semantic caching on your own traffic.