Try Bifrost Enterprise free for 14 days. Request access

Semantic Caching for LLMs: How to Cut Token Spend with AI Gateways

Gateway caching runs in two modes: direct hash matching with no embeddings, and semantic matching by meaning. This guide compares five gateways, explains what each mode costs, and shows which workloads see real savings.

Semantic Caching for LLMs: How to Cut Token Spend with AI Gateways
Semantic caching matches LLM requests by meaning rather than exact text, enabling AI gateways to serve cached responses for semantically similar prompts. This can reduce token spend and latency dramatically. This article breaks down how semantic caching works at the gateway layer, then compares five platforms: Bifrost, Cloudflare AI Gateway, LiteLLM, Kong AI Gateway, and Apache APISIX.

TL;DR

  • Caching at the gateway runs in two modes: direct (hash) matching replays a byte-identical request with no embeddings, and semantic matching serves a cached answer when a new request is close enough in meaning.
  • The two can run together, direct first and semantic on a miss, or direct alone with no embedding provider configured at all.
  • Semantic mode adds an embedding call to every lookup, so the saving is the model call minus that embedding call; direct hits avoid the provider entirely.
  • Caching only engages for requests that carry a cache key, which means the hit rate is a design decision rather than something that happens automatically.
  • This guide compares Bifrost, Cloudflare AI Gateway, LiteLLM, Kong AI Gateway, and Apache APISIX, then covers the thresholds, isolation, and metrics that decide whether caching pays off.

Bifrost, the open-source AI gateway by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability.

Production LLM applications have a recurring cost problem: a large portion of the requests sent to model providers are semantically redundant. A customer support bot answering "How do I reset my password?" processes nearly identical intent whether the user types "password reset help," "I forgot my login credentials," or "can't get into my account." Each variation triggers a fresh API call, burns tokens, and adds latency.

Traditional exact-match caching only helps when prompts are character-for-character identical, which is rare in natural language. Semantic caching solves this by comparing the meaning of incoming requests against previously cached ones using vector embeddings and similarity search. When a match exceeds a configurable similarity threshold, the cached response is returned instantly and no LLM call is made.

The impact depends on which mode does the work. A direct (hash) hit is a lookup with no embedding step, so it returns in milliseconds against seconds for a full inference call. A semantic hit first embeds the incoming request, so the saving is the model call minus one embedding call, and the net gain is still large for expensive models but is not free. Either way, the ceiling on savings is the share of traffic that repeats, which is why hit rate is the number to measure before and after, alongside the other levers in enterprise AI gateways for controlling AI costs.


How It Works at the Gateway Layer

The most effective place to implement semantic caching is at the AI gateway layer. A centralized gateway ensures every request across all services benefits from a shared cache, improving hit rates as usage scales.

The typical flow involves converting the incoming prompt into a vector embedding, running a similarity search against stored embeddings in a vector database, evaluating whether cosine similarity exceeds a defined threshold (commonly 0.90-0.98), and either returning the cached response or forwarding the request to the LLM provider and caching the new result.

The key tuning parameter is the similarity threshold. A strict threshold (0.98) minimizes false positives but limits hit rates. A relaxed threshold (0.85) maximizes savings but risks returning generic answers for subtly different queries.

Three design decisions matter as much as the threshold:

DecisionWhy it mattersPractical default
Which requests carry a cache keyCaching only engages when a request is keyed, so an unkeyed workload sees no hits at allKey the high-repetition paths first: support answers, docs Q&A, classification
Direct, semantic, or bothDirect needs no embedding provider; semantic pays one embedding call per lookupStart direct-only, add semantic where wording varies
Cache scope per tenantA shared cache across tenants can return one customer's answer to anotherScope by tenant or user ID from the start

Streaming responses need care too: a cached reply has to be replayed in the same chunk order the client expects, or the application sees a different shape on a hit than on a miss.


1. Bifrost

Platform Overview

Bifrost is a high-performance, open-source AI gateway built in Go by Maxim AI. It unifies access to 25+ providers (OpenAI, Anthropic, AWS Bedrock, Google Vertex, Azure, and more) through a single OpenAI-compatible API, delivering approximately 11 microseconds of gateway overhead at 5,000 requests per second.

Features

Bifrost ships semantic caching as a first-class, built-in plugin with a dual-layer system: exact hash matching plus vector similarity search. Cache hits return in roughly 5 milliseconds, and the system supports multiple vector store backends including Weaviate, Redis/Valkey, Qdrant, and Pinecone. Teams can tune the similarity threshold per use case, and governance features enable multi-tenant cache isolation using tenant or user IDs, preventing data leakage in SaaS deployments. Cached responses also support full streaming with proper chunk ordering.

Beyond caching, Bifrost provides automatic fallbacks, adaptive load balancing, MCP gateway support, budget management with virtual keys, and native Prometheus-based observability. It integrates directly with Maxim's AI evaluation and observability platform for end-to-end production monitoring.

Best For

Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

Get started in seconds with npx -y @maximhq/bifrost or via GitHub.


2. Cloudflare AI Gateway

Platform Overview

Cloudflare AI Gateway is a managed proxy service that leverages Cloudflare's global edge network to add caching, rate limiting, retries, and analytics with a single line of code.

Features

Provides exact-match caching from its edge network with configurable TTL and per-request cache control. Supports 20+ providers with real-time analytics and cost tracking. Core features are free on all plans. However, Cloudflare currently does not support semantic caching; only character-identical requests trigger cache hits.

Best For

Teams already on Cloudflare that need lightweight, free observability and caching for AI traffic with limited prompt variability.


3. LiteLLM

Platform Overview

LiteLLM is an open-source Python-based gateway providing unified access to 100+ LLM providers through OpenAI-compatible APIs.

Features

Supports exact-match caching via Redis and in-memory backends, along with a semantic caching option using embedding-based similarity search. Also provides virtual key management, spend tracking, rate limiting, and basic load balancing.

Best For

Python-centric teams that need wide provider coverage and quick unification of LLM calls for development and moderate-scale production.


4. Kong AI Gateway

Platform Overview

Kong AI Gateway extends Kong's API management platform with AI-specific capabilities including prompt engineering guardrails, multi-LLM routing, and token-level rate limiting.

Features

Provides semantic caching through a dedicated plugin that uses vector embeddings and configurable similarity thresholds. Supports both open-source and enterprise tiers with an extensive plugin ecosystem.

Best For

Organizations already running Kong as their API gateway that want to extend existing infrastructure to manage LLM traffic.


5. Apache APISIX AI Gateway

Platform Overview

Apache APISIX is a fully open-source API gateway under the Apache 2.0 license that extended into AI traffic through a dedicated plugin set. Unlike platforms that reserve caching for a commercial tier, its semantic caching ships in the open-source distribution.

Features

The ai-cache plugin, introduced in APISIX 3.18.0, covers exact-match, semantic, and streaming response caching for LLM traffic, with semantic matching that compares prompts by meaning rather than literal text. Prometheus counters for cache hits, misses, and bypasses, plus an embedding-latency histogram, make cache effectiveness measurable at the gateway.

Best For

Best for: teams that want semantic caching in a fully open-source gateway and already run APISIX as their API layer, or are willing to adopt it as one.


Semantic Caching Compared at a Glance

GatewaySemantic cachingCache latencyDeploymentGateway overhead
BifrostYes, built-in dual-layer (hash + vector)~5 ms cache hitsOpen source, self-hosted11 µs gateway overhead
Cloudflare AI GatewayNo, exact-match onlyEdge cacheManaged onlyNot published
LiteLLMYes, embedding-based (Redis)Redis-backedOpen source, self-hostedPython runtime overhead
Kong AI GatewayYes, dedicated pluginVector-store backedOpen-core, self-hosted or SaaSNot published
Apache APISIXYes, ai-cache plugin (3.18+)Vector-store backedOpen source, self-hostedNot published

When Caching Pays Off, and When It Does Not

Caching is not a universal win. It pays off where the same intent arrives repeatedly and the answer does not depend on fresh data or per-user context. It does not pay off where every prompt is unique, where responses must reflect live state, or where a near-miss answer carries real risk.

WorkloadExpected hit rateRecommended mode
Support and FAQ assistantsHigh: users ask the same things in different wordsDirect, then semantic
Documentation and knowledge Q&AHigh for common topicsDirect, then semantic
Classification and moderation from templatesHigh, because prompts are templatedDirect only
Agentic coding sessionsLow: context changes on every turnUsually neither
Personalized recommendations or live dataLow, and a stale hit is a correctness bugNeither

Agent traffic is the clearest negative case: coding sessions change context every turn, so the savings come from code execution with MCP instead. A stale answer is the risk that deserves the most attention. Cached responses are frozen at write time, so any workload where the correct answer changes (pricing, availability, account state) needs either a short TTL or no caching at all on those paths.


Choosing the Right Gateway

If you need the lowest possible overhead with built-in semantic caching and self-hosted deployment, Bifrost is the strongest option. If you are on Cloudflare, their gateway offers solid free exact-match caching. LiteLLM is ideal for Python-heavy teams. Kong makes sense for extending existing API infrastructure. And Apache APISIX fits teams wanting caching tightly integrated with model serving.

Whichever gateway you choose, measure the result rather than assuming it. Four numbers decide whether caching is paying off: hit rate by mode, latency on hits versus misses, embedding calls per hit in semantic mode, and provider spend before and after. Gateway request logs and Prometheus metrics carry all four, and the reporting options across gateways are compared in enterprise AI gateways for LLM observability.

A caching rollout that works usually follows the same order: key the highest-repetition paths, enable direct mode, measure the hit rate for a week, then turn on semantic mode for the paths where wording varies and compare the net saving against the embedding cost. Tenant scoping goes in from the start, not later, since it is the one mistake that leaks data rather than just wasting money. The governance controls around it are covered in governing LLM usage in the enterprise.

Caching sits alongside the other cost levers rather than replacing them: routing cheap work to cheaper models, capping spend per team, and cutting tool-definition overhead for agents. Where prompt caching is available from the provider itself, as Anthropic documents, it works on a different axis: it discounts repeated prefixes inside one conversation, while gateway caching replaces whole responses across conversations. The two compose, and neither substitutes for the other.

To see how caching, routing, and governance fit together on your own traffic, book a demo with the Bifrost team.


Frequently Asked Questions

What is semantic caching for LLMs?

Semantic caching stores model responses and returns one when a new prompt means the same thing as an earlier prompt, even if the wording differs. It embeds each prompt as a vector and compares against stored vectors, rather than hashing the literal text. Bifrost's semantic caching runs as a built-in gateway plugin.

How is semantic caching different from exact-match caching?

Exact-match caching only hits when a later request is byte-identical, which almost never happens with natural language. Semantic caching compares meaning, so "reset my password" and "how do I change my password" resolve to the same entry. That difference is what turns cache hit rates from negligible into meaningful savings.

Does semantic caching reduce response quality?

Not when the similarity threshold is set correctly. Too loose a threshold risks returning a near-but-wrong answer; too strict lowers the hit rate. Tune it per workload and keep it stricter for correctness-sensitive tasks. The threshold is the main control over the accuracy-versus-savings trade-off.

How much can semantic caching save on token costs?

Savings scale with prompt overlap. Support assistants, FAQ bots, and document Q&A see high hit rates and large savings because users ask similar questions; workloads where every prompt is unique save little. Cache hits also return far faster than a model call, so latency improves alongside cost.

Which AI gateways support true semantic caching?

Bifrost, LiteLLM, Kong AI Gateway, and Apache APISIX offer embedding-based semantic caching. Cloudflare AI Gateway provides only exact-match caching. If near-duplicate prompts are common in your workload, exact-match caching will capture very little of the available savings.