Caching at the Gateway
Model calls are slow and expensive, and a surprising fraction of them are repeats or near-repeats. Caching at the gateway turns those into instant, free responses — and because the gateway sees all traffic, it's the one place a cache benefits every application at once. Beyond exact-match caching, semantic caching catches queries that mean the same thing in different words, which is where the real savings live.
The gateway sees every model call, which makes it the natural place to cache. This post covers caching as a first-class gateway capability: exact caching for identical requests, semantic caching for equivalent ones, and the correctness questions (when not to cache) that caching always raises. Done well, it’s one of the highest-ROI features a gateway offers.
Why cache model calls
Two properties of model calls make caching unusually valuable: - They’re expensive — you pay per token, on every call. A cache hit costs nothing. - They’re slow — generation latency is high (often seconds). A cache hit is near-instant.
So a cache hit saves both money and latency, and meaningfully: in many workloads a real share of requests are repeats — the same question asked repeatedly, the same document summarized, common prompts recurring across users. Every such repeat served from cache is a provider call avoided. Caching is the gateway feature with perhaps the clearest, most direct ROI: it directly cuts the two costs (money and time) that matter most for LLM systems. And because it lives at the gateway, one cache serves all apps — a repeat generated by one app can be a hit for another.
Exact caching
The simplest form: exact caching keys on the precise request (the prompt, model, and parameters) and returns the stored response when an identical request recurs. If the same prompt with the same settings comes in twice, the second is served from cache.
Exact caching is easy to reason about and safe (identical input → identical cached output), and it catches genuinely repeated calls — common in automated pipelines, retries, repeated document processing, and identical user queries. Its limitation is that it only hits on byte-identical requests: change one word, one whitespace, or a parameter, and it misses. For free-form natural language, exact repeats are rarer than equivalent ones — which is what semantic caching targets.
Semantic caching
Semantic caching is the powerful, LLM-specific idea: cache based on meaning, not exact text. “What is the capital of France?” and “France’s capital city?” are different strings but the same question, and should hit the same cache entry.
How it works, using the same embedding machinery as vector search: 1. Embed the incoming request into a vector. 2. Search the cache for a semantically similar prior request (nearest-neighbor within a similarity threshold). 3. If a close-enough match exists, return its cached response (a semantic hit); otherwise call the provider and store the new request’s embedding + response.
Semantic caching dramatically raises the hit rate for natural-language traffic, because it catches the many ways people phrase the same intent. It’s essentially applying the vector-search stack (embeddings + similarity search) to caching. But it introduces a genuine risk that exact caching doesn’t: the similarity threshold is a correctness knob. Set it too loose and you’ll return a cached answer for a query that’s similar but not actually equivalent — a wrong answer served fast. Set it too tight and you lose most of the benefit. Tuning that threshold — and choosing where semantic caching is safe to use at all — is the central engineering decision, and it’s why semantic caching suits some workloads (FAQs, common queries) far better than others (anything where subtle differences change the correct answer).
When NOT to cache: the correctness questions
Caching always raises “is a stale/reused answer correct here?”, and getting this wrong is worse than not caching. Cases where caching is dangerous or needs care: - Non-deterministic or intentionally-varied output — if you want variety (creative generation, sampling at high temperature), caching defeats the purpose by returning the same thing. - Personalized or context-dependent responses — a prompt that includes user-specific or session-specific context isn’t safely shareable across users; the cache key must include everything that affects the answer, or you leak one user’s response to another. This is a real correctness and privacy hazard. - Time-sensitive answers — anything depending on current state (retrieved live data, “today’s …”) goes stale; needs TTLs or exclusion. - Freshness requirements — even cacheable content may need expiry (TTL) so it doesn’t serve outdated answers forever.
The disciplines: make the cache key include everything that affects the response (prompt, model, parameters, and any injected context or user scope), set TTLs appropriate to freshness needs, allow per-request opt-out (some calls should bypass cache), and be conservative with semantic caching where correctness is subtle. The rule of thumb: cache aggressively where answers are stable and shared, cache carefully or not at all where they’re personalized, varied, or time-sensitive.
Caching in the request flow
Caching sits at the front of the gateway’s request handling (as the sequence diagram in the reliability post showed): on each request, check the cache first; a hit returns immediately, skipping routing, the provider call, and its cost entirely; a miss proceeds to route and call, then stores the result for next time. This placement is why caching compounds with everything else — a hit avoids not just the provider cost but the routing and reliability machinery too. Positioned at the gateway, one well-tuned cache cuts cost and latency across every application, which is exactly the leverage the choke-point design (post 1) was meant to provide.
Key takeaways
- Caching model calls saves both money and latency (the two costs that matter most for LLMs), and because the gateway sees all traffic, one cache benefits every app — a repeat from one app can be a hit for another. Highest-ROI gateway feature.
- Exact caching keys on the precise request (prompt + model + params) — safe and simple, but only hits byte-identical requests, so it misses the equivalent-but-reworded calls common in natural language.
- Semantic caching keys on meaning via embeddings + nearest-neighbor search — catching different phrasings of the same intent and dramatically raising hit rate — but the similarity threshold is a correctness knob: too loose returns a wrong-but-similar answer, too tight loses the benefit.
- Know when not to cache: intentionally-varied/creative output, personalized/context-dependent responses (the cache key must include user scope, or you leak across users — a privacy hazard), and time-sensitive answers; use TTLs and per-request opt-out.
- Cache at the front of request handling (check first, hit returns instantly skipping routing + provider + cost) — cache aggressively where answers are stable and shared, carefully or not at all where personalized/varied/time-sensitive.