Revision note (2026-09-15). The earlier version of this article was a Redis setex wrapper around a time.sleep(1) stand-in model, and its only measurement was that the second call returned "almost instantly", which is true of anything compared with a one-second sleep. It said caching helps with "very similar" requests while implementing a SHA-256 exact-match key that misses on any change; it advised "fallback to direct inference" in the troubleshooting section while the code had no exception handling, so a Redis outage takes the endpoint down (redis-py raises redis.exceptions.ConnectionError from get); it relied on volatile-lru eviction without setting maxmemory, whose default is 0 (no limit); it listed "TTLs" and eviction as reasons to prefer Redis over Memcached, which has both; its setex call draws a DeprecationWarning from redis-py 8.1.0; it required Python 3.8+ (end-of-life since October 2024) and Redis 6+; it never mentioned the request stampede that an exact-match cache creates on every cold key; and its four sources had no links, one of them TorchServe, whose repository was archived in August 2025. This version replaces the wrapper with a seeded simulation in which every number below is produced by the code shown. The workload is synthetic; nothing here is a benchmark of a real model.
Three things called "model caching"
The phrase covers at least three mechanisms that share nothing but the word:
- Response caching: store the model's output under a key derived from the input, serve it again on an identical input. This is what the earlier article built and what the lab measures first.
- Semantic caching: embed the input, store the output under the embedding, serve it again for any input whose embedding is within a similarity threshold. The lab measures this second, including the rate at which it serves a wrong answer.
- Prefix or KV caching inside the serving engine: for transformer models, reuse the attention key/value tensors already computed for a shared prompt prefix. vLLM's documentation describes hashing each KV block by the tokens in the block plus the tokens of the prefix before it, notes that this reduces the prefill phase only, not token generation, and states that it does not change model outputs. It lives inside the engine, keyed on tokens, and is not something you build with Redis. It was checked against the documentation only; the lab does not run an LLM.
Weight caching (keeping model files on local disk or in memory between cold starts) is a deployment concern, not a request-path cache, and is out of scope here.
Tested versions: Python 3.14.7, NumPy 2.4.4, redis-py 8.1.0, Redis 8.2.9 from the redis:8.2-alpine image in Docker 27. The lab is examples/ai-model-caching in the companion repository; python3 -m unittest -v test_caching_lab.py runs 22 tests (3 need REDIS_URL) and python3 caching_lab.py --seed 0 --json results-seed0.json reproduces the tables.
The lab
The "model" is two dense 2048 by 2048 float32 layers with tanh in between, applied to a 16-float payload: a real matrix-vector product that costs about half a millisecond on the laptop used, deterministic for a given weight seed, so a version change is a different weight seed. The request stream draws from 2,000 distinct inputs with Zipf popularity (exponent 1.0) over one simulated hour, and every request carries its own simulated arrival time so TTL expiry is deterministic and does not depend on how fast the script runs. The key is the earlier article's recipe:
def cache_key(payload: dict, model_version: str | None) -> str:
canonical = json.dumps(payload, sort_keys=True, separators=(",", ":"))
raw = canonical if model_version is None else f"{model_version}:{canonical}"
return "cache:model:" + hashlib.sha256(raw.encode("utf-8")).hexdigest()
Tests pin what this key does and does not normalize: key order does not change it, but {"x": 1} and {"x": 1.0} are different keys (json.dumps writes 1 and 1.0), list order changes it, and the model version changes it. Two clients that format floats differently will never share a cache entry.
The exact cache is a dict with the semantics SET key value EX seconds gives you: an entry is gone once now >= stored_at + ttl. It has no size bound, on purpose, because neither did the earlier article's Redis (see the operational notes below).
Exact-match cache: the hit rate belongs to the traffic, not the cache
10,000 requests, 2,000 possible inputs, Zipf 1.0, one simulated hour. Seed 0 saw 1,461 distinct inputs, seed 1 saw 1,458.
| Seed | Policy | Hit rate | Model calls | p50 | p95 |
|---|---|---|---|---|---|
| 0 | no cache | 0.0000 | 10,000 | 536 us | 963 us |
| 0 | exact, no TTL | 0.8539 | 1,461 | 4.9 us | 532 us |
| 0 | exact, TTL 600 s | 0.6989 | 3,011 | 6.7 us | 655 us |
| 0 | exact, TTL 60 s | 0.4433 | 5,567 | 500 us | 762 us |
| 1 | no cache | 0.0000 | 10,000 | 504 us | 641 us |
| 1 | exact, no TTL | 0.8542 | 1,458 | 4.8 us | 531 us |
| 1 | exact, TTL 600 s | 0.7010 | 2,990 | 5.7 us | 561 us |
| 1 | exact, TTL 60 s | 0.4366 | 5,634 | 482 us | 675 us |
Three things to read off this table. The no-TTL hit rate is exactly one minus distinct-inputs over requests; the cache did not create it, the traffic did. With a fixed repeat rate instead of Zipf the hit rate is the repeat rate (seed 0: repeat rate 0.2, 0.5, 0.8 gave hit rates 0.2012, 0.4960, 0.8058), which is a test, not a result. Second, p50 collapses to a dict lookup as soon as most requests hit, but p95 stays at the model's latency in every row, because more than 5 percent of requests miss in every row. A cache moves the median; it moves the tail only when the miss rate falls below the percentile you care about. Third, a TTL is a hit-rate tax paid for freshness: 600 seconds cost 15 points of hit rate here and 60 seconds cost 41, with nothing gained in return since the model never changed in this scenario.
What a hit costs when the cache is Redis
The in-process dict makes a hit nearly free (0.1 us). The earlier article's cache was Redis, and a Redis hit is a network round trip plus deserialization. Measured with redis_roundtrip.py against Redis 8.2.9 in Docker on the same laptop, 500 samples:
| Operation | p50 | p95 |
|---|---|---|
| model call | 508 us | 649 us |
Redis SET ... EX (a miss also pays this) | 425 us | 961 us |
Redis GET hit plus json.loads | 279 us | 353 us |
| dict hit | 0.1 us | 0.4 us |
With these numbers a hit saves 229 us and a miss costs 425 us more than not caching. Mean latency is below the no-cache latency only when the hit rate exceeds (933 minus 508) over (933 minus 279), about 65 percent. Below that, the Redis cache makes this model slower on average. These figures go through Docker Desktop's port mapping on macOS and would be lower with a native Redis or a Unix socket, but the shape of the argument does not depend on them: a remote cache in front of a sub-millisecond model needs a high hit rate to pay for itself, and the earlier article's one-second sleep hid that entirely. The Redis INFO stats counters keyspace_hits and keyspace_misses give you the hit rate; you have to measure the two latencies yourself.
A redeploy: the version belongs in the key, and it costs a cold start
At the midpoint of the simulated hour the model is replaced by a different weight seed, TTL 600 seconds. Without the version in the key, the cache keeps serving the old model's answers until each entry expires. With it, every key is new after the switch.
| Seed | Key | Hit rate | Stale hits | Model calls | Hit rate, first 5 min after switch |
|---|---|---|---|---|---|
| 0 | version not in key | 0.6928 | 283 | 3,072 | 0.7077 |
| 0 | version in key | 0.6838 | 0 | 3,162 | 0.5834 |
| 1 | version not in key | 0.6967 | 272 | 3,033 | 0.6741 |
| 1 | version in key | 0.6892 | 0 | 3,108 | 0.5728 |
Stale hits are hits after the switch on an entry the old model produced. The overall hit rates barely differ, which is the trap: a dashboard showing hit rate alone would not distinguish these two runs, and only one of them served 283 wrong answers. The price of correctness is the post-switch window, where the hit rate drops by 12 points while the popular keys are re-filled. If that window is unacceptable, warm the new version's keys from the old version's key list before cutting over; the lab does not do that, and a TTL alone does not help, since it shortens the stale window only by shortening every other window as well.
The stampede on a cold key
An exact-match cache with the earlier article's shape, "get, and if missing, compute and set", has a window between the get and the set during which every other request for the same key also misses. The lab starts 16 threads on a barrier against the same new key:
| Seed | Threads | Model calls, naive | Model calls, single-flight |
|---|---|---|---|
| 0 | 16 | 16 | 1 |
| 1 | 16 | 14 | 1 |
The naive count varies with scheduling (the test asserts only that it is greater than 1); the single-flight count is asserted to be exactly 1. In-process the fix is a lock per key:
def get_or_compute(self, key, compute):
value = self.store.get(key)
if value is not None:
return value
with self._meta:
lock = self.locks.setdefault(key, threading.Lock())
with lock:
value = self.store.get(key)
if value is None:
value = compute()
self.store[key] = value
return value
Across processes the same idea needs a shared lock or a "computing" marker in Redis with its own expiry, and the choice of whether waiters block or fall through to the model is a latency-versus-cost decision the article cannot make for you. What it can say is that every redeploy in the previous section is a stampede on every popular key at once.
Semantic cache: the hit rate is not the metric
An exact-match cache misses on any change to the input. A semantic cache embeds the input and serves the stored answer of the nearest stored embedding when the cosine similarity clears a threshold. The lab makes 2,000 inputs in 500 families of four siblings, each sibling a distinct input with its own answer. Every request's embedding is its input's vector plus a random direction of length 0.25 (a paraphrase); half of the requests are verbatim repeats and half carry a distinct phrasing field, so the exact key misses them although the correct answer is the same. Measured geometry (seed 0): a request and a stored paraphrase of the same input sit at cosine 0.941 (5th percentile 0.933); a request and the closest stored sibling sit at 0.765 (95th percentile 0.802). That gap is a parameter of the lab, not a property of real embeddings, and it is what makes the table below clean.
| Seed | Policy | Hit rate | False hits | Model calls | p50 | p95 | Lookup p50 at end |
|---|---|---|---|---|---|---|---|
| 0 | exact, same stream | 0.3942 | 0 | 6,058 | 416 us | 542 us | |
| 0 | semantic, threshold 0.70 | 0.9462 | 2,128 (21.3%) | 538 | 6.0 us | 455 us | 6.1 us |
| 0 | semantic, threshold 0.75 | 0.9072 | 1,051 (10.5%) | 928 | 8.5 us | 492 us | 9.2 us |
| 0 | semantic, threshold 0.80 | 0.8587 | 58 (0.6%) | 1,413 | 11.7 us | 513 us | 12.9 us |
| 0 | semantic, threshold 0.85 | 0.8555 | 0 | 1,445 | 12.6 us | 662 us | 13.7 us |
| 0 | semantic, threshold 0.90 | 0.8555 | 0 | 1,445 | 12.0 us | 569 us | 13.3 us |
| 0 | semantic, threshold 0.95 | 0.4606 | 0 | 5,394 | 531 us | 848 us | 41.5 us |
| 1 | exact, same stream | 0.3902 | 0 | 6,098 | 476 us | 841 us | |
| 1 | semantic, threshold 0.70 | 0.9460 | 2,130 (21.3%) | 540 | 6.0 us | 494 us | 6.1 us |
| 1 | semantic, threshold 0.80 | 0.8560 | 57 (0.6%) | 1,440 | 11.9 us | 615 us | 13.1 us |
| 1 | semantic, threshold 0.85 | 0.8529 | 0 | 1,471 | 12.0 us | 546 us | 13.3 us |
| 1 | semantic, threshold 0.95 | 0.4581 | 0 | 5,419 | 536 us | 887 us | 45.7 us |
A false hit is a hit whose stored answer came from a different input than the one requested: the user got somebody else's answer. Read the 0.70 row as a dashboard would: hit rate 95 percent, model calls down 95 percent, p50 6 us. Every one of those numbers is better than the exact cache's, and 21 percent of responses were wrong. Nothing in the hit rate, the cost saving, or the latency reveals it; only labeling each hit with the identity of the input that produced it does, which in production means sampling hits and re-running the model on them. At 0.85 and 0.90 the semantic cache serves no wrong answers and more than doubles the exact cache's hit rate (0.8555 versus 0.3942) because paraphrases now hit; at 0.95 it has thrown away most true hits as well, because paraphrases sit at 0.94 in this geometry. The workable window is between the sibling similarity and the paraphrase similarity, and its width is a property of your embedding model and your data that this lab cannot tell you.
The last column is the cost the threshold does not show: the lookup is a matrix product against every stored embedding, and it grows with the cache. It is 6 us at 538 entries and 41 us at 5,394 entries here; a real cache of a million entries needs a vector index, which trades exactness of the nearest-neighbor search for speed and adds its own recall error on top of the threshold's. Redis's redisvl library packages this as SemanticCache with a distance_threshold on cosine distance (0 identical, 2 opposite) and TTL support, and its guide says the right threshold "is not a fixed quantity" and depends on the embedding model, the inputs and the use case. That was checked against the documentation only; the lab's lookup is brute force.
Operational notes on the Redis side, checked
- Nothing bounds the cache by default. The
redis:8.2-alpineimage starts withmaxmemory 0andmaxmemory-policy noeviction(asserted by a test withCONFIG GET). The eviction documentation states that 0 means no limit on 64-bit systems, and thatvolatile-*policies "behave likenoevictionif no keys have an associated expiration". Setmaxmemory, then pick a policy; the earlier article namedvolatile-lruwithout setting the limit that makes any policy run. - Outage behavior is your code, not Redis's. With nothing listening, redis-py raises
redis.exceptions.ConnectionError("Connection refused") fromget(asserted by a test). A wrapper that promises to fall back to the model has to catchredis.exceptions.RedisErroraround both the get and the set, count the failures, and call the model; the earlier article promised it in prose and did not do it in code. SET key value EX seconds, notSETEX. redis-py 8.1.0 emitsDeprecationWarning: Call to deprecated setex. (Use 'set' instead.) -- Deprecated since version 2.6.12on every call; the Redis documentation describesSETEXas equivalent toSETwithEX. The lab usesr.set(key, value, ex=ttl).SCAN, notKEYS, to look at what is cached. TheKEYSdocumentation says "Don't useKEYSin your regular application code" and points atSCAN; the earlier article's verification step ranKEYS cache:model:*, which is fine for a one-off but is the command people copy into a metrics job.StrictRedisis a plain alias. redis-py 8.1.0'sclient.pycontainsStrictRedis = Redis; it works, and it signals code written for redis-py 2.x.- Memcached has TTLs too. Its protocol gives every storage command an
exptimein seconds (up to 30 days, beyond that a Unix timestamp) and it evicts by LRU. Reasons to choose Redis over it exist (data structures, persistence, the vector search used above); per-item expiry is not one of them.
What this lab is not
It is a matrix product behind a dict, on one laptop, with Gaussian clusters standing in for embeddings and a simulated clock standing in for an hour. It measures no real model, no real embedding model, no vector index, no Redis under memory pressure, and no serving engine's KV cache. Its value is that the four effects above, a p95 that a cache does not move, a hit rate that hides stale answers after a redeploy, a cold key that fans out into N model calls, and a semantic threshold that buys hit rate with wrong answers, are each reproduced by a script that runs in about 35 seconds and by tests that pin them.
Reproduce it
python3 -m unittest -v test_caching_lab.py # 22 tests (3 skipped without REDIS_URL)
python3 caching_lab.py --seed 0 --json results-seed0.json
python3 caching_lab.py --seed 1 --json results-seed1.json
docker run -d --name dd-redis-b -p 6381:6379 redis:8.2-alpine
REDIS_URL=redis://127.0.0.1:6381/0 python3 -m unittest -v test_caching_lab.py
REDIS_URL=redis://127.0.0.1:6381/0 python3 redis_roundtrip.py
Sources
- Redis: Key eviction for the
maxmemorydefault of 0, the policy list, thevolatile-*note, and thekeyspace_hits/keyspace_misseshit-rate formula. - Redis: KEYS for "Don't use
KEYSin your regular application code" and the pointer toSCAN. - Redis: SETEX for "This command is equivalent to
SET key value EX seconds". - redis-py 8.1.0 documentation; the
StrictRedis = Redisalias and thesetexdeprecation warning were observed in the installed package. - redisvl user guide: Semantic caching for LLMs for
SemanticCache,distance_threshold, cosine distance and the threshold note. - vLLM: Automatic prefix caching and vLLM design: Prefix caching for block hashing, the prefill-only effect and unchanged outputs.
- Memcached protocol for
exptimeand LRU. - TorchServe repository, archived 2025-08-07 with the notice "This project is no longer actively maintained."
- Python release status: 3.8 end-of-life 2024-10-07, 3.9 end-of-life 2025-10-31.
- NumPy 2.4 documentation
