Caching AI Model Inference: What a Hit Costs, When the Answer Goes Stale, and How a Semantic Cache Lies

Lab run with Python 3.14.7, NumPy 2.4.4, redis-py 8.1.0 and Redis 8.2.9 (redis:8.2-alpine in Docker) for seeds 0 and 1, 22 unit tests; vLLM prefix caching, redisvl SemanticCache, Memcached and TorchServe claims checked against documentation only, not run

Revision note (2026-09-15). The earlier version of this article was a Redis setex wrapper around a time.sleep(1) stand-in model, and its only measurement was that the second call returned "almost instantly", which is true of anything compared with a one-second sleep. It said caching helps with "very similar" requests while implementing a SHA-256 exact-match key that misses on any change; it advised "fallback to direct inference" in the troubleshooting section while the code had no exception handling, so a Redis outage takes the endpoint down (redis-py raises redis.exceptions.ConnectionError from get); it relied on volatile-lru eviction without setting maxmemory, whose default is 0 (no limit); it listed "TTLs" and eviction as reasons to prefer Redis over Memcached, which has both; its setex call draws a DeprecationWarning from redis-py 8.1.0; it required Python 3.8+ (end-of-life since October 2024) and Redis 6+; it never mentioned the request stampede that an exact-match cache creates on every cold key; and its four sources had no links, one of them TorchServe, whose repository was archived in August 2025. This version replaces the wrapper with a seeded simulation in which every number below is produced by the code shown. The workload is synthetic; nothing here is a benchmark of a real model.

Three things called "model caching"

The phrase covers at least three mechanisms that share nothing but the word:

  • Response caching: store the model's output under a key derived from the input, serve it again on an identical input. This is what the earlier article built and what the lab measures first.
  • Semantic caching: embed the input, store the output under the embedding, serve it again for any input whose embedding is within a similarity threshold. The lab measures this second, including the rate at which it serves a wrong answer.
  • Prefix or KV caching inside the serving engine: for transformer models, reuse the attention key/value tensors already computed for a shared prompt prefix. vLLM's documentation describes hashing each KV block by the tokens in the block plus the tokens of the prefix before it, notes that this reduces the prefill phase only, not token generation, and states that it does not change model outputs. It lives inside the engine, keyed on tokens, and is not something you build with Redis. It was checked against the documentation only; the lab does not run an LLM.

Weight caching (keeping model files on local disk or in memory between cold starts) is a deployment concern, not a request-path cache, and is out of scope here.

Tested versions: Python 3.14.7, NumPy 2.4.4, redis-py 8.1.0, Redis 8.2.9 from the redis:8.2-alpine image in Docker 27. The lab is examples/ai-model-caching in the companion repository; python3 -m unittest -v test_caching_lab.py runs 22 tests (3 need REDIS_URL) and python3 caching_lab.py --seed 0 --json results-seed0.json reproduces the tables.

The lab

The "model" is two dense 2048 by 2048 float32 layers with tanh in between, applied to a 16-float payload: a real matrix-vector product that costs about half a millisecond on the laptop used, deterministic for a given weight seed, so a version change is a different weight seed. The request stream draws from 2,000 distinct inputs with Zipf popularity (exponent 1.0) over one simulated hour, and every request carries its own simulated arrival time so TTL expiry is deterministic and does not depend on how fast the script runs. The key is the earlier article's recipe:

def cache_key(payload: dict, model_version: str | None) -> str:
    canonical = json.dumps(payload, sort_keys=True, separators=(",", ":"))
    raw = canonical if model_version is None else f"{model_version}:{canonical}"
    return "cache:model:" + hashlib.sha256(raw.encode("utf-8")).hexdigest()

Tests pin what this key does and does not normalize: key order does not change it, but {"x": 1} and {"x": 1.0} are different keys (json.dumps writes 1 and 1.0), list order changes it, and the model version changes it. Two clients that format floats differently will never share a cache entry.

The exact cache is a dict with the semantics SET key value EX seconds gives you: an entry is gone once now >= stored_at + ttl. It has no size bound, on purpose, because neither did the earlier article's Redis (see the operational notes below).

Exact-match cache: the hit rate belongs to the traffic, not the cache

10,000 requests, 2,000 possible inputs, Zipf 1.0, one simulated hour. Seed 0 saw 1,461 distinct inputs, seed 1 saw 1,458.

SeedPolicyHit rateModel callsp50p95
0no cache0.000010,000536 us963 us
0exact, no TTL0.85391,4614.9 us532 us
0exact, TTL 600 s0.69893,0116.7 us655 us
0exact, TTL 60 s0.44335,567500 us762 us
1no cache0.000010,000504 us641 us
1exact, no TTL0.85421,4584.8 us531 us
1exact, TTL 600 s0.70102,9905.7 us561 us
1exact, TTL 60 s0.43665,634482 us675 us

Three things to read off this table. The no-TTL hit rate is exactly one minus distinct-inputs over requests; the cache did not create it, the traffic did. With a fixed repeat rate instead of Zipf the hit rate is the repeat rate (seed 0: repeat rate 0.2, 0.5, 0.8 gave hit rates 0.2012, 0.4960, 0.8058), which is a test, not a result. Second, p50 collapses to a dict lookup as soon as most requests hit, but p95 stays at the model's latency in every row, because more than 5 percent of requests miss in every row. A cache moves the median; it moves the tail only when the miss rate falls below the percentile you care about. Third, a TTL is a hit-rate tax paid for freshness: 600 seconds cost 15 points of hit rate here and 60 seconds cost 41, with nothing gained in return since the model never changed in this scenario.

What a hit costs when the cache is Redis

The in-process dict makes a hit nearly free (0.1 us). The earlier article's cache was Redis, and a Redis hit is a network round trip plus deserialization. Measured with redis_roundtrip.py against Redis 8.2.9 in Docker on the same laptop, 500 samples:

Operationp50p95
model call508 us649 us
Redis SET ... EX (a miss also pays this)425 us961 us
Redis GET hit plus json.loads279 us353 us
dict hit0.1 us0.4 us

With these numbers a hit saves 229 us and a miss costs 425 us more than not caching. Mean latency is below the no-cache latency only when the hit rate exceeds (933 minus 508) over (933 minus 279), about 65 percent. Below that, the Redis cache makes this model slower on average. These figures go through Docker Desktop's port mapping on macOS and would be lower with a native Redis or a Unix socket, but the shape of the argument does not depend on them: a remote cache in front of a sub-millisecond model needs a high hit rate to pay for itself, and the earlier article's one-second sleep hid that entirely. The Redis INFO stats counters keyspace_hits and keyspace_misses give you the hit rate; you have to measure the two latencies yourself.

A redeploy: the version belongs in the key, and it costs a cold start

At the midpoint of the simulated hour the model is replaced by a different weight seed, TTL 600 seconds. Without the version in the key, the cache keeps serving the old model's answers until each entry expires. With it, every key is new after the switch.

SeedKeyHit rateStale hitsModel callsHit rate, first 5 min after switch
0version not in key0.69282833,0720.7077
0version in key0.683803,1620.5834
1version not in key0.69672723,0330.6741
1version in key0.689203,1080.5728

Stale hits are hits after the switch on an entry the old model produced. The overall hit rates barely differ, which is the trap: a dashboard showing hit rate alone would not distinguish these two runs, and only one of them served 283 wrong answers. The price of correctness is the post-switch window, where the hit rate drops by 12 points while the popular keys are re-filled. If that window is unacceptable, warm the new version's keys from the old version's key list before cutting over; the lab does not do that, and a TTL alone does not help, since it shortens the stale window only by shortening every other window as well.

The stampede on a cold key

An exact-match cache with the earlier article's shape, "get, and if missing, compute and set", has a window between the get and the set during which every other request for the same key also misses. The lab starts 16 threads on a barrier against the same new key:

SeedThreadsModel calls, naiveModel calls, single-flight
016161
116141

The naive count varies with scheduling (the test asserts only that it is greater than 1); the single-flight count is asserted to be exactly 1. In-process the fix is a lock per key:

def get_or_compute(self, key, compute):
    value = self.store.get(key)
    if value is not None:
        return value
    with self._meta:
        lock = self.locks.setdefault(key, threading.Lock())
    with lock:
        value = self.store.get(key)
        if value is None:
            value = compute()
            self.store[key] = value
        return value

Across processes the same idea needs a shared lock or a "computing" marker in Redis with its own expiry, and the choice of whether waiters block or fall through to the model is a latency-versus-cost decision the article cannot make for you. What it can say is that every redeploy in the previous section is a stampede on every popular key at once.

Semantic cache: the hit rate is not the metric

An exact-match cache misses on any change to the input. A semantic cache embeds the input and serves the stored answer of the nearest stored embedding when the cosine similarity clears a threshold. The lab makes 2,000 inputs in 500 families of four siblings, each sibling a distinct input with its own answer. Every request's embedding is its input's vector plus a random direction of length 0.25 (a paraphrase); half of the requests are verbatim repeats and half carry a distinct phrasing field, so the exact key misses them although the correct answer is the same. Measured geometry (seed 0): a request and a stored paraphrase of the same input sit at cosine 0.941 (5th percentile 0.933); a request and the closest stored sibling sit at 0.765 (95th percentile 0.802). That gap is a parameter of the lab, not a property of real embeddings, and it is what makes the table below clean.

SeedPolicyHit rateFalse hitsModel callsp50p95Lookup p50 at end
0exact, same stream0.394206,058416 us542 us
0semantic, threshold 0.700.94622,128 (21.3%)5386.0 us455 us6.1 us
0semantic, threshold 0.750.90721,051 (10.5%)9288.5 us492 us9.2 us
0semantic, threshold 0.800.858758 (0.6%)1,41311.7 us513 us12.9 us
0semantic, threshold 0.850.855501,44512.6 us662 us13.7 us
0semantic, threshold 0.900.855501,44512.0 us569 us13.3 us
0semantic, threshold 0.950.460605,394531 us848 us41.5 us
1exact, same stream0.390206,098476 us841 us
1semantic, threshold 0.700.94602,130 (21.3%)5406.0 us494 us6.1 us
1semantic, threshold 0.800.856057 (0.6%)1,44011.9 us615 us13.1 us
1semantic, threshold 0.850.852901,47112.0 us546 us13.3 us
1semantic, threshold 0.950.458105,419536 us887 us45.7 us

A false hit is a hit whose stored answer came from a different input than the one requested: the user got somebody else's answer. Read the 0.70 row as a dashboard would: hit rate 95 percent, model calls down 95 percent, p50 6 us. Every one of those numbers is better than the exact cache's, and 21 percent of responses were wrong. Nothing in the hit rate, the cost saving, or the latency reveals it; only labeling each hit with the identity of the input that produced it does, which in production means sampling hits and re-running the model on them. At 0.85 and 0.90 the semantic cache serves no wrong answers and more than doubles the exact cache's hit rate (0.8555 versus 0.3942) because paraphrases now hit; at 0.95 it has thrown away most true hits as well, because paraphrases sit at 0.94 in this geometry. The workable window is between the sibling similarity and the paraphrase similarity, and its width is a property of your embedding model and your data that this lab cannot tell you.

The last column is the cost the threshold does not show: the lookup is a matrix product against every stored embedding, and it grows with the cache. It is 6 us at 538 entries and 41 us at 5,394 entries here; a real cache of a million entries needs a vector index, which trades exactness of the nearest-neighbor search for speed and adds its own recall error on top of the threshold's. Redis's redisvl library packages this as SemanticCache with a distance_threshold on cosine distance (0 identical, 2 opposite) and TTL support, and its guide says the right threshold "is not a fixed quantity" and depends on the embedding model, the inputs and the use case. That was checked against the documentation only; the lab's lookup is brute force.

Operational notes on the Redis side, checked

  • Nothing bounds the cache by default. The redis:8.2-alpine image starts with maxmemory 0 and maxmemory-policy noeviction (asserted by a test with CONFIG GET). The eviction documentation states that 0 means no limit on 64-bit systems, and that volatile-* policies "behave like noeviction if no keys have an associated expiration". Set maxmemory, then pick a policy; the earlier article named volatile-lru without setting the limit that makes any policy run.
  • Outage behavior is your code, not Redis's. With nothing listening, redis-py raises redis.exceptions.ConnectionError ("Connection refused") from get (asserted by a test). A wrapper that promises to fall back to the model has to catch redis.exceptions.RedisError around both the get and the set, count the failures, and call the model; the earlier article promised it in prose and did not do it in code.
  • SET key value EX seconds, not SETEX. redis-py 8.1.0 emits DeprecationWarning: Call to deprecated setex. (Use 'set' instead.) -- Deprecated since version 2.6.12 on every call; the Redis documentation describes SETEX as equivalent to SET with EX. The lab uses r.set(key, value, ex=ttl).
  • SCAN, not KEYS, to look at what is cached. The KEYS documentation says "Don't use KEYS in your regular application code" and points at SCAN; the earlier article's verification step ran KEYS cache:model:*, which is fine for a one-off but is the command people copy into a metrics job.
  • StrictRedis is a plain alias. redis-py 8.1.0's client.py contains StrictRedis = Redis; it works, and it signals code written for redis-py 2.x.
  • Memcached has TTLs too. Its protocol gives every storage command an exptime in seconds (up to 30 days, beyond that a Unix timestamp) and it evicts by LRU. Reasons to choose Redis over it exist (data structures, persistence, the vector search used above); per-item expiry is not one of them.

What this lab is not

It is a matrix product behind a dict, on one laptop, with Gaussian clusters standing in for embeddings and a simulated clock standing in for an hour. It measures no real model, no real embedding model, no vector index, no Redis under memory pressure, and no serving engine's KV cache. Its value is that the four effects above, a p95 that a cache does not move, a hit rate that hides stale answers after a redeploy, a cold key that fans out into N model calls, and a semantic threshold that buys hit rate with wrong answers, are each reproduced by a script that runs in about 35 seconds and by tests that pin them.

Reproduce it

python3 -m unittest -v test_caching_lab.py                      # 22 tests (3 skipped without REDIS_URL)
python3 caching_lab.py --seed 0 --json results-seed0.json
python3 caching_lab.py --seed 1 --json results-seed1.json
docker run -d --name dd-redis-b -p 6381:6379 redis:8.2-alpine
REDIS_URL=redis://127.0.0.1:6381/0 python3 -m unittest -v test_caching_lab.py
REDIS_URL=redis://127.0.0.1:6381/0 python3 redis_roundtrip.py

Sources