-
Caching AI Model Inference: What a Hit Costs, When the Answer Goes Stale, and How a Semantic Cache Lies
A reproducible NumPy lab that puts an exact-match cache, a TTL, a model redeploy, a thread stampede and a cosine-similarity semantic cache in front of a deterministic model, measures hit rate, model calls and p50/p95 on a synthetic request stream, and shows the one number a semantic cache must report that a hit rate hides:…

-
API Key Rotation and Revocation: A Verifier That Stores Hashes, Overlaps Two Keys, and Cuts a Key Off on the Next Request
A tested Python 3.14 / Flask 3.1 key store: public id plus 256-bit secret, SHA-256 at rest, hmac.compare_digest, rotation with a grace window that both keys survive and one does not, revocation that beats the window, and a log that never sees the secret. Plus what the earlier version’s code did when run.

-
Kafka Client Quotas on 4.3.1: What Throttling Actually Does to a Producer, Measured
Byte-rate and request quotas on a Kafka 4.3.1 broker: the Admin API as it really exists, a 30 MB run that went from 0.36 s to 12.1 s with zero retries, the burst allowance nobody mentions, a user quota that silently ignored an unauthenticated client, and the JMX names that do exist.

-
Spring Boot 3 Cache Invalidation with Redis: @CacheEvict, @CachePut, TTL, Null Entries, and What Pub/Sub Is Actually For
Tested on Spring Boot 3.5.16 and Redis 7.4.11: the exact keys and TTLs Spring writes, why a null result can throw, why a second instance needs no pub/sub message to see an eviction, and the customizer ordering trap behind spring.cache.cache-names.

-
Cache-Aside with Caffeine: Which Read Path Prevents a Stampede, and What Expiry, Refresh, and Stats Really Do
Tested on Caffeine 3.2.4 and Java 17: the getIfPresent/put helper let 32 of 32 concurrent misses hit the database, cache.get(key, fn) let one through, and the published async example did not compile.

-
Compressing a Model for the Edge: What int8 and Pruning Actually Cost, Measured
A reproducible NumPy and onnxruntime lab that quantizes a small MLP to int8 per-tensor and per-channel, prunes it to 98 percent, and measures bytes, held-out accuracy and latency at every step, including the one weight rewrite that leaves the float model identical and drops int8 accuracy from 82 to 12 percent, and the reason a…

-
Backpressure, Rate Limiting and Load Shedding in an API Gateway: Three Different Things, Measured
A dependency-free Java gateway with a token bucket, a bounded queue and a downstream timeout, each measured under 50 concurrent requests; a JDK Flow API run that shows what backpressure actually is; and the published Node.js example run as printed, which answered 500 to every admitted request.

-
Kafka Producer Metrics in Java: What the Client Reports, How Micrometer Exposes It to Prometheus, and a Per-Record Latency Timer That Compiles
All 110 metrics a kafka-clients 4.3.1 producer reports, checked against the docs and JMX; the exact Prometheus text Micrometer’s KafkaClientMetrics produces; and a send-to-ack interceptor built on the headers-aware onAcknowledgement overload, replacing sample code that never compiled.
