Revision note (2026-09-15). The earlier version of this article had a UserCacheService example that does not compile (buildAsync(this::loadUserFromDbAsync) with a loader returning CompletableFuture<User>; javac rejects it on Caffeine 3.1.8 and 3.2.4 alike), described Caffeine as a "Java 8+" library although the 3.x line requires Java 11, pinned version 3.1.8 while 3.2.4 is the current release, showed a cache.stats() monitoring snippet on a cache built without recordStats() (every counter stays zero), presented the getIfPresent / load / put helper as the read path and only mentioned in passing that it lets concurrent misses through, and cited a Martin Fowler page that does not exist. This version is rebuilt around a small project with 15 tests; every number below comes from those tests.
The pattern in one paragraph
Cache-aside means the application owns the cache: on a read it checks the cache first, loads from the primary store on a miss and stores the result; on a write it updates the primary store and invalidates (or replaces) the cached entry. The store never populates the cache on its own. Microsoft's Azure Architecture Center has the standard description of the pattern, including its consistency limits; the link is in the sources.
Tested versions: Caffeine 3.2.4 (the release entry in its Maven Central metadata, published 2026-05-03), Java 17 (Amazon Corretto 17.0.14), Gradle 8.8, JUnit 5.10.2. Caffeine 3.x needs Java 11 or later; the 2.x line is the one for Java 8 (release notes for 3.0.0).
Three read paths, one stampede test
The lab's "database" is a map with a call counter and a 200 ms sleep. Thirty-two threads are released by a latch to read the same key that is not yet cached. The question is how many times the database is loaded for that one key.
The helper the earlier article presented:
User get(String id) {
User value = cache.getIfPresent(id);
if (value == null) {
value = database.load(id);
if (value != null) {
cache.put(id, value);
}
}
return value;
}
Result: 32 threads, 32 database loads. Every thread saw null, every thread loaded, every thread wrote. On a single thread the helper is fine; under concurrent misses it is the stampede.
The same logic with the mapping-function overload:
User get(String id) {
return cache.get(id, database::load);
}
Result: 32 threads, 1 database load. Cache.get(key, mappingFunction) computes the value at most once per key at a time; the other callers wait for that computation and reuse it. The Caffeine wiki says to prefer this form over getIfPresent plus put for exactly this reason.
A LoadingCache built with the loader (Caffeine.newBuilder().build(database::load) and cache.get(id)): 32 threads, 1 load. An AsyncLoadingCache (buildAsync(database::load) and cache.get(id)): 32 threads, 1 load, and all 32 callers received the same CompletableFuture instance.
| Read path | Database loads, 32 concurrent misses on one key |
|---|---|
getIfPresent, then load, then put | 32 |
Cache.get(key, database::load) | 1 |
LoadingCache.get(key) | 1 |
AsyncLoadingCache.get(key) | 1 |
The manual test asserts only "more than one" because a thread that starts late enough can find the value already cached; in both recorded runs the count was 32. Latency was not measured and is not the point; the load count is.
So the cache-aside read path in Caffeine is get(key, fn) or a loading cache. getIfPresent and put remain useful for the write side (replace after a write) and for lookups that must not trigger a load.
The async service, corrected
Caffeine.buildAsync has two overloads. One takes a CacheLoader<K, V>, a plain K -> V function that Caffeine runs on the cache's executor. The other takes an AsyncCacheLoader<K, V>, a (K, Executor) -> CompletableFuture<V> function. A one-argument method reference that returns CompletableFuture<User> matches neither as AsyncLoadingCache<String, User>, which is the compile error quoted in the lab README. The version that compiles and passes the read, update, invalidate, read test:
public final class AsyncUserCacheService {
private final FakeDatabase database;
private final AsyncLoadingCache<String, User> userCache;
public AsyncUserCacheService(FakeDatabase database, Executor executor) {
this.database = database;
this.userCache = Caffeine.newBuilder()
.maximumSize(10_000)
.expireAfterWrite(Duration.ofMinutes(15))
.refreshAfterWrite(Duration.ofMinutes(5))
.executor(executor)
.buildAsync((String id, Executor exec) ->
CompletableFuture.supplyAsync(() -> database.load(id), exec));
}
public CompletableFuture<User> getUser(String id) {
return userCache.get(id);
}
public void updateUser(User user) {
database.put(user);
userCache.synchronous().invalidate(user.id());
}
}
If you do not need control over where the load runs, buildAsync(database::load) with a loader that returns User is the shorter correct form.
Expiry and refresh are different things
All of these were run with a manual Ticker, so no test sleeps.
expireAfterWrite(10 min): the entry is present at 9 minutes, and gone at 10 minutes plus one nanosecond even though it was read at minute 9. expireAfterAccess(10 min): three reads at 9-minute intervals keep it alive; it is gone 10 minutes after the last read. Expiry blocks the next reader behind a fresh load.
refreshAfterWrite(5 min) does not remove anything. It makes the entry eligible for a reload, and the reload is started by the next read of that key, not by a timer. In the test the reload was blocked on a latch: the read that triggered it returned the old value immediately, every read while the reload was in flight returned the old value, exactly one reload was started, and once the latch opened the new value appeared. Two entries were made eligible and only the one that was read again was reloaded; the other was later removed by expireAfterWrite(15 min). A reload that throws keeps the old value (Caffeine logs the exception). getIfPresent triggers a refresh as well, not only get.
One detail that matters for tests: with a same-thread executor (Runnable::run) the reload finishes inside the triggering call, so that call already returns the new value. The "old value while refreshing" behavior is only visible with a real executor, which is the default (ForkJoinPool.commonPool()) unless you set executor(...).
So the combination the earlier article used, expireAfterWrite(15 min) plus refreshAfterWrite(5 min), means: hot keys are reloaded in the background and readers never wait, cold keys are not reloaded and simply expire.
Statistics need recordStats()
The earlier article's monitoring snippet was:
CacheStats stats = cache.stats();
System.out.println("Hit rate: " + stats.hitRate());
On a cache built without recordStats(), after 10 real requests, stats() returned hitCount=0, missCount=0, loadSuccessCount=0 and hitRate() returned 1.0, because zero hits out of zero requests is defined as 1.0. That is the worst kind of dashboard: a perfect hit rate on a cache that records nothing. With recordStats() the same 10 requests gave 9 hits, 1 miss, 1 successful load, hit rate 0.9.
Size eviction is enforced by maintenance
maximumSize(100), then 1,000 sequential put calls: estimatedSize() peaked at 1,000 and was still 1,000 right after the loop. After cleanUp() it was 100 and evictionCount() was 900. Caffeine runs eviction as a maintenance step, usually asynchronously on the executor, so a bound is not a hard limit at every instant. Do not write tests that assert estimatedSize() <= maximumSize without cleanUp(), and do not size the cache as if the bound were exact.
Which entries are evicted is decided by Caffeine's admission and eviction policy, which is based on TinyLFU (the project README links the papers). This lab did not measure hit-rate quality against any workload, and this article makes no claim about it.
Nulls
A loader that returns null (the row does not exist) gets null back from get, nothing is stored, and every later read of that key goes to the database again. CacheStats counts a null result under loadFailureCount, as its Javadoc says. cache.put(key, null) throws NullPointerException. So the earlier advice "avoid caching nulls" is not a choice you get to make with Caffeine: it does not store them. If absent rows are hot, cache a sentinel or an Optional deliberately, with a shorter expiry than real rows.
What this does not cover
- A real database, connection pool, or network; the lab's store is a map that sleeps.
- Hit-rate quality of the eviction policy, weighted eviction (
maximumWeight), customExpiry, reference-based eviction, removal listeners. - Metrics integrations (Micrometer, Dropwizard, Prometheus); the wiki's Statistics page lists them.
- Consistency across instances. Caffeine is in-process; each JVM has its own cache and its own invalidations.
Reproduce it
gradle test --no-daemon --no-build-cache --rerun-tasks --console=plain # 15 tests
The lab is examples/java-caffeine-cache-aside in the site's repository; the README holds the full test output and the javac output for the earlier example.
