Revision note (2026-09-15). The earlier version of this article said virtual threads were "still incubating", required "JDK 19 or later with preview features enabled", should be compiled with --enable-preview --release 19, and that production users should wait for an LTS release. Virtual threads were a preview in JDK 19 (JEP 425) and are final in JDK 21 (JEP 444). It also recommended -Xss for virtual thread stacks, which JEP 444 says live in the heap, and never mentioned pinning, the one JDK 21 behavior that turns "blocking code scales" into a stall. Every number below comes from the lab run described at the end; the earlier version had none.
Environment
- JDK 21.0.3 (Amazon Corretto), Gradle 8.8, no
--enable-preview - Apple M3 Pro, 11 cores (
Runtime.availableProcessors() = 11), 18 GB, macOS; loopback networking only - Per-process thread limit on this Mac:
kern.num_taskthreads = 4096 - Other builds ran on the same machine during the test run, so timings are indicative. Two full runs gave the same ordering with different absolute values; both are in the lab README.
The echo server, unchanged in spirit
The server accepts on a platform thread and hands every connection to an executor. With Executors.newVirtualThreadPerTaskExecutor() each connection gets its own virtual thread and the handler is written as blocking code:
public final class EchoServer implements AutoCloseable {
private final ServerSocket serverSocket;
private final ExecutorService executor;
private final Thread acceptThread;
public EchoServer(int port, ExecutorService executor) throws IOException {
this.serverSocket = new ServerSocket(port);
this.executor = executor;
this.acceptThread = Thread.ofPlatform().name("echo-accept").daemon(true).unstarted(this::acceptLoop);
}
private void acceptLoop() {
while (!serverSocket.isClosed()) {
try {
Socket client = serverSocket.accept();
executor.submit(() -> handleClient(client));
} catch (SocketException closed) {
return;
} catch (IOException e) {
System.err.println("accept failed: " + e.getMessage());
}
}
}
static void handleClient(Socket socket) {
try (socket;
BufferedReader in = new BufferedReader(new InputStreamReader(socket.getInputStream()));
PrintWriter out = new PrintWriter(socket.getOutputStream(), true)) {
out.println("Welcome to the Virtual Thread Echo Server!");
String line;
while ((line = in.readLine()) != null) {
if (line.equalsIgnoreCase("bye")) {
out.println("Goodbye!");
break;
}
out.println("Echo: " + line);
}
} catch (IOException e) {
System.err.println("Error handling client: " + e.getMessage());
}
}
}
Compile and run it like any other Java 21 program. gradle run --args='8093' then:
$ printf "hellonworldnbyen" | nc localhost 8093
Welcome to the Virtual Thread Echo Server!
Echo: hello
Echo: world
Goodbye!
2,000 concurrent connections
The test connects 2,000 client sockets before any of them sends a byte, then every client sends 20 lines and checks 20 echoes (40,000 round trips), then bye. A sampler thread records the peak number of platform threads in the JVM during the run.
virtual-thread server: 2000 clients x 20 messages in 817 ms; platform threads before=26 peak=28; carrier threads=11; availableProcessors=11
platform-thread-per-connection server: 2000 clients x 20 messages in 614 ms
Two things to read from this:
- Serving 2,000 open connections on virtual threads added two platform threads to the JVM. The carrier pool had 11 threads, one per core, which is the default (
carrierPoolSizeDefaultsToAvailableProcessorsobserved at most 11 carriers while 1,000 virtual threads slept). - The platform-thread-per-connection server was faster on this loopback echo: 614 ms against 817 ms (run 1: 438 ms against 748 ms). Virtual threads did not make the server faster and nobody should expect them to. Each round trip costs the same syscalls either way; what changes is how many blocked handlers you can afford to have at once.
Where the win is, measured
10,000 tasks that each block for 10 ms, on a fixed pool of 200 platform threads and on a virtual thread per task:
10,000 x sleep(10ms): fixed pool of 200 platform threads = 587 ms, virtual thread per task = 99 ms
The pool cannot do better than 10000 / 200 x 10 ms = 500 ms; it queues. Virtual threads run all 10,000 sleeps concurrently on 11 carriers because a sleeping virtual thread is unmounted and costs no carrier. The same 10,000 tasks made CPU-bound instead:
10,000 x CPU-bound task (11 cores): warm-up 527 ms, fixed pool of 11 = 406 ms, virtual thread per task = 404 ms
No difference. There are 11 cores and virtual threads do not add any.
Where the ceiling is: platform threads on this machine
A child JVM (-Xmx512m) starts N threads that all park on a latch, then reports the platform thread count and the process RSS from ps:
kind=platform count=2000 startedMs=117 platformThreadsAlive=2007 rssBeforeKb=45968 rssAfterKb=201232 rssDeltaMb=151
kind=virtual count=10000 startedMs=32 platformThreadsAlive=18 rssBeforeKb=46016 rssAfterKb=93200 rssDeltaMb=46
kind=virtual count=100000 startedMs=300 platformThreadsAlive=18 rssBeforeKb=45152 rssAfterKb=278320 rssDeltaMb=227
Asking for 10,000 platform threads failed at thread number 4,069:
[0.450s][warning][os,thread] Failed to start thread "Unknown thread" - pthread_create failed (EAGAIN) for attributes: stacksize: 2048k, guardsize: 16k, detached.
Exception in thread "main" java.lang.OutOfMemoryError: unable to create native thread: possibly out of memory or process/resource limits reached
That is the macOS per-process limit (kern.num_taskthreads = 4096), not a JVM limit; Linux boxes have different limits. 100,000 virtual threads started in 300 ms with 18 platform threads alive. RSS deltas include everything the JVM touched, so read them as orders of magnitude, not per-thread costs.
Pinning: the JDK 21 rule you have to know
JEP 444 lists two situations in which a virtual thread cannot be unmounted while blocking: inside a synchronized block or method, and inside a native method or foreign function. The thread is then pinned to its carrier, and the carrier is blocked with it.
The lab makes this visible with exactly one carrier (-Djdk.virtualThreadScheduler.parallelism=1 -Djdk.virtualThreadScheduler.maxPoolSize=1) and the JDK's own pinning trace (-Djdk.tracePinnedThreads=short). Thread A takes a lock and sleeps 300 ms; thread B is started right after A holds the lock and records when it got to run:
--- synchronized ---
VirtualThread[#20,A]/runnable@ForkJoinPool-1-worker-1 reason:MONITOR
com.devdrunk.vthreads.PinningDemo.lambda$main$0(PinningDemo.java:32) <== monitors:1
mode=synchronized B started after 302 ms
--- reentrantlock ---
mode=reentrantlock B started after 0 ms
With synchronized, B waited for the whole sleep although it never touched the lock: the only carrier was held by a pinned, sleeping A. With ReentrantLock, A unmounted while sleeping and B ran immediately. Production JVMs have as many carriers as cores, so one pinned thread costs one core's worth of scheduling, not the whole application; a synchronized block around a slow I/O call in a hot path costs more. jdk.tracePinnedThreads prints the stack of every pinned park, which is the way to find them.
JEP 491 (JDK 24) changes this: blocking to acquire a monitor unmounts the virtual thread. This lab ran on JDK 21, where the pinning is real.
Facts about the executor
Executors.newVirtualThreadPerTaskExecutor() threads report isVirtual() == true, are always daemon threads, and have an empty name unless you set one through a Thread.Builder. The lab pins all three in virtualThreadFactsFromTheExecutor. Empty names matter for thread dumps and logs: use Thread.ofVirtual().name("conn-", 0).factory() with Executors.newThreadPerTaskExecutor if you want numbered names.
What this does not cover
- Stack size. JEP 444 states that virtual thread stacks are stored in the heap as stack-chunk objects. What
-Xssdoes or does not do to them was not measured here, so this article makes no claim about it. - The scheduler's compensation (
jdk.virtualThreadScheduler.maxPoolSize, default 256) for blocking file I/O orObject.wait(). The pinning demo caps the pool at 1 on purpose. - Anything beyond loopback: no TLS, no real protocol, no slow clients.
- Memory per thread. Only process RSS deltas were recorded.
- JDK 24+ behavior; the JDK here is 21.0.3.
Reproduce it
gradle test --no-daemon # 13 tests, prints every number above
gradle run --no-daemon --args='8093'
printf "hellonbyen" | nc localhost 8093
The Gradle toolchain asks for Java 21; Gradle auto-detects SDKMAN installations.
