Revision note (2026-09-15). The earlier version of this article defined backpressure as an umbrella for rate limiting, queueing and circuit breaking, measured nothing, and shipped one Node.js/Express example. That example, run exactly as printed, answered HTTP 500 to every request that got past its token bucket (TypeError: next is not a function); with the one broken line fixed, its priority queue never held more than one element and its Server Busy branch never ran. Two of its six references now return 404. This version keeps the topic and replaces the content with a small Java gateway whose three admission controls are measured one at a time, plus a run of the JDK Flow API that shows what the word backpressure means when it is used precisely.
Three words that are not synonyms
- Backpressure is a signal from the consumer to the producer about how much it may send. In Reactive Streams it is
Subscription.request(n); in TCP it is the receive window; in HTTP/2 it is stream and connection flow control. The producer slows down, buffers, or drops; the consumer decides the pace. - Rate limiting is a policy decided in advance: at most N requests per second per key, enforced before any work is attempted. The answer to the excess is
429 Too Many Requests, ideally withRetry-After. - Load shedding (with a concurrency limit) is a bound on work in progress plus waiting. Whatever does not fit is rejected immediately with
503 Service Unavailable, again withRetry-After.
An HTTP/1.1 gateway has no backpressure channel to its clients except not reading from the socket. A 429 or a 503 is a rejection, not a demand signal. That distinction matters because the three mechanisms fail differently: a rate limit protects against a known budget being exceeded, a concurrency limit protects the downstream from whatever arrives, and only backpressure keeps a producer from ever generating what cannot be consumed.
The lab
Everything below was run on Java 17 (Amazon Corretto 17.0.14) with Gradle 8.8 and JUnit 5.10.2, on one laptop. The gateway uses only the JDK: com.sun.net.httpserver.HttpServer to accept, java.net.http.HttpClient to forward, a ThreadPoolExecutor as the concurrency limit, and java.util.concurrent.Flow for the backpressure demonstration. There is no framework, so every number is the mechanism's, not a library's. The code is in the repository under examples/api-gateway-backpressure.
The protected service is SlowDownstream on port 8103: GET /work?ms=N sleeps N milliseconds and answers 200, accepts unlimited concurrency on purpose, and reports maxInFlight on /stats. The gateway on 8102 applies, in order:
- a token bucket (capacity, refill per second) on the acceptor thread; no token,
429withRetry-After; - a fixed pool of
workersthreads in front of anArrayBlockingQueue(queueCapacity); when the queue is full,RejectedExecutionExceptionbecomes503 queue fullwithRetry-After: 1, decided on the acceptor thread without waiting; - a queue-wait deadline: when a worker picks a request up, if it waited longer than
maxQueueWaitMillisit gets503 queue wait exceededinstead of being forwarded late; - a per-request downstream timeout on the
HttpClientcall;504when it fires.
The core of the admission path:
private void admit(HttpExchange exchange) throws IOException {
if (!bucket.tryAcquire()) {
long retryAfterSeconds = Math.max(1, (bucket.millisUntilNextToken() + 999) / 1000);
exchange.getResponseHeaders().set("Retry-After", Long.toString(retryAfterSeconds));
Http.respond(exchange, 429, "rate limitedn");
return;
}
long enqueuedNanos = System.nanoTime();
try {
workers.execute(() -> process(exchange, enqueuedNanos));
} catch (RejectedExecutionException ex) {
exchange.getResponseHeaders().set("Retry-After", "1");
Http.respond(exchange, 503, "queue fulln");
}
}
workers is new ThreadPoolExecutor(n, n, 0, MILLISECONDS, new ArrayBlockingQueue<>(queueCapacity), new AbortPolicy()); with queueCapacity 0 it is a SynchronousQueue, so nothing waits at all. Every integration test fires 50 concurrent GET /work?ms=200 through the gateway with the rate limiter effectively off, so one control is under test at a time.
Bounded queue: shed early, cap the downstream
Four workers, queue of 8, 50 concurrent requests of 200 ms each:
12 x 200 (latency ms min=230 p50=439 p99=644 max=644)
38 x 503 (latency ms min=65 p50=70 p99=79 max=79)
gateway {"received":50,"shed":38,"forwarded":12,"maxQueueDepth":8}
downstream {"maxInFlight":4,"started":12,"completed":12}
Four in flight plus eight waiting is twelve served, in three rounds of about 200, 400 and 600 ms. The other 38 were told no within 79 ms, before a single unit of work had finished. The downstream never saw more than four concurrent requests, which is the whole point: its protection came from the gateway, not from itself.
With the queue at zero the shape is the same, only tighter: 4 served at 214 ms, 46 rejected with a p50 of 5 ms.
With a queue of 64 nobody is rejected:
50 x 200 (latency ms min=213 p50=1455 p99=2689 max=2689)
gateway {"shed":0,"forwarded":50,"maxQueueDepth":46}
downstream {"maxInFlight":4,"started":50,"completed":50}
Same downstream protection (maxInFlight 4), but 50 requests through 4 workers is 13 rounds, so p50 is 1.45 s and the last client waited 2.7 s. The queue did not remove the cost; it converted rejections into latency. Whether that is right depends on whether the client is still there at 2.7 s.
The queue-wait deadline, or rejecting late
A deadline on queue wait sounds like a way to have the big queue without the latency. Queue 64, deadline 500 ms:
12 x 200 (latency ms min=205 p50=411 p99=625 max=625)
38 x 503 (latency ms min=614 p50=620 p99=624 max=624)
gateway {"shed":0,"queueTimedOut":38,"forwarded":12,"maxQueueDepth":46}
The same 12 were served as with queue 8. The 38 rejections, though, arrived after about 620 ms of waiting, because the deadline is only checked when a worker finally picks the request up. That is the worst of both: the client waited and then got nothing, and the queue held 46 entries the whole time. If the wait budget is known, size the queue so it cannot be exceeded and shed on arrival instead. With workers W, budget B and service time S, the longest wait for the last queued request is about queueCapacity x S / W, so queueCapacity <= W x B / S; here 4 x 500 / 200 = 10. The queue-of-8 run above is that configuration, and its rejections took 70 ms.
Timeouts protect the caller, not the callee
Four workers, downstream timeout 300 ms, four concurrent requests of 1000 ms:
4 x 504 (latency ms min=307 p50=307 p99=307 max=307)
downstream after 1 s: {"inFlight":0,"maxInFlight":4,"started":4,"completed":4}
Every client got 504 at 307 ms. One second later the downstream reported all four units of work completed. The timeout freed the gateway's worker thread and the client's wait; it did not stop the work, and a client that retries on 504 doubles the load on the thing that was already slow. A timeout is part of a concurrency limit's accounting, not a substitute for one.
Rate limiting: the token bucket
Bucket of 20 with 10 tokens per second refill, 60 requests fired at once through a single HTTP client:
20 x 200 (latency ms min=5 p50=8 p99=13 max=13)
40 x 429 (latency ms min=3 p50=5 p99=8 max=8)
gateway {"received":60,"rateLimited":40,"forwarded":20}
Exactly 20 admitted, 40 rejected with Retry-After of at least 1 second, and the downstream saw exactly 20. A separate unit test drives a bucket of 100 with 50 per second refill in a tight loop for two seconds: 199 admitted out of 71,268,283 attempts (100 burst plus two seconds of refill; the missing one is the final partial token).
From the curl session, the same bucket configuration under seq 1 150 | xargs -P 150 curl admitted 119 and rejected 31, not 100 and 50. Starting 150 curl processes takes long enough for the bucket to refill about 20 tokens meanwhile, and a second run of the same command admitted 124. The arrival rate was set by process start-up, not by the limiter. When a rate limiter looks wrong in a load test, check what the load tool's arrival pattern really was before touching the limiter.
The bucket in the lab is one object in one process. A gateway with several instances needs either a per-instance budget (divide the global limit) or a shared counter, and the second has its own failure modes that this lab does not cover.
What backpressure looks like when it is real
SubmissionPublisher (JDK 9+) implements Reactive Streams demand. A subscriber that requests one item at a time and takes 10 ms per item, a buffer of 16, a producer with 100 items:
for (int i = 0; i < ITEMS; i++) {
publisher.submit(i); // blocks while the subscriber's buffer (16) is full
}
Flow submit(): producer blocked for 1018 ms to hand over 100 items (buffer 16, consumer 10 ms/item); all consumed after 1227 ms
The producer's loop took about a second because after the first 16 items it could only proceed as fast as the subscriber called request(1). Nothing was lost. The same producer using offer() with a drop handler:
Flow offer(): producer finished in 0 ms; 83 of 100 items dropped, 17 delivered
offer() never blocks; with no demand it drops, and the producer finished instantly having thrown away 83 items. That is the choice backpressure gives you: slow the producer or drop, decided by the consumer's demand. The gateway's 429 and 503 are the second option imposed on a client that never expressed demand at all; the first option does not exist in HTTP/1.1 above the socket.
What the earlier Node.js example actually did
The previous version's Express gateway was run byte for byte on Node.js 22.22.2 with Express 5.2.1, changed only by adding counters, a /__stats route and moving the port. Three bursts of concurrent GET /downstream:
N=200 500 x 101 429 x 99 stats {"enqueues":101,"maxQueueLength":1,"serverBusy":0}
N=200 500 x 9 429 x 191
N=300 500 x 103 429 x 197 stats {"enqueues":213,"maxQueueLength":1,"serverBusy":0}
server log: 213 x TypeError: next is not a function at processQueue
Every request that passed the token bucket got a 500. PriorityQueue.enqueue stores { request: { req, res, next }, priority } and processQueue destructures { req, res, next } from the wrapper, so next is undefined. With that one line fixed:
N=200 200 x 76 503 "Service Unavailable" x 25 429 x 99 stats {"maxQueueLength":1,"serverBusy":0}
N=300 200 x 76 503 x 26 429 x 198 stats {"enqueues":215,"maxQueueLength":1,"serverBusy":0}
maxQueueLength never exceeded 1 across 215 enqueues and Server Busy was never sent, because queueMiddleware calls processQueue() synchronously right after enqueue and processQueue drains the queue in the same tick. The priority sort, the MAX_QUEUE_LENGTH check and the 503 branch were dead code. The 503s in the fixed run are the circuit breaker's response to the simulated 20 percent random failures, and the breaker itself never opened in these runs: it resets its failure count on every success, so it needs five consecutive failures, which a 20 percent rate rarely produces. The token bucket was the only part that worked as described.
Choosing the numbers
- Start from the downstream: the number of concurrent requests it can serve at acceptable latency is
workers. In the lab it was 4 because the service was a sleep; for a real service measure it. - Decide the longest wait a client will tolerate, B, and the typical service time, S. Set
queueCapacityat mostworkers x B / S, so that every queued request is served inside the budget and everything else is shed on arrival withRetry-After. - Set the downstream timeout to something the client will wait for, and remember that it does not cancel the downstream work. Pair it with the concurrency limit so that a slow downstream costs you worker slots, not unbounded threads.
- Rate limits are a separate budget per client or key, applied first. They do not protect the downstream from an unexpected number of clients; the concurrency limit does.
- Return
429for rate limits and503for shedding, both withRetry-After(RFC 6585 section 4 and RFC 9110 sections 15.6.4 and 10.2.3), so clients and tooling can tell "you are over your budget" from "we are full".
What this does not cover
- Distributed rate limiting with a shared store; the bucket is per process.
- Circuit breakers; the earlier example's breaker was run but no breaker was built or measured here. The same site has a separate article on Resilience4j.
- Priority queues, tenant fairness, adaptive concurrency limits (Vegas or gradient style), and CoDel-style queue management.
- Real gateway products (Spring Cloud Gateway, Envoy, NGINX, Kong). The Java gateway is a teaching device on
HttpServer, not something to deploy. - HTTP/2 flow control and TCP receive-window behaviour at the socket level; stated from the RFCs, not measured.
- Sustained load or large worker counts; every run was 50 to 150 concurrent requests on one laptop, and the latency numbers are wall-clock measurements that move by a few milliseconds between runs (the lab README shows two runs).
Reproduce it
cd examples/api-gateway-backpressure
gradle test --no-daemon --console=plain # 13 tests; each prints its [measured] line
gradle -q runDemo # downstream on 8103, gateway on 8102
seq 1 50 | xargs -P 50 -I{} curl -s -o /dev/null -w "%{http_code} %{time_total}n" "localhost:8102/work?ms=200" | sort | uniq -c
curl -s localhost:8102/gateway/stats; curl -s localhost:8103/stats
cd article-node-example && npm install && cat README.md # the earlier Node.js example, as published
Sources
- Reactive Streams 1.0.4: Subscription
- The Reactive Manifesto glossary: Back-Pressure
- JDK 17 API: SubmissionPublisher
- Node.js Learn: Backpressuring in Streams
- RFC 6585 section 4: 429 Too Many Requests
- RFC 9110 section 15.6.4: 503 Service Unavailable
- RFC 9110 section 10.2.3: Retry-After
- RFC 9113 section 5.2: HTTP/2 flow control
- Cloudflare: How we built rate limiting capable of scaling to millions of domains
- Google SRE book: Handling Overload
