Envoy as a Sidecar for a Java Service: A Configuration That Loads, and What Retries, Timeouts, Outlier Ejection and Circuit Breakers Actually Do

Envoy v1.39.1 in Docker 27.4.0 in front of a Java 17 HttpServer; 12 JUnit tests run twice, envoy –mode validate on the earlier configuration on v1.39.1 and v1.23-latest, curl session and access log recorded

Revision note (2026-09-15). The earlier version of this article had one deliverable, a local Envoy proxying to a Java service, and it could not be built from the text. Its envoy.yaml put access_log under static_resources, where the field does not exist, and listed the router filter without typed_config; Envoy rejects it for both reasons, on the v1.23-latest image the article recommended as much as on 1.39.1. Its docker run command pointed the cluster at 127.0.0.1, which inside a container is the container, so even a corrected file returned 503 with the Java service receiving nothing. It told readers to use the admin interface without configuring one, recommended an Envoy version that has been end of life since July 2023, and described the mTLS snippet as securing traffic "without modifying your Java services" although the snippet makes the Java service the TLS server. This version is built from a lab where every behaviour below was observed through Envoy's own stats and access log.

What was run

  • Envoy envoyproxy/envoy:v1.39.1 (/server_info reports 1.39.1/Clean/RELEASE/BoringSSL; released 2026-08-27), in Docker 27.4.0 on Docker Desktop for macOS, container name dd-envoy-sidecar, listener published on host port 10000, admin on 10001.
  • A Java 17 (Amazon Corretto 17.0.14) service on host port 8104 built on com.sun.net.httpserver.HttpServer with no framework, so that every behaviour observed through the proxy is Envoy's. It answers GET /, echoes request headers on /headers, returns any status on /status/{code}, sleeps on /slow?ms=N, and fails a configurable number of times on /flaky.
  • 12 JUnit 5.10.2 tests driven by Gradle 8.8 that start the container, poll the admin /ready endpoint, and assert on responses and on /stats counters. They passed 12 of 12 in two separate runs. The code and the pasted output are in the repository under examples/java-envoy-sidecar.

Spring Boot is not in the lab on purpose; anything Envoy does to a JDK HttpServer it does to a Spring Boot application on the same port.

Why the earlier configuration does not load

envoy --mode validate (CLI reference: "Validate the JSON configuration and then exit") on the earlier file, byte for byte, on both images:

$ docker run --rm -v $PWD/article-verbatim.yaml:/etc/envoy/envoy.yaml:ro envoyproxy/envoy:v1.39.1 --mode validate -c /etc/envoy/envoy.yaml
[critical][main] error initializing configuration '/etc/envoy/envoy.yaml': Protobuf message (type envoy.config.bootstrap.v3.Bootstrap
  reason INVALID_ARGUMENT: invalid JSON in envoy.config.bootstrap.v3.Bootstrap @ static_resources:
  message envoy.config.bootstrap.v3.Bootstrap.StaticResources, near 1:314 (offset 313): no such field: 'access_log') has unknown fields

$ ... envoyproxy/envoy:v1.23-latest --mode validate -c /etc/envoy/envoy.yaml
[critical][main] error initializing configuration '/etc/envoy/envoy.yaml': Protobuf message (type envoy.config.bootstrap.v3.Bootstrap
  reason INVALID_ARGUMENT:(static_resources) access_log: Cannot find field.) has unknown fields

Bootstrap.StaticResources holds listeners, clusters and secrets. Access logging is a property of the HTTP connection manager (access_log (repeated config.accesslog.v3.AccessLog) in the HCM proto). Moving it there, the next error is the same on both images:

[critical][main] error initializing configuration '/etc/envoy/envoy.yaml':
  Didn't find a registered implementation for 'envoy.filters.http.router' with type URL: ''

An HTTP filter entry needs typed_config with the filter's @type; a bare name is not enough. With both fixed, the file validates (configuration '/etc/envoy/envoy.yaml' OK) on 1.23 and 1.39.1. It still does not work with the earlier docker run -p 8080:8080 ... command, because its cluster endpoint is 127.0.0.1:8081:

$ curl -s -i localhost:10002/           # Envoy in Docker, Java service listening on the host
HTTP/1.1 503 Service Unavailable
upstream connect error or disconnect/reset before headers. reset reason: remote connection failure
envoy access log: "GET / HTTP/1.1" 503 UF 0 98 4 - ... "127.0.0.1:10002" "127.0.0.1:8104"

The Java service's hit counter for / stayed at 0. Inside the container, 127.0.0.1 is the container. On Docker Desktop the host is reachable as host.docker.internal; on Linux, docker run --add-host=host.docker.internal:host-gateway provides the same name (or use --network host and skip -p). The lab passes the --add-host flag on every platform.

Three more things the earlier text got wrong, for the record: Envoy 1.23.0 was released on 2022-07-15 and reached end of life on 2023-07-15 (the project's RELEASES.md), and the v1.23-latest tag was last pushed on 2023-07-25; there is no admin interface unless the bootstrap has an admin block (with the corrected earlier file running, /proc/net/tcp in the container showed only port 8080 listening); and envoy -c envoy.yaml on a local binary fails for the same two syntax reasons as Docker.

A configuration that loads

The lab's envoy.yaml, trimmed to the parts that matter for one route (the full file adds routes for the experiments below and is in the repository):

admin:
  address:
    socket_address: { address: 0.0.0.0, port_value: 10001 }
static_resources:
  listeners:
    - name: listener_10000
      address:
        socket_address: { address: 0.0.0.0, port_value: 10000 }
      filter_chains:
        - filters:
            - name: envoy.filters.network.http_connection_manager
              typed_config:
                "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
                stat_prefix: ingress_http
                access_log:
                  - name: envoy.access_loggers.file
                    typed_config:
                      "@type": type.googleapis.com/envoy.extensions.access_loggers.file.v3.FileAccessLog
                      path: /dev/stdout
                      log_format:
                        text_format_source:
                          inline_string: "[%START_TIME%] "%REQ(:METHOD)% %REQ(X-ENVOY-ORIGINAL-PATH?:PATH)%" %RESPONSE_CODE% flags=%RESPONSE_FLAGS% upstream=%UPSTREAM_HOST% attempt=%REQ(X-ENVOY-ATTEMPT-COUNT)% duration_ms=%DURATION% rq_id=%REQ(X-REQUEST-ID)%n"
                route_config:
                  name: local_route
                  virtual_hosts:
                    - name: java
                      domains: ["*"]
                      include_request_attempt_count: true
                      include_attempt_count_in_response: true
                      routes:
                        - match: { prefix: "/" }
                          route: { cluster: java_service }
                http_filters:
                  - name: envoy.filters.http.router
                    typed_config:
                      "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router
  clusters:
    - name: java_service
      connect_timeout: 1s
      type: STRICT_DNS
      dns_lookup_family: V4_ONLY
      lb_policy: ROUND_ROBIN
      load_assignment:
        cluster_name: java_service
        endpoints:
          - lb_endpoints:
              - endpoint:
                  address:
                    socket_address: { address: host.docker.internal, port_value: 8104 }

Start it with:

docker run --rm --name dd-envoy-sidecar --add-host=host.docker.internal:host-gateway 
  -p 10000:10000 -p 10001:10001 -v "$PWD/envoy.yaml:/etc/envoy/envoy.yaml:ro" 
  envoyproxy/envoy:v1.39.1 -c /etc/envoy/envoy.yaml --log-level warn --file-flush-interval-msec 200

--file-flush-interval-msec 200 is there because the file access logger flushes every 10 seconds by default (CLI reference), which makes docker logs look empty during a short experiment. Every field above is listed with its v1.39.1 API reference entry in the lab README.

What the Java service sees

$ curl -s -i localhost:10000/
HTTP/1.1 200 OK
x-served-by: java-service:8104
x-envoy-upstream-service-time: 53
x-envoy-attempt-count: 1
server: envoy

Hello from Java behind Envoy

$ curl -s localhost:10000/headers
{"accept":"*/*","host":"localhost:10000","user-agent":"curl/8.7.1","x-envoy-attempt-count":"1",
 "x-envoy-expected-rq-timeout-ms":"15000","x-forwarded-proto":"http","x-request-id":"a6cb1fec-..."}

Envoy generated x-request-id, added x-forwarded-proto, told the upstream how long the router will wait (x-envoy-expected-rq-timeout-ms: 15000, the default route timeout of 15 s), and added the attempt count because the virtual host asks for it. It did not add x-forwarded-for: the HCM headers reference says Envoy appends to XFF only when use_remote_address is true (the default is false, which the docs describe as right for an internal service node in a mesh) or skip_xff_append is false. A Java service that logs the client IP from XFF behind a sidecar with the defaults will log nothing.

Retries

A route with retry_policy: { retry_on: "5xx", num_retries: 3, per_try_timeout: 1s } in front of /flaky, which the test tells to fail twice:

$ curl -s -X POST localhost:8104/control/flaky/2; curl -s -i localhost:10000/retry/flaky
HTTP/1.1 200 OK
x-envoy-attempt-count: 3

flaky ok (x-envoy-attempt-count=3)
access log: "GET /retry/flaky" 200 flags=- attempt=3 duration_ms=71

Two 503s from the Java service were retried and the client saw a 200 on attempt 3; the service's own counter recorded three calls, cluster.java_service.upstream_rq_retry rose by 2 and upstream_rq_retry_success by 1. With ten failures queued instead:

GET /retry/flaky -> 503 flaky failure (x-envoy-attempt-count=4, failures left after this one=6)
access log: "GET /retry/flaky" 503 flags=URX attempt=4 duration_ms=125
stats: upstream_rq_retry +3, upstream_rq_retry_limit_exceeded +1

num_retries: 3 means four attempts; the fourth attempt's response is what the client gets, unchanged. Retries are transparent to the Java service except that it does the work four times, which is why a retried route should be idempotent and why the same site's article on idempotent REST APIs exists.

Route timeout

timeout: 1s on /slow, against a 3 s sleep:

$ curl -s -i -w "time_total=%{time_total}sn" "localhost:10000/slow?ms=3000"
HTTP/1.1 504 Gateway Timeout
upstream request timeout
time_total=1.004587s
access log: "GET /slow?ms=3000" 504 flags=UT duration_ms=1000

cluster.java_service.upstream_rq_timeout rose by 1. The Java thread kept sleeping; Envoy's timeout ends the client's wait, not the upstream's work, and the router reference notes that a request that exceeds the route timeout is not retried by the 5xx policy.

Outlier detection and the two defaults that hide it

A cluster with outlier_detection: { consecutive_5xx: 3, interval: 1s, base_ejection_time: 5s, max_ejection_percent: 100 } and common_lb_config: { healthy_panic_threshold: { value: 0 } }:

$ for i in 1 2 3; do curl -s -o /dev/null -w "%{http_code} " localhost:10000/outlier/status/503; done
503 503 503
$ curl -s -i localhost:10000/outlier/
HTTP/1.1 503 Service Unavailable
no healthy upstream
access log: "GET /outlier/" 503 flags=UH upstream=- duration_ms=0
stats: outlier_detection.ejections_active: 1, ejections_enforced_consecutive_5xx: 1

After the third consecutive 5xx the only host was ejected, and the next request was answered by Envoy itself (no healthy upstream, flag UH, no upstream host) without touching the Java service, whose counter for / stayed at 0. The test polled ejections_active and saw it return to 0 after about 5.7 s, consistent with base_ejection_time: 5s checked every interval: 1s; the next request was a 200 again.

Two settings had to be changed from their defaults to see any of that with a single host:

  • max_ejection_percent defaults to 10 percent (outlier_detection.proto). With one host, 10 percent of the cluster is nobody, so nothing is ever ejected.
  • healthy_panic_threshold defaults to 50 percent (cluster.proto; the panic threshold page). With 0 of 1 hosts healthy, the cluster is below the threshold, and in panic mode Envoy routes to all hosts, ejected or not. The lab keeps a second cluster identical except for the panic threshold:
3 x GET /outlier-panic/status/503 -> ejections_active: 1
GET /outlier-panic/ -> 200 Hello from Java behind Envoy   (lb_healthy_panic +1, host still ejected)

The ejected host served the request anyway and the Java service's counter went to 1. Both defaults are sensible for a cluster of many hosts and invisible in a two-container demo, which is where most people first try outlier detection and conclude it does nothing.

Circuit breakers are per worker thread, eventually

A cluster with circuit_breakers.thresholds: [{ priority: DEFAULT, max_connections: 1, max_pending_requests: 1, max_requests: 1024, max_retries: 3, track_remaining: true }], and 20 concurrent GET /tight/slow?ms=1000. First with Envoy started with --concurrency 1:

{200 slow response after 1000 ms=2, 503 upstream connect error or disconnect/reset before headers. reset reason: overflow=18}
upstream_rq_pending_overflow=18

Exactly one in flight plus one pending: 2 served, 18 rejected immediately with flag UO, and upstream_rq_pending_overflow counted every rejection. The same run against the default worker count on this machine ("concurrency": 11 in /server_info) served 2 in the JUnit run and 3 in the curl session (17 overflow). The circuit-breaking page explains why the number can move: worker threads share the limits, but "since the implementation is eventually consistent, races between threads may allow limits to be potentially exceeded". The test asserts a range for that case and an exact count only for the single-worker one. If you need a hard bound of one, put it in the upstream, not in the sidecar.

When the Java service is down

$ kill <java service>; curl -s -i localhost:10000/
HTTP/1.1 503 Service Unavailable
upstream connect error or disconnect/reset before headers. reset reason: remote connection failure
access log: "GET /" 503 flags=UF upstream=<host-gateway>:8104 duration_ms=2

Restarting the service on 8104 made the next request a 200 without touching Envoy: STRICT_DNS re-resolves host.docker.internal continuously and the connection pool reconnects. The response flags in the access log (UF connection failure, UT timeout, UH no healthy host, UO overflow, URX retry limit) are the fastest way to tell which of these mechanisms produced a 503 or 504; the JSON field %RESPONSE_FLAGS% is documented on the access log formatter page.

About the mTLS step

The earlier article's TLS snippet adds an UpstreamTlsContext to the cluster: Envoy becomes a TLS client to java-service:8443 and presents service.crt. In that layout the Java service is the TLS server. It has to terminate TLS on 8443 and, for this to be mutual, request and verify the client certificate itself, which is the opposite of "without modifying your Java services". Sidecar mTLS keeps the Java service on plain HTTP behind its own Envoy, whose listener carries a DownstreamTlsContext with require_client_certificate: true (tls.proto: "Envoy will reject connections without a valid client certificate"); the calling side's Envoy carries the UpstreamTlsContext. That was checked against the v1.39.1 reference only; no certificates were generated and no TLS run is part of this lab.

What this does not cover

  • mTLS, any control plane (Istio, Consul, xDS), gRPC, active health checks, or Prometheus scraping of /stats/prometheus.
  • Linux Docker networking; host.docker.internal was tested on Docker Desktop for macOS. The --add-host=host.docker.internal:host-gateway flag is what the Docker documentation describes for Linux and is in the lab script, but it was not run on Linux.
  • Spring Boot itself. The lab's service is a JDK HttpServer so that every observed behaviour is Envoy's.
  • The exact over-admission of the circuit breaker with many workers (2 and 3 of 20 were observed; the test asserts a range), and the default 10 s access-log flush delay (taken from the CLI reference, not measured).

Reproduce it

cd examples/java-envoy-sidecar
gradle test --no-daemon --console=plain     # 12 tests; pulls envoyproxy/envoy:v1.39.1 on first run
gradle -q runService                        # Java service on 8104
sh envoy/run-envoy.sh                       # Envoy on 10000, admin on 10001, container dd-envoy-sidecar
curl -s -i localhost:10000/
curl -s "localhost:10001/stats?filter=java_service"
docker logs dd-envoy-sidecar
docker rm -f dd-envoy-sidecar

Sources