Revision note (2026-09-15). The earlier version of this article used placeholder images (myapp:v1.0.0) that cannot be pulled, verified everything with kubectl get endpoints (the Endpoints API is deprecated since Kubernetes 1.33), claimed the selector patch "atomically switches" traffic with no measurement, and simulated a failure by deleting a green pod, which exercises the Deployment controller rather than the readiness probe it then talked about. This version keeps the same idea (two Deployments, one Service, flip the selector) and replaces every claim with something observed on a cluster: a request loop across the switch, a green pod failing readiness, a full-green failure with a revert, and what kubectl patch does to a selector key it does not mention.
The setup that was actually run
Kubernetes v1.34.11 on kind v0.33.0 (one node, kube-proxy in iptables mode as kind ships it), kubectl client 1.30.5, application image ghcr.io/stefanprodan/podinfo at 6.9.0 for blue and 6.9.1 for green. podinfo answers GET / with JSON that includes version and hostname, so a request loop can tell which pod answered. It also has POST /readyz/disable, which makes its readiness probe fail without killing the process; that is how "a failing pod" is produced below. The lab, its scripts and the full log are in the repository under examples/k8s-progressive-delivery/blue-green/runtime.
Blue is a plain Deployment with three replicas and two probes:
apiVersion: apps/v1
kind: Deployment
metadata:
name: myapp-blue
namespace: bluegreen
labels:
app: myapp
version: blue
spec:
replicas: 3
selector:
matchLabels:
app: myapp
version: blue
template:
metadata:
labels:
app: myapp
version: blue
tier: web
spec:
containers:
- name: myapp
image: ghcr.io/stefanprodan/podinfo:6.9.0
args: ["./podinfo", "--port=9898", "--level=info"]
env:
- name: PODINFO_UI_COLOR
value: "blue"
ports:
- name: http
containerPort: 9898
readinessProbe:
httpGet:
path: /readyz
port: http
periodSeconds: 2
failureThreshold: 2
livenessProbe:
httpGet:
path: /healthz
port: http
initialDelaySeconds: 5
periodSeconds: 10
Green is the same file with myapp-green, version: green, image tag 6.9.1, and no tier label. The tier: web label on blue is there on purpose; it is the prop for the patch section. The Service starts on blue:
apiVersion: v1
kind: Service
metadata:
name: myapp-service
namespace: bluegreen
spec:
type: ClusterIP
selector:
app: myapp
version: blue
ports:
- name: http
port: 80
targetPort: http
Read endpoints from EndpointSlice, not Endpoints
The Endpoints API is deprecated as of Kubernetes 1.33 and every read of it returns a warning. The replacement is the EndpointSlice API, and it carries the fact you care about in a blue-green switch, the per-endpoint ready condition:
kubectl -n bluegreen get endpointslice -l kubernetes.io/service-name=myapp-service
-o jsonpath='{range .items[*].endpoints[*]}{.targetRef.name}{"t"}{.addresses[0]}{"tready="}{.conditions.ready}{"tserving="}{.conditions.serving}{"n"}{end}'
Right after applying the manifests:
myapp-blue-5c7fcb656-7jn5f 10.244.0.25 ready=true serving=true
myapp-blue-5c7fcb656-bvprw 10.244.0.24 ready=true serving=true
myapp-blue-5c7fcb656-qkpnc 10.244.0.22 ready=true serving=true
Green pods are running and ready but are not in the slice, because the selector does not match them.
Where the request loop has to run
kubectl port-forward svc/myapp-service cannot observe a switch. The port-forward picks one pod when it starts and keeps sending to that pod, so it would keep showing blue after the selector moved to green. The loop in this lab is a small Python script executed inside a client pod in the same namespace; it opens a new connection per request to http://myapp-service.bluegreen/, every 50 ms, and records the response time, status, version and hostname. Twenty warm-up requests before the switch all returned 6.9.0 from the three blue pods.
The switch, measured
With the loop running for 30 s, the switch was done 8 s in:
kubectl -n bluegreen patch svc myapp-service -p '{"spec":{"selector":{"app":"myapp","version":"green"}}}'
The recorded result:
patch sent at 1789453752.214, returned at 1789453752.304
change,1789453744.198,200,6.9.0,myapp-blue-5c7fcb656-qkpnc
change,1789453752.303,200,6.9.1,myapp-green-9dc55f485-dp2x9
summary,total=552,ok=552,failed=0,by_version={"6.9.0": 150, "6.9.1": 402}
552 requests, none failed, 150 answered by blue and 402 by green. The first green answer arrived within the measurement's 0.1 s resolution of the patch call (the two clocks were compared with one kubectl exec round trip, which is about that long), and the EndpointSlice read immediately after the patch already listed only the three green pods. So on this cluster the switch was as fast as it can be measured with this method, and no request was lost. What this does not show: a keep-alive connection that was already open to a blue pod stays on blue until it is closed, because a selector change does not touch established connections. The loop here opens a fresh connection per request, so that case was not exercised.
Readiness gating, with a pod that is actually unready
The earlier text "simulated failure" by deleting a green pod. That tests the Deployment controller (it makes a new pod), not the readiness probe. To test the probe, make it fail:
POST http://<green pod ip>:9898/readyz/disable
With periodSeconds: 2 and failureThreshold: 2 the kubelet needs two failed probes; the endpoint's ready condition flipped to false 4.3 s after the call, and the pod showed 0/1 while still Running:
myapp-green-9dc55f485-9zx5m 10.244.0.27 ready=false serving=false
myapp-green-9dc55f485-dp2x9 10.244.0.26 ready=true serving=true
myapp-green-9dc55f485-tx9l9 10.244.0.23 ready=true serving=true
Note the pod stays in the slice with ready=false; it is not removed. Ten seconds of requests after that:
summary,total=183,ok=183,failed=0,by_version={"6.9.1": 183},hosts=myapp-green-9dc55f485-dp2x9,myapp-green-9dc55f485-tx9l9
No request reached the unready pod and none failed. That is the guarantee readiness gives you in a blue-green setup: a green pod that is not ready gets nothing, so bringing green up early costs nothing and lets you inspect it before the switch.
What kubectl patch does to your selector
kubectl patch with -p and no --type sends a strategic merge patch. For a map like spec.selector this merges keys: keys named in the patch are set, keys not named are kept. The earlier article called this an atomic replacement. It is atomic as an API write, but it is not a replacement, and the difference bites when the selector has a key the patch does not name.
The lab has a second Service, myapp-service-extra-key, whose selector is {app: myapp, version: blue, tier: web}. Blue pods have tier: web, green pods do not. Applying the article's patch to it:
$ kubectl -n bluegreen patch svc myapp-service-extra-key -p '{"spec":{"selector":{"app":"myapp","version":"green"}}}'
service/myapp-service-extra-key patched
$ kubectl -n bluegreen get svc myapp-service-extra-key -o jsonpath='{.spec.selector}'
{"app":"myapp","tier":"web","version":"green"}
tier: web survived, no pod matches all three keys, the slice went to zero endpoints, and 8 of 8 requests in the next two seconds failed with connection errors. The Service was down. Two ways out, both observed:
# remove the stale key: a merge patch removes a key when its value is null
kubectl -n bluegreen patch svc myapp-service-extra-key -p '{"spec":{"selector":{"tier":null}}}'
# -> {"app":"myapp","version":"green"}, three green endpoints back
# or replace the whole map with a JSON patch
kubectl -n bluegreen patch svc myapp-service-extra-key --type=json
-p '[{"op":"replace","path":"/spec/selector","value":{"app":"myapp","version":"blue","tier":"web"}}]'
The practical rule: keep the selector to exactly the keys you switch on, or use the JSON patch replace so the command states the entire selector.
Failure after the switch, then revert
The sequence that the earlier article described but did not run: traffic is on green, green goes bad, the operator switches back. In the lab, "bad" is all three green pods failing readiness at once (/readyz/disable on each), with the 50 ms request loop running for 40 s.
all green endpoints ready=false at 1789453807.475 (4.2 s after readiness was disabled)
myapp-green-9dc55f485-9zx5m 10.244.0.27 ready=false serving=false
myapp-green-9dc55f485-dp2x9 10.244.0.26 ready=false serving=false
myapp-green-9dc55f485-tx9l9 10.244.0.23 ready=false serving=false
From that moment the Service had no ready endpoint and every request failed. The script waited five seconds, standing in for a person deciding, and then reverted:
kubectl -n bluegreen patch svc myapp-service -p '{"spec":{"selector":{"app":"myapp","version":"blue"}}}'
rollback patch sent at 1789453812.623, returned at 1789453812.703
change,1789453807.279,ERR:URLError,-,-
change,1789453812.999,200,6.9.0,myapp-blue-5c7fcb656-7jn5f
summary,total=644,ok=634,failed=10,by_version={"6.9.0": 467, "6.9.1": 167}
failed_window,1789453807.279,1789453811.910
first response served by blue after the rollback: 0.26 s after the patch was sent
Three things worth reading off this. First, the outage was not caused by the switch; it ran from the moment the last green endpoint became unready until the revert, and its length was the 4.2 s the probes took to notice plus the 5 s decision delay. Second, the revert itself took 0.26 s from kubectl patch to the first blue response. Third, the request loop lost 10 of 644 requests, and all 10 are inside that window; the five failed samples printed were immediate connection errors, and the later ones took about a second each before failing, which is consistent with a dropped SYN being retried but was not investigated further. One request that started 0.8 s before the revert was answered by blue after it.
The revert works only because blue was still there with three ready pods. The kubectl delete deployment myapp-blue step that the earlier article listed under "reclaim resources" is exactly the step that removes your rollback; do it after you no longer need one.
Rolling update is not "partial downtime"
The earlier trade-off table said a rolling update risks "possible partial downtime". With readiness probes and maxUnavailable: 0 a rolling update never takes a ready pod away before a replacement is ready; the real difference from blue-green is that both versions serve at once for the duration of the rollout and there is no single moment of switch. Blue-green buys the single moment, costs a second full set of pods, and (as measured above) delivers the switch in well under a second on a plain ClusterIP Service.
What was not measured
- An ingress controller or cloud load balancer in front of the Service. Those keep their own endpoint views and their own connection pools; the timings above end at kube-proxy.
- Established keep-alive connections across the switch.
- kube-proxy in IPVS or nftables mode, and multi-node clusters, where endpoint propagation to every node adds time.
- Anything about session state; podinfo is stateless.
Sources
- kind v0.33.0 release: https://github.com/kubernetes-sigs/kind/releases/tag/v0.33.0
- Endpoints deprecation (Kubernetes 1.33): https://kubernetes.io/blog/2025/04/24/endpoints-deprecation/
- EndpointSlice API: https://kubernetes.io/docs/concepts/services-networking/endpoint-slices/
- kubectl patch (strategic merge, JSON merge, JSON patch): https://kubernetes.io/docs/tasks/manage-kubernetes-objects/update-api-object-kubectl-patch/
- Probes: https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/
- podinfo: https://github.com/stefanprodan/podinfo
