Revision note (2026-09-15). The earlier version of this article told readers to hand-write an Istio DestinationRule with version: v1 / version: v2 subsets and a VirtualService for Flagger to "update gradually". Flagger does not work that way: it generates its own services, DestinationRules and VirtualService and overwrites a pre-existing VirtualService of the same name, and nothing in the old text ever created a pod with a version: v2 label. The Canary referenced a load tester that was never installed, the text gave no way to produce a failure, and the versions (Istio 1.15, Kubernetes 1.16) were years out of support. This version is built from a recorded run: Flagger v1.45.0 with Istio 1.31.0 on a kind cluster, one healthy candidate promoted, one failing candidate rolled back, and one configuration mistake found along the way. The lab, scripts, and the unedited log are in the repository under examples/k8s-progressive-delivery/canary-flagger-istio/runtime.
What Flagger actually creates
Tested environment: kind v0.33.0 with node image kindest/node:v1.34.11 (Kubernetes 1.34.11), Istio 1.31.0 installed with istioctl install --set profile=minimal (istiod only, no ingress gateway), Prometheus from Istio's samples/addons/prometheus.yaml (15 s scrape interval), Flagger v1.45.0 from kubectl apply -k https://github.com/fluxcd/flagger//kustomize/istio?ref=v1.45.0, and podinfo plus the Flagger load tester from the kustomize/podinfo and kustomize/tester overlays of the same tag. Docker had 8 GB; nothing was evicted.
You give Flagger one Deployment (podinfo, two replicas) and a Canary resource. On first reconcile it builds everything else:
$ kubectl -n test get deploy,svc,virtualservice,destinationrule,canary
deployment.apps/podinfo 0/0 0 0 3m16s
deployment.apps/podinfo-primary 2/2 2 2 33s
service/podinfo ClusterIP 10.96.65.68 9898/TCP 3s
service/podinfo-canary ClusterIP 10.96.43.61 9898/TCP 33s
service/podinfo-primary ClusterIP 10.96.96.126 9898/TCP 33s
virtualservice.networking.istio.io/podinfo ["mesh"] ["podinfo"] 3s
destinationrule.networking.istio.io/podinfo-canary podinfo-canary 3s
destinationrule.networking.istio.io/podinfo-primary podinfo-primary 3s
canary.flagger.app/podinfo Initialized 0
$ kubectl -n test get svc podinfo -o jsonpath='{.spec.selector}'
{"app":"podinfo-primary"}
Read that carefully, because it is the part the old article got backwards. Your Deployment podinfo is scaled to zero and becomes the canary; a copy named podinfo-primary serves production; the Service podinfo (which Flagger re-points if you created it yourself) selects the primary; and the VirtualService podinfo routes between the podinfo-primary and podinfo-canary services with weights 100 and 0. There are no version subsets and nothing for you to write in Istio.
The Canary used in the run
apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
name: podinfo
namespace: test
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: podinfo
progressDeadlineSeconds: 60
service:
port: 9898
targetPort: 9898
analysis:
interval: 30s
threshold: 5
maxWeight: 50
stepWeight: 10
metrics:
- name: request-success-rate
thresholdRange:
min: 99
interval: 1m
- name: request-duration
thresholdRange:
max: 500
interval: 1m
webhooks:
- name: acceptance-test
type: pre-rollout
url: http://flagger-loadtester.test/
timeout: 30s
metadata:
type: bash
cmd: "curl -sd 'test' http://podinfo-canary:9898/token | grep token"
- name: load-test-canary
type: rollout
url: http://flagger-loadtester.test/
timeout: 5s
metadata:
cmd: "hey -z 30s -q 10 -c 2 http://podinfo-canary.test:9898/"
- name: load-test-weighted
type: rollout
url: http://flagger-loadtester.test/
timeout: 5s
metadata:
cmd: "hey -z 30s -q 10 -c 2 http://podinfo.test:9898/"
Points that matter:
- The load tester is a separate Deployment (
flagger/kustomize/tester), and the rollout webhooks call it every interval. Without it a canary in a test cluster gets no traffic, and Flagger halts with "no values found" until it hitsthreshold. The old article's Canary referenced aflagger-loadtesterit never installed. request-durationhasinterval: 1mhere. The run was first done with30s, and that is the mistake described below.- Two load generators: one straight at
podinfo-canary(like Flagger's tutorial) and one atpodinfo, which goes through the VirtualService so the weighted split shows up in Prometheus. threshold: 5means five failed checks in total, not five in a row. Every "Halt" event counts.
A healthy candidate, promoted
Trigger: kubectl -n test set image deployment/podinfo podinfod=ghcr.io/stefanprodan/podinfo:6.9.0. A watcher sampled the Canary phase, the VirtualService route weights and the ready pods every five seconds. Condensed to the transitions (full samples in the lab):
06:28:03 phase=Progressing weight=0 vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=/2
06:28:25 phase=Progressing weight=0 vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=2/2
06:28:35 phase=Progressing weight=10 vs[podinfo-primary=90 podinfo-canary=10 ] primary=2/2 canary=2/2
06:29:34 phase=Progressing weight=20 vs[podinfo-primary=80 podinfo-canary=20 ] primary=2/2 canary=2/2
06:30:33 phase=Progressing weight=30 vs[podinfo-primary=70 podinfo-canary=30 ] primary=2/2 canary=2/2
06:31:08 phase=Progressing weight=40 vs[podinfo-primary=60 podinfo-canary=40 ] primary=2/2 canary=2/2
06:32:06 phase=Progressing weight=50 vs[podinfo-primary=50 podinfo-canary=50 ] primary=2/2 canary=2/2
06:33:05 phase=Promoting weight=50 vs[podinfo-primary=50 podinfo-canary=50 ] primary=1/2 canary=2/2
06:33:37 phase=Finalising weight=0 vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=2/2
06:34:04 phase=Succeeded weight=0 vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=/0
388 seconds from set image to Succeeded. The sequence is: scale the canary Deployment up and wait for it to be ready, run the pre-rollout webhook, then advance the VirtualService weight by stepWeight per interval while the metrics pass, and at maxWeight copy the canary's pod template into podinfo-primary, wait for that Deployment's own rolling update (the 1/2 and 3/2 samples), set the route back to 100/0 and scale the canary to zero. At the end podinfo-primary runs 6.9.0 and podinfo is 0/0 with the same image.
Prometheus confirms the split. Only the weighted generator (20 requests per second) reaches the primary, so the primary's request rate should be 20 x (1 – weight). Sampled at 30 s: 15.8, 16.2, 14.6, 11.7, 9.7 and 7.1 requests per second as the weight went 10, 20, 30, 40, 50 (expected 18, 16, 14, 12, 10; the samples are noisy because each hey run lasts 30 s and the sampling points fall on its edges). Over the run: 4831 requests to the canary and 2341 to the primary, all HTTP 200.
The mistake: the healthy rollout finished with four failed checks
Status:
Failed Checks: 4
Phase: Succeeded
Events:
Warning Halt advancement no values found for istio metric request-success-rate probably podinfo.test is not receiving traffic
Warning Halt advancement no values found for istio metric request-duration probably podinfo.test is not receiving traffic (x3)
Four halts against threshold: 5. One more and a perfectly healthy candidate would have been rolled back, and the events would have said "failed checks threshold reached", which is what an on-call engineer would read as "the new version is bad".
The first halt is the first check after weight 10: the canary had been receiving traffic for less than one Prometheus scrape interval, so the 1-minute success-rate window was empty. That one happened in all three recorded rollouts. The other three are the request-duration metric with interval: 30s. Flagger's built-in Istio query for that metric takes a rate over the metric's interval, the Istio addon Prometheus scrapes every 15 seconds, and a 30-second window with scrape jitter regularly contains fewer than the two samples a rate needs. Result: "no values found", which Flagger counts as a failed check.
The fix is to make the metric window at least a few scrape intervals long. The same rollout with request-duration at interval: 1m (the Canary shown above) was run as a second healthy candidate (6.10.0):
06:50:06 phase=Progressing weight=10 failed=0
06:50:33 phase=Progressing weight=10 failed=1 <- first-check "no values" for request-success-rate
06:51:05 phase=Progressing weight=20 failed=1
06:51:38 phase=Progressing weight=30 failed=1
06:52:05 phase=Progressing weight=40 failed=1
06:52:37 phase=Progressing weight=50 failed=1
06:53:04 phase=Promoting weight=50 failed=1
06:54:03 phase=Succeeded weight=0 failed=1
One failed check, none from request-duration, and 304 seconds instead of 388 because no step had to be retried. One run per configuration is not a proof, but it matches the explanation, and the general rule stands whatever the provider: know your scrape interval, and give every metric query a window that is a multiple of it. The threshold you choose should leave room for the first-check halt.
A failing candidate, rolled back
Trigger: set image to podinfo:6.9.2, and at the same time, from inside the load tester pod:
hey -z 8m -q 10 -c 2 http://podinfo-canary.test:9898/status/500
podinfo returns HTTP 500 on that path, so the canary's sidecar reports a low success rate while the primary sees only 200s. This is Flagger's own way of simulating a bad release, and it is what the old article's "simulate a failure and observe Flagger reverting" was missing.
06:38:33 phase=Progressing weight=0 failed=0 vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=/2
06:39:38 phase=Progressing weight=10 failed=0 vs[podinfo-primary=90 podinfo-canary=10 ] primary=2/2 canary=2/2
06:40:05 phase=Progressing weight=10 failed=1 vs[podinfo-primary=90 podinfo-canary=10 ]
06:40:37 phase=Progressing weight=10 failed=2
06:41:04 phase=Progressing weight=10 failed=3
06:41:36 phase=Progressing weight=10 failed=4
06:42:04 phase=Progressing weight=10 failed=5
06:42:36 phase=Failed weight=0 failed=0 vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=/0
From the Flagger log, one line per 30-second check:
06:40:03 Halt podinfo.test advancement success rate 0.00% < 99%
06:40:33 Halt podinfo.test advancement success rate 17.21% < 99%
06:41:03 Halt podinfo.test advancement success rate 22.84% < 99%
06:41:33 Halt podinfo.test advancement success rate 20.27% < 99%
06:42:03 Halt podinfo.test advancement success rate 23.02% < 99%
06:42:33 Rolling back podinfo.test failed checks threshold reached 5
06:42:33 Canary failed! Scaling down podinfo.test
250 seconds from set image to Failed. The weight never went past 10, so at most a tenth of the weighted traffic touched the bad version for about three minutes; the VirtualService is back at 100/0, the canary Deployment is at zero replicas, and podinfo-primary still runs 6.9.0. Prometheus for the window: 12158 HTTP 500 and 5598 HTTP 200 on the canary workload, 3712 requests on the primary, all 200. The primary never served an error.
Note what "rollback" means here: Flagger does not roll the primary back, because the primary was never changed. It stops sending traffic to the canary and scales it down. The failed image stays in the podinfo Deployment spec; the next set image to a different tag starts a new analysis.
How to watch a rollout
kubectl -n test get canary podinfoshows the phase and weight;kubectl -n test describe canary podinfohas the events, but Kubernetes merges similar events, so the per-check values are in Flagger's log (kubectl -n istio-system logs deploy/flagger).- The live weights are in the VirtualService:
kubectl -n test get virtualservice podinfo -o jsonpath='{.spec.http[0].route[*].weight}'. kubectl rollout status deployment/podinfo, which the old article recommended, reports on the canary Deployment's own pod rollout, which finishes within seconds of the scale-up; it says nothing about the analysis.- The weighted split itself can be checked in Prometheus with
sum by (destination_workload) (rate(istio_requests_total{reporter="destination",destination_workload_namespace="test"}[30s])).
What went wrong in the lab, and what it teaches
Two stages of the run script were botched and are left in the log. The first version of the failing-candidate stage did not wait for the phase to leave Succeeded before waiting for a terminal phase, so it returned immediately while Flagger was still detecting the new revision; the 6.9.1 rollout it started ran unwatched and was rolled back on the same 500s. Then, killing the local kubectl exec ... hey did not stop hey inside the load tester pod, so the 500 generator was still running when the next healthy candidate started, and that one was rolled back too, with success rates between 0 and 62 percent. Both are worth knowing if you script this: the Canary phase is stateful across rollouts, and load generators live in the cluster, not in your terminal.
What was not verified
- An ingress gateway in front of the mesh (
service.gateways/hostsin the Canary). The run kept all traffic inside the mesh. - Istio
demoordefaultprofiles; onlyminimalwas installed. - The
request-durationmetric failing on its own; podinfo answers in a few milliseconds, so only the "no values" case was seen. - An HPA on the target (
autoscalerRef); the HPA from the podinfo overlay was removed because kind has no metrics-server. - Session affinity, header-based A/B routing, mirroring, and Flagger's non-Istio providers.
- Multi-node clusters and promotion timing with large images.
Sources
- Flagger v1.45.0 Istio tutorial: https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/tutorials/istio-progressive-delivery.md
- Flagger v1.45.0, how it works (canary service, primary copy): https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/usage/how-it-works.md
- Flagger v1.45.0 webhooks and metrics: https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/usage/webhooks.md and https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/usage/metrics.md
- Flagger v1.45.0 install with Kustomize: https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/install/flagger-install-on-kubernetes.md
- Istio 1.31.0 release: https://github.com/istio/istio/releases/tag/1.31.0
- kind v0.33.0 release (node image digests): https://github.com/kubernetes-sigs/kind/releases/tag/v0.33.0
- podinfo: https://github.com/stefanprodan/podinfo
