Canary Releases With Automated Rollback on Kubernetes: Flagger and Istio, Run and Recorded

Run on kind v0.33.0 (kindest/node v1.34.11) with Istio 1.31.0 (minimal profile), Flagger v1.45.0, podinfo 6.0.0 to 6.10.0; every number comes from the recorded run in examples/k8s-progressive-delivery/canary-flagger-istio/runtime

Revision note (2026-09-15). The earlier version of this article told readers to hand-write an Istio DestinationRule with version: v1 / version: v2 subsets and a VirtualService for Flagger to "update gradually". Flagger does not work that way: it generates its own services, DestinationRules and VirtualService and overwrites a pre-existing VirtualService of the same name, and nothing in the old text ever created a pod with a version: v2 label. The Canary referenced a load tester that was never installed, the text gave no way to produce a failure, and the versions (Istio 1.15, Kubernetes 1.16) were years out of support. This version is built from a recorded run: Flagger v1.45.0 with Istio 1.31.0 on a kind cluster, one healthy candidate promoted, one failing candidate rolled back, and one configuration mistake found along the way. The lab, scripts, and the unedited log are in the repository under examples/k8s-progressive-delivery/canary-flagger-istio/runtime.

What Flagger actually creates

Tested environment: kind v0.33.0 with node image kindest/node:v1.34.11 (Kubernetes 1.34.11), Istio 1.31.0 installed with istioctl install --set profile=minimal (istiod only, no ingress gateway), Prometheus from Istio's samples/addons/prometheus.yaml (15 s scrape interval), Flagger v1.45.0 from kubectl apply -k https://github.com/fluxcd/flagger//kustomize/istio?ref=v1.45.0, and podinfo plus the Flagger load tester from the kustomize/podinfo and kustomize/tester overlays of the same tag. Docker had 8 GB; nothing was evicted.

You give Flagger one Deployment (podinfo, two replicas) and a Canary resource. On first reconcile it builds everything else:

$ kubectl -n test get deploy,svc,virtualservice,destinationrule,canary
deployment.apps/podinfo              0/0     0            0           3m16s
deployment.apps/podinfo-primary      2/2     2            2           33s
service/podinfo                      ClusterIP   10.96.65.68    9898/TCP   3s
service/podinfo-canary               ClusterIP   10.96.43.61    9898/TCP   33s
service/podinfo-primary              ClusterIP   10.96.96.126   9898/TCP   33s
virtualservice.networking.istio.io/podinfo   ["mesh"]   ["podinfo"]   3s
destinationrule.networking.istio.io/podinfo-canary    podinfo-canary    3s
destinationrule.networking.istio.io/podinfo-primary   podinfo-primary   3s
canary.flagger.app/podinfo   Initialized   0
$ kubectl -n test get svc podinfo -o jsonpath='{.spec.selector}'
{"app":"podinfo-primary"}

Read that carefully, because it is the part the old article got backwards. Your Deployment podinfo is scaled to zero and becomes the canary; a copy named podinfo-primary serves production; the Service podinfo (which Flagger re-points if you created it yourself) selects the primary; and the VirtualService podinfo routes between the podinfo-primary and podinfo-canary services with weights 100 and 0. There are no version subsets and nothing for you to write in Istio.

The Canary used in the run

apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: podinfo
  namespace: test
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: podinfo
  progressDeadlineSeconds: 60
  service:
    port: 9898
    targetPort: 9898
  analysis:
    interval: 30s
    threshold: 5
    maxWeight: 50
    stepWeight: 10
    metrics:
    - name: request-success-rate
      thresholdRange:
        min: 99
      interval: 1m
    - name: request-duration
      thresholdRange:
        max: 500
      interval: 1m
    webhooks:
    - name: acceptance-test
      type: pre-rollout
      url: http://flagger-loadtester.test/
      timeout: 30s
      metadata:
        type: bash
        cmd: "curl -sd 'test' http://podinfo-canary:9898/token | grep token"
    - name: load-test-canary
      type: rollout
      url: http://flagger-loadtester.test/
      timeout: 5s
      metadata:
        cmd: "hey -z 30s -q 10 -c 2 http://podinfo-canary.test:9898/"
    - name: load-test-weighted
      type: rollout
      url: http://flagger-loadtester.test/
      timeout: 5s
      metadata:
        cmd: "hey -z 30s -q 10 -c 2 http://podinfo.test:9898/"

Points that matter:

  • The load tester is a separate Deployment (flagger/kustomize/tester), and the rollout webhooks call it every interval. Without it a canary in a test cluster gets no traffic, and Flagger halts with "no values found" until it hits threshold. The old article's Canary referenced a flagger-loadtester it never installed.
  • request-duration has interval: 1m here. The run was first done with 30s, and that is the mistake described below.
  • Two load generators: one straight at podinfo-canary (like Flagger's tutorial) and one at podinfo, which goes through the VirtualService so the weighted split shows up in Prometheus.
  • threshold: 5 means five failed checks in total, not five in a row. Every "Halt" event counts.

A healthy candidate, promoted

Trigger: kubectl -n test set image deployment/podinfo podinfod=ghcr.io/stefanprodan/podinfo:6.9.0. A watcher sampled the Canary phase, the VirtualService route weights and the ready pods every five seconds. Condensed to the transitions (full samples in the lab):

06:28:03 phase=Progressing  weight=0   vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=/2
06:28:25 phase=Progressing  weight=0   vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=2/2
06:28:35 phase=Progressing  weight=10  vs[podinfo-primary=90 podinfo-canary=10 ] primary=2/2 canary=2/2
06:29:34 phase=Progressing  weight=20  vs[podinfo-primary=80 podinfo-canary=20 ] primary=2/2 canary=2/2
06:30:33 phase=Progressing  weight=30  vs[podinfo-primary=70 podinfo-canary=30 ] primary=2/2 canary=2/2
06:31:08 phase=Progressing  weight=40  vs[podinfo-primary=60 podinfo-canary=40 ] primary=2/2 canary=2/2
06:32:06 phase=Progressing  weight=50  vs[podinfo-primary=50 podinfo-canary=50 ] primary=2/2 canary=2/2
06:33:05 phase=Promoting    weight=50  vs[podinfo-primary=50 podinfo-canary=50 ] primary=1/2 canary=2/2
06:33:37 phase=Finalising   weight=0   vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=2/2
06:34:04 phase=Succeeded    weight=0   vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=/0

388 seconds from set image to Succeeded. The sequence is: scale the canary Deployment up and wait for it to be ready, run the pre-rollout webhook, then advance the VirtualService weight by stepWeight per interval while the metrics pass, and at maxWeight copy the canary's pod template into podinfo-primary, wait for that Deployment's own rolling update (the 1/2 and 3/2 samples), set the route back to 100/0 and scale the canary to zero. At the end podinfo-primary runs 6.9.0 and podinfo is 0/0 with the same image.

Prometheus confirms the split. Only the weighted generator (20 requests per second) reaches the primary, so the primary's request rate should be 20 x (1 – weight). Sampled at 30 s: 15.8, 16.2, 14.6, 11.7, 9.7 and 7.1 requests per second as the weight went 10, 20, 30, 40, 50 (expected 18, 16, 14, 12, 10; the samples are noisy because each hey run lasts 30 s and the sampling points fall on its edges). Over the run: 4831 requests to the canary and 2341 to the primary, all HTTP 200.

The mistake: the healthy rollout finished with four failed checks

Status:
  Failed Checks:           4
  Phase:                   Succeeded
Events:
  Warning  Halt advancement no values found for istio metric request-success-rate probably podinfo.test is not receiving traffic
  Warning  Halt advancement no values found for istio metric request-duration probably podinfo.test is not receiving traffic   (x3)

Four halts against threshold: 5. One more and a perfectly healthy candidate would have been rolled back, and the events would have said "failed checks threshold reached", which is what an on-call engineer would read as "the new version is bad".

The first halt is the first check after weight 10: the canary had been receiving traffic for less than one Prometheus scrape interval, so the 1-minute success-rate window was empty. That one happened in all three recorded rollouts. The other three are the request-duration metric with interval: 30s. Flagger's built-in Istio query for that metric takes a rate over the metric's interval, the Istio addon Prometheus scrapes every 15 seconds, and a 30-second window with scrape jitter regularly contains fewer than the two samples a rate needs. Result: "no values found", which Flagger counts as a failed check.

The fix is to make the metric window at least a few scrape intervals long. The same rollout with request-duration at interval: 1m (the Canary shown above) was run as a second healthy candidate (6.10.0):

06:50:06 phase=Progressing  weight=10  failed=0
06:50:33 phase=Progressing  weight=10  failed=1     <- first-check "no values" for request-success-rate
06:51:05 phase=Progressing  weight=20  failed=1
06:51:38 phase=Progressing  weight=30  failed=1
06:52:05 phase=Progressing  weight=40  failed=1
06:52:37 phase=Progressing  weight=50  failed=1
06:53:04 phase=Promoting    weight=50  failed=1
06:54:03 phase=Succeeded    weight=0   failed=1

One failed check, none from request-duration, and 304 seconds instead of 388 because no step had to be retried. One run per configuration is not a proof, but it matches the explanation, and the general rule stands whatever the provider: know your scrape interval, and give every metric query a window that is a multiple of it. The threshold you choose should leave room for the first-check halt.

A failing candidate, rolled back

Trigger: set image to podinfo:6.9.2, and at the same time, from inside the load tester pod:

hey -z 8m -q 10 -c 2 http://podinfo-canary.test:9898/status/500

podinfo returns HTTP 500 on that path, so the canary's sidecar reports a low success rate while the primary sees only 200s. This is Flagger's own way of simulating a bad release, and it is what the old article's "simulate a failure and observe Flagger reverting" was missing.

06:38:33 phase=Progressing  weight=0   failed=0  vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=/2
06:39:38 phase=Progressing  weight=10  failed=0  vs[podinfo-primary=90 podinfo-canary=10 ] primary=2/2 canary=2/2
06:40:05 phase=Progressing  weight=10  failed=1  vs[podinfo-primary=90 podinfo-canary=10 ]
06:40:37 phase=Progressing  weight=10  failed=2
06:41:04 phase=Progressing  weight=10  failed=3
06:41:36 phase=Progressing  weight=10  failed=4
06:42:04 phase=Progressing  weight=10  failed=5
06:42:36 phase=Failed       weight=0   failed=0  vs[podinfo-primary=100 podinfo-canary=0 ] primary=2/2 canary=/0

From the Flagger log, one line per 30-second check:

06:40:03 Halt podinfo.test advancement success rate 0.00% < 99%
06:40:33 Halt podinfo.test advancement success rate 17.21% < 99%
06:41:03 Halt podinfo.test advancement success rate 22.84% < 99%
06:41:33 Halt podinfo.test advancement success rate 20.27% < 99%
06:42:03 Halt podinfo.test advancement success rate 23.02% < 99%
06:42:33 Rolling back podinfo.test failed checks threshold reached 5
06:42:33 Canary failed! Scaling down podinfo.test

250 seconds from set image to Failed. The weight never went past 10, so at most a tenth of the weighted traffic touched the bad version for about three minutes; the VirtualService is back at 100/0, the canary Deployment is at zero replicas, and podinfo-primary still runs 6.9.0. Prometheus for the window: 12158 HTTP 500 and 5598 HTTP 200 on the canary workload, 3712 requests on the primary, all 200. The primary never served an error.

Note what "rollback" means here: Flagger does not roll the primary back, because the primary was never changed. It stops sending traffic to the canary and scales it down. The failed image stays in the podinfo Deployment spec; the next set image to a different tag starts a new analysis.

How to watch a rollout

  • kubectl -n test get canary podinfo shows the phase and weight; kubectl -n test describe canary podinfo has the events, but Kubernetes merges similar events, so the per-check values are in Flagger's log (kubectl -n istio-system logs deploy/flagger).
  • The live weights are in the VirtualService: kubectl -n test get virtualservice podinfo -o jsonpath=&#39;{.spec.http[0].route[*].weight}&#39;.
  • kubectl rollout status deployment/podinfo, which the old article recommended, reports on the canary Deployment's own pod rollout, which finishes within seconds of the scale-up; it says nothing about the analysis.
  • The weighted split itself can be checked in Prometheus with sum by (destination_workload) (rate(istio_requests_total{reporter=&quot;destination&quot;,destination_workload_namespace=&quot;test&quot;}[30s])).

What went wrong in the lab, and what it teaches

Two stages of the run script were botched and are left in the log. The first version of the failing-candidate stage did not wait for the phase to leave Succeeded before waiting for a terminal phase, so it returned immediately while Flagger was still detecting the new revision; the 6.9.1 rollout it started ran unwatched and was rolled back on the same 500s. Then, killing the local kubectl exec ... hey did not stop hey inside the load tester pod, so the 500 generator was still running when the next healthy candidate started, and that one was rolled back too, with success rates between 0 and 62 percent. Both are worth knowing if you script this: the Canary phase is stateful across rollouts, and load generators live in the cluster, not in your terminal.

What was not verified

  • An ingress gateway in front of the mesh (service.gateways / hosts in the Canary). The run kept all traffic inside the mesh.
  • Istio demo or default profiles; only minimal was installed.
  • The request-duration metric failing on its own; podinfo answers in a few milliseconds, so only the "no values" case was seen.
  • An HPA on the target (autoscalerRef); the HPA from the podinfo overlay was removed because kind has no metrics-server.
  • Session affinity, header-based A/B routing, mirroring, and Flagger's non-Istio providers.
  • Multi-node clusters and promotion timing with large images.

Sources

  • Flagger v1.45.0 Istio tutorial: https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/tutorials/istio-progressive-delivery.md
  • Flagger v1.45.0, how it works (canary service, primary copy): https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/usage/how-it-works.md
  • Flagger v1.45.0 webhooks and metrics: https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/usage/webhooks.md and https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/usage/metrics.md
  • Flagger v1.45.0 install with Kustomize: https://github.com/fluxcd/flagger/blob/v1.45.0/docs/gitbook/install/flagger-install-on-kubernetes.md
  • Istio 1.31.0 release: https://github.com/istio/istio/releases/tag/1.31.0
  • kind v0.33.0 release (node image digests): https://github.com/kubernetes-sigs/kind/releases/tag/v0.33.0
  • podinfo: https://github.com/stefanprodan/podinfo