Automating AI Model Deployment: A Pipeline You Can Run on kind, and Two Ways the Usual Workflow Silently Does Nothing

Pipeline run on kind v0.33.0 (kindest/node v1.34.11), Docker 27.4.0, kubectl v1.30.5, Python 3.14.7, NumPy 2.4.4, registry:2.8.3; 15 unit tests; Terraform 1.16.2 with hashicorp/aws 6.64.0 and actionlint 1.7.12 used for static validation only, nothing applied to AWS

Revision note (2026-09-15). The earlier version of this article was an overview with two Terraform snippets and one GitHub Actions workflow that had never been run. Checked against current tooling: the Terraform set acl = "private" on aws_s3_bucket, which the AWS provider (6.64.0) reports as deprecated, and interpolated a Secrets Manager value into EC2 user_data, where AWS documents it is readable by anyone on the instance. The workflow used actions/checkout@v3 and actions/setup-python@v4, which actionlint rejects because their node16 runtime is no longer executed; targeted Python 3.9, which reached end of life in October 2025; logged in to ECR with no credentials step; ran terraform apply -auto-approve on every push with no backend, so the state was lost with the runner; put the kubeconfig's contents into the KUBECONFIG variable, which kubectl treats as a list of file paths; and deployed a :latest tag with kubectl set image, which does nothing once the Deployment already points at :latest. This version is rebuilt around a pipeline that actually runs, on a local kind cluster, with every output below pasted from the recorded run. The AWS and GitHub Actions parts could not be executed here and are labelled as statically validated.

What a deployment pipeline has to prove

For a model service, "deployed" means one thing: the pods answering requests run the model file the pipeline just produced, and you can show it. Everything else in this article is in service of that check. Two things in the original workflow made it impossible: a mutable image tag, so the Deployment could not say which model it referenced, and no verification step, so a no-op rollout looked identical to a real one.

The lab that replaces the original snippets is at examples/ai-model-deploy-pipeline in the site's repository. Tested environment: Docker 27.4.0 (linux/arm64), kind v0.33.0 with node image kindest/node:v1.34.11 (pinned by digest, containerd 2.3.4), kubectl v1.30.5, Python 3.14.7 with NumPy 2.4.4 on the host, registry:2.8.3 as a local registry on 127.0.0.1:5001, python:3.13-slim as the base image. The whole run, including cluster creation and teardown, took about three minutes.

The pieces

The model is a two-feature logistic regression trained with NumPy from a seed, serialised as JSON rather than pickle so that loading it cannot execute code. The point of the model is not accuracy; it is that every training run with a different seed produces a different, hashable file:

$ python3 model/train.py v1 build/model.json --seed 0
wrote build/model.json model_version=v1 train_accuracy=0.888 sha256=fb02f93891349b125df0c31918c900d6ccb4fbc09e11f7d8d8f7f13795c71f25

The service is a standard-library HTTP server with three routes. GET /version is the one the pipeline depends on: it returns the model version, the SHA-256 of the model file the process actually loaded, the image tag it was built with, and the pod hostname. POST /predict takes {"features": [x0, x1]} and returns a probability and a label; GET /healthz is the readiness probe. The server also handles SIGTERM, because a PID 1 process that ignores it holds each old pod for the full 30-second grace period during a rollout.

Fifteen unit tests run before anything is built: five on the trainer (same seed gives a byte-identical artifact, different seed gives different weights, accuracy above 0.8, JSON round trip, zero samples rejected) and ten on the server (version payload, predictions on both sides of the boundary and on it, 400 for the wrong feature count, non-numeric features and malformed JSON, 404 for unknown paths, SIGTERM exits 0 in under three seconds, missing or inconsistent model file fails at startup). Both suites passed in both pipeline runs:

$ python3 -m unittest discover -s app -p test_*.py
Ran 10 tests in 1.191s
OK
$ python3 -m unittest discover -s model -p test_*.py
Ran 5 tests in 0.029s
OK

The Dockerfile copies the server and build/model.json into python:3.13-slim (by digest), runs as an unprivileged user, and passes a build argument through to /version:

FROM python:3.13-slim@sha256:9d2e5553305c7c7b0097999bb17187c69b921ccd6bc9d40e4bb5ebe652c00285
ARG IMAGE_TAG=unknown
ENV IMAGE_TAG=${IMAGE_TAG} MODEL_PATH=/app/model.json PORT=8080 PYTHONUNBUFFERED=1
WORKDIR /app
COPY app/server.py /app/server.py
COPY build/model.json /app/model.json
RUN useradd --system --uid 10001 model && chown -R model:model /app
USER 10001
EXPOSE 8080
CMD ["python", "/app/server.py"]

The tag is the contract

The pipeline script tags each image g<git short sha>-m<first 12 hex of sha256(model.json)>. Code and model are both in the name, and neither can change without the name changing:

GIT_SHA=$(git rev-parse --short=7 HEAD)
MODEL_SHA=$(shasum -a 256 build/model.json | cut -c1-12)
TAG="g${GIT_SHA}-m${MODEL_SHA}"
docker build --build-arg "IMAGE_TAG=${TAG}" -t "localhost:5001/ai-model:${TAG}" .
docker push "localhost:5001/ai-model:${TAG}"

Deploying is kubectl apply on the first run and kubectl set image afterwards, followed by rollout status with a timeout. Then comes the step the original workflow did not have. A verifier runs inside the cluster (a python:3.13-slim pod) and polls the Service's /version every 100 ms for up to 20 seconds, until 20 consecutive answers from at least two distinct pods report the expected model hash and image tag. It prints every pod that answered and marks any that returned something else as stale, so a half-finished rollout is visible rather than averaged away.

Two runs and a rollback

First run, model v1:

=== pipeline v1 (2026-09-15T08:25:00Z)
image tag: g2748079-mfb02f9389134
pushed: localhost:5001/ai-model@sha256:fba4794912efc2f839aa5f7a41dc8fed841db6a76a5500c810e9bddd6e50e673
deployment.apps/ai-model created
service/ai-model created
deployment "ai-model" successfully rolled out
REVISION  CHANGE-CAUSE
1         <none>
ai-model-79c6bc88d9-kmnmr:   8 answers, last at +2.00s  model_version=v1 image_tag=g2748079-mfb02f9389134 sha=fb02f9389134  [expected]
ai-model-79c6bc88d9-cwlm9:  12 answers, last at +2.42s  model_version=v1 image_tag=g2748079-mfb02f9389134 sha=fb02f9389134  [expected]
predict [1.0, -0.5] -> {"probability": 0.9952161064141438, "label": 1}
VERIFY OK: expected tag g2748079-mfb02f9389134, model sha fb02f9389134; connection errors=4

The four connection errors are the first requests, sent before the Service had ready endpoints; the verifier counts them and keeps polling.

Second run, model v2 (a different seed and 3000 samples instead of 2000):

=== pipeline v2 (2026-09-15T08:25:16Z)
wrote build/model.json model_version=v2 train_accuracy=0.894 sha256=81fc74abbb1fbc1f0140c130b3487fc56f2b65eadf3eb05d50861c49b0b82fe9
image tag: g2748079-m81fc74abbb1f
$ kubectl -n ai-model set image deploy/ai-model model=localhost:5001/ai-model:g2748079-m81fc74abbb1f
deployment.apps/ai-model image updated
deployment "ai-model" successfully rolled out
REVISION  CHANGE-CAUSE
1         <none>
2         <none>
ai-model-d88dcf99f-mllfg:  13 answers, last at +5.43s  model_version=v2 image_tag=g2748079-m81fc74abbb1f sha=81fc74abbb1f  [expected]
ai-model-d88dcf99f-xspxs:  10 answers, last at +5.54s  model_version=v2 image_tag=g2748079-m81fc74abbb1f sha=81fc74abbb1f  [expected]
predict [1.0, -0.5] -> {"probability": 0.9948569008229757, "label": 1}
VERIFY OK: expected tag g2748079-m81fc74abbb1f, model sha 81fc74abbb1f; connection errors=2

The prediction for the same input moved from 0.99522 to 0.99486: a small change, but a real one, and the only reason we know the new model is serving is that the hash and the number both changed.

Rollback is one command, and the same verifier proves it:

$ kubectl -n ai-model rollout undo deploy/ai-model
deployment.apps/ai-model rolled back
deployment "ai-model" successfully rolled out
REVISION  CHANGE-CAUSE
2         <none>
3         <none>
$ kubectl -n ai-model get deploy ai-model -o jsonpath={.spec.template.spec.containers[0].image}{"n"}
localhost:5001/ai-model:g2748079-mfb02f9389134
ai-model-79c6bc88d9-g84f9:  14 answers, last at +5.37s  model_version=v1 image_tag=g2748079-mfb02f9389134 sha=fb02f9389134  [expected]
ai-model-79c6bc88d9-swn7x:   8 answers, last at +5.48s  model_version=v1 image_tag=g2748079-mfb02f9389134 sha=fb02f9389134  [expected]
predict [1.0, -0.5] -> {"probability": 0.9952161064141438, "label": 1}
VERIFY OK: expected tag g2748079-mfb02f9389134, model sha fb02f9389134; connection errors=3

Revision 1 vanished from the history because rollout undo re-creates the old template as revision 3, reusing the old ReplicaSet (same 79c6bc88d9 prefix in the pod names). The prediction is bit-identical to run 1. Because the template names the exact image, "roll back" means "run that model again", not "run whatever :latest pointed at back then".

What the original workflow does: :latest three times

The original pipeline built my-ai-model:latest, pushed it, and ran kubectl set image deployment/... :latest on every push. The lab repeats that with three different models.

The first switch to :latest rolls out, because the pod template changed from the immutable tag to :latest (revision 4) and the verifier confirms model latest-A. Then a second model, latest-B, is built and pushed under the same tag, and the same command runs:

registry now serves :latest as sha256:1707e54b3e630f6157b63346d5cd57f06cd25cec6a3a77b56d7228d1da6cf49e
$ kubectl -n ai-model set image deploy/ai-model model=localhost:5001/ai-model:latest; echo "exit=$?"
exit=0
$ kubectl -n ai-model rollout status deploy/ai-model --timeout=60s
deployment "ai-model" successfully rolled out
revision after second :latest deploy: 4 (was 4)
NAME                        CREATED                PHASE     DELETION   PULL_POLICY    IMAGE_ID
ai-model-765d866cf5-dsndf   2026-09-15T08:25:48Z   Running   <none>     IfNotPresent   localhost:5001/ai-model@sha256:272a055c...
ai-model-765d866cf5-kg8rz   2026-09-15T08:25:47Z   Running   <none>     IfNotPresent   localhost:5001/ai-model@sha256:272a055c...
ai-model-765d866cf5-dsndf:  90 answers, last at +19.73s  model_version=latest-A image_tag=latest sha=efb63f2b02a2  [STALE]
ai-model-765d866cf5-kg8rz: 101 answers, last at +19.94s  model_version=latest-A image_tag=latest sha=efb63f2b02a2  [STALE]
VERIFY FAILED: expected tag latest, model sha 08b3924db72d; connection errors=0

set image exited 0 without printing image updated, rollout status reported success, the revision stayed at 4, the two pods are the same two pods, and for the full 20 seconds they answered with latest-A. Every step a CI job would check was green. The Kubernetes documentation states the rule: a Deployment's rollout is triggered if and only if the Deployment's pod template is changed. Setting the image field to the value it already has is not a change.

The next thing people try is kubectl rollout restart. In the lab it did not help either:

template imagePullPolicy=IfNotPresent
$ kubectl -n ai-model rollout restart deploy/ai-model
deployment.apps/ai-model restarted
deployment "ai-model" successfully rolled out
ai-model-77b4dfd4f7-z8f4p:  88 answers, last at +19.75s  model_version=latest-A image_tag=latest sha=efb63f2b02a2  [STALE]
ai-model-77b4dfd4f7-8j5pn: 103 answers, last at +19.96s  model_version=latest-A image_tag=latest sha=efb63f2b02a2  [STALE]
VERIFY FAILED: expected tag latest, model sha 08b3924db72d; connection errors=0

New pods, old model. The Deployment was created with a non-:latest tag, so the API server defaulted imagePullPolicy to IfNotPresent, and set image changes only the image field. The node already had an image called localhost:5001/ai-model:latest, so it used it. The :latest special case in the docs (omit the policy and use :latest, get Always) applies when the object is created, not when the tag is changed later.

Patching the policy to Always is itself a template change, so it rolled out and the new pods pulled the current :latest, which was latest-B; the verifier passed. With Always in place, a third model pushed as :latest still did nothing on its own and needed an explicit rollout restart before the pods served latest-C. The final history had revisions 2 through 7, and revisions 4 to 7 all read image: localhost:5001/ai-model:latest. Nothing in the history says which model each revision ran, so rollout undo from there cannot promise anything.

Conclusion for a model pipeline: tag with something that changes when the artifact changes (commit SHA, or commit plus model hash as here), let set image do the rollout, and verify from inside the cluster that the served model is the one you built. If a registry enforces it, turn on tag immutability so a tag can never be re-pushed; ECR has that setting.

KUBECONFIG holds paths, not a kubeconfig

The original workflow's last step had env: KUBECONFIG: ${{ secrets.KUBECONFIG }}, the file's contents in a variable. The Kubernetes docs define the variable as a colon-delimited list of kubeconfig files. The lab does what the workflow does, with the kind cluster's kubeconfig (19 lines, 5613 bytes) and an empty HOME so no ~/.kube/config can hide the result. kubectl's error echoes the "paths" it tried to open, which are fragments of the file including the base64 certificate and key, so the output below is redacted:

$ KUBECONFIG="$(cat build/kubeconfig)" HOME=build/emptyhome kubectl -n ai-model get deploy ai-model
error loading config file " <base64 redacted>
    server": open  <base64 redacted>
    server: file name too long
error loading config file " <base64 redacted>
    client-key-data": open  <base64 redacted>
    client-key-data: file name too long
error loading config file " <base64 redacted>": open  <base64 redacted>: file name too long
exit code: 1
$ KUBECONFIG=build/kubeconfig HOME=build/emptyhome kubectl -n ai-model get deploy ai-model
NAME       READY   UP-TO-DATE   AVAILABLE   AGE
ai-model   2/2     2            2           115s
exit code: 0

kubectl split the value on every : (a kubeconfig has several, starting with server: https://) and tried to open each fragment as a file. With a shorter value, the first review of this article saw kubectl config view print clusters: null instead of an error: an empty configuration, so kubectl set image had no cluster to address. Either way the step cannot reach the cluster. Write the secret to a file and point KUBECONFIG at the path, or, for EKS, let aws eks update-kubeconfig write the file for you. The corrected workflow below does the latter.

The cloud parts: statically validated, not run

There is no AWS account in this lab, so the Terraform and the GitHub Actions workflow were checked with terraform init, terraform fmt -check, terraform validate (Terraform 1.16.2, hashicorp/aws 6.64.0) and actionlint 1.7.12. That proves they are well-formed against the provider schema and the workflow grammar. It does not prove that terraform apply succeeds, that the IAM policy is sufficient, or that the OIDC trust is configured. Treat this section as a reviewed starting point.

What the original Terraform gets wrong

Running terraform validate on the article's own block, with the provider pinned to 6.64.0:

Warning: Argument is deprecated

  with aws_s3_bucket.model_artifacts,
  on main.tf line 16, in resource "aws_s3_bucket" "model_artifacts":
  16:   acl    = "private"

acl is deprecated. Use the aws_s3_bucket_acl resource instead.
Success! The configuration is valid, but there were some validation warnings

private is also the default, so the line did nothing but trigger the warning. The corrected version drops it and adds aws_s3_bucket_public_access_block, aws_s3_bucket_ownership_controls with BucketOwnerEnforced, and versioning, so old model artifacts are recoverable.

The secrets snippet was worse. It wrote export DB_PASSWORD=${data.aws_secretsmanager_secret_version.db_password.secret_string} into user_data. AWS's instance metadata documentation says user data "is not protected by authentication or cryptographic methods" and that "you should not store sensitive data, such as passwords or long-lived encryption keys, as user data". The value would also land in the Terraform state. The corrected configuration gives the instance an IAM role that may call secretsmanager:GetSecretValue on that one secret and passes only the secret's name:

data "aws_iam_policy_document" "read_secret" {
  statement {
    actions   = ["secretsmanager:GetSecretValue"]
    resources = [data.aws_secretsmanager_secret.db_password.arn]
  }
}

resource "aws_instance" "gpu" {
  ami                  = var.gpu_ami_id
  instance_type        = "g4dn.xlarge"
  iam_instance_profile = aws_iam_instance_profile.inference.name

  metadata_options {
    http_tokens = "required" # IMDSv2 only
  }

  user_data = <<-EOT
    #!/bin/bash
    echo "DB_PASSWORD_SECRET_ID=${data.aws_secretsmanager_secret.db_password.name}" >> /etc/environment
  EOT
}

The third problem was not in the snippet but in how the workflow used it: terraform init and terraform apply -auto-approve on every push, with no backend block. Terraform's default is the local backend, which writes terraform.tfstate next to the configuration, on the runner's disk. When the job ends the state is gone, and the next run has no record of the instance and bucket it created; it will try to create them again, and the fixed bucket name will collide. The corrected versions.tf pins the provider and declares an S3 backend with use_lockfile = true:

terraform {
  required_version = ">= 1.6.0, < 2.0.0"
  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "6.64.0"
    }
  }
  backend "s3" {
    bucket       = "example-terraform-state"
    key          = "ai-model-deploy/terraform.tfstate"
    region       = "us-west-2"
    use_lockfile = true
  }
}

The corrected files validate cleanly (terraform fmt -check exit 0, terraform validate prints Success! The configuration is valid.). Nothing was planned or applied.

What the original workflow gets wrong

actionlint on the article's workflow, verbatim:

ci-cd-pipeline.yml:14:15: the runner of "actions/checkout@v3" action is too old to run on GitHub Actions. update the action's version to fix this issue [action]
ci-cd-pipeline.yml:17:15: the runner of "actions/setup-python@v4" action is too old to run on GitHub Actions. update the action's version to fix this issue [action]

Those two majors, and aws-actions/amazon-ecr-login@v1, declare using: node16 in their action.yml; GitHub moved actions to Node 20 in 2024 and the current majors (checkout@v7, setup-python@v7, amazon-ecr-login@v2, checked on 2026-09-15) declare node24. python-version: 3.9 targets a release that reached end of life in October 2025.

Two problems actionlint cannot see. The Log in to AWS ECR step ran with no credentials at all: the AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY block was attached to the later Terraform step, and env on a step is scoped to that step. The ECR login action's README requires a preceding configure-aws-credentials step. And the credentials themselves were long-lived access keys in GitHub secrets, where the same README recommends assuming a role through GitHub's OIDC provider.

The corrected workflow (actionlint exit 0):

permissions:
  id-token: write   # OIDC token for configure-aws-credentials
  contents: read

jobs:
  build-and-deploy:
    runs-on: ubuntu-24.04
    steps:
      - uses: actions/checkout@v7
      - uses: actions/setup-python@v7
        with:
          python-version: "3.13"
      - name: Unit tests
        run: |
          python -m unittest discover -s app -p 'test_*.py'
          python -m unittest discover -s model -p 'test_*.py'
      - name: Configure AWS credentials (OIDC)
        uses: aws-actions/configure-aws-credentials@v6
        with:
          role-to-assume: arn:aws:iam::123456789012:role/github-ai-model-deploy
          aws-region: ${{ env.AWS_REGION }}
      - name: Log in to ECR
        id: ecr
        uses: aws-actions/amazon-ecr-login@v2
      - name: Build and push
        id: build
        env:
          REGISTRY: ${{ steps.ecr.outputs.registry }}
          TAG: ${{ github.sha }}
        run: |
          python model/train.py "$TAG" build/model.json
          docker build --build-arg "IMAGE_TAG=$TAG" -t "$REGISTRY/$ECR_REPOSITORY:$TAG" .
          docker push "$REGISTRY/$ECR_REPOSITORY:$TAG"
          echo "image=$REGISTRY/$ECR_REPOSITORY:$TAG" >> "$GITHUB_OUTPUT"
      - name: Point kubectl at the cluster
        run: aws eks update-kubeconfig --name "$EKS_CLUSTER" --region "$AWS_REGION"
      - name: Roll out and wait
        env:
          IMAGE: ${{ steps.build.outputs.image }}
        run: |
          kubectl -n ai-model set image deployment/ai-model "model=$IMAGE"
          kubectl -n ai-model rollout status deployment/ai-model --timeout=180s
      - name: Roll back if the rollout failed
        if: failure()
        run: kubectl -n ai-model rollout undo deployment/ai-model

The role ARN and cluster name are placeholders. terraform apply is deliberately absent from the deploy job: infrastructure changes and model releases have different blast radii and different reviewers, and a plan that is applied automatically on every model commit is how an unrelated push resizes your GPU instance. Run Terraform from its own workflow, on changes to its own directory, with a plan step that someone reads.

What this does not cover

  • Anything on AWS was not executed: no terraform plan or apply, no ECR push, no EKS cluster, no OIDC trust. The Terraform and the workflow are validated, not tested.
  • The kind cluster has one node. "The image was already cached on the node" is the only cache scenario exercised; a multi-node cluster adds the case where some nodes have the stale :latest and some do not.
  • Progressive delivery. The rollout here is the Deployment default (one surge pod, one unavailable); canary or blue-green strategies and model-quality gates before promotion are separate topics.
  • Model quality. The accuracy figures printed by the trainer are training-set accuracy on synthetic data and mean nothing beyond "the file changed".
  • Training in CI. The lab trains in the pipeline because it takes a fraction of a second; real training runs belong in an orchestrator, and the pipeline should pick up a registered artifact by hash.

Reproduce it

Requirements: Docker, kind, kubectl, Python 3 with NumPy, git. The scripts use 127.0.0.1:5001, a kind cluster named dd-kind-b, and a registry container named dd-registry-a.

cd examples/ai-model-deploy-pipeline
./run-all.sh          # cluster + registry, pipeline v1, v2, rollback, both demos, teardown; ~3 min
./terraform/validate.sh && ./github/validate.sh && ./original/validate.sh   # static checks only

Sources