Topic archive
AI Engineering
AI tools, prompt engineering, and practical AI usage guides.
-
Temperature Scaling Under Data Drift: What Recalibration Fixes and What It Cannot
A reproducible NumPy lab: temperature scaling never changes accuracy, a temperature fitted on clean data stops working after covariate drift, and under concept drift the best temperature runs away because the model is wrong.

-
Shadow Testing an AI Model: The Metrics, and the Two Ways the Shadow Lies
A reproducible NumPy lab that scores a production and a candidate model on the same request stream, computes the paired comparison a shadow run should produce, and shows two cases where the shadow verdict is wrong: label delay under a traffic-mix shift, and labels that only exist for what production approved.

-
Automating AI Model Deployment: A Pipeline You Can Run on kind, and Two Ways the Usual Workflow Silently Does Nothing
A train-build-push-rollout-verify-rollback pipeline run against a kind cluster and a local registry, with recorded output. It shows why a model image needs an immutable tag (three different models pushed as :latest gave a green pipeline and stale pods), why KUBECONFIG must be a path, and what in the Terraform and GitHub Actions parts could only…

-
Caching AI Model Inference: What a Hit Costs, When the Answer Goes Stale, and How a Semantic Cache Lies
A reproducible NumPy lab that puts an exact-match cache, a TTL, a model redeploy, a thread stampede and a cosine-similarity semantic cache in front of a deterministic model, measures hit rate, model calls and p50/p95 on a synthetic request stream, and shows the one number a semantic cache must report that a hit rate hides:…

-
Compressing a Model for the Edge: What int8 and Pruning Actually Cost, Measured
A reproducible NumPy and onnxruntime lab that quantizes a small MLP to int8 per-tensor and per-channel, prunes it to 98 percent, and measures bytes, held-out accuracy and latency at every step, including the one weight rewrite that leaves the float model identical and drops int8 accuracy from 82 to 12 percent, and the reason a…

-
Building Scalable Multi-Modal AI Pipelines Combining Vision and Language Models
Explore building scalable AI pipelines integrating vision and language models with practical tips, code examples, and best practices.

-
Implementing AI-Driven Feature Engineering Pipelines with Automated Data Transformation
Learn how to build scalable AI-driven feature engineering pipelines automating data transformation, feature selection, and model integration using Python with practical code examples and troubleshooting guidance.

-
Implementing Federated Learning Workflows for Privacy-Preserving AI Deployment
A practical engineering guide to implementing federated learning workflows using TensorFlow Federated, offering an end-to-end example, testing strategies, failure handling, and privacy considerations.
