Revision note (2026-09-15). The earlier version of this article was a Flask endpoint that started one daemon thread per request, wrote the candidate's output to a log file, and stopped. It never logged the production output next to the candidate output, carried no request id, and proposed no metric, so the "offline analysis" it promised had nothing to compare. Its thread-per-request design has no bound, and Python daemon threads are killed at interpreter exit, which loses whatever the shadow was still scoring. It also targeted Python 3.8 and Flask 1.x (both end-of-life) and listed three sources without links that could not be located. This version replaces the endpoint with a small, seeded simulation in which every number below is produced by the code shown.
What shadow testing can and cannot tell you
Shadow testing sends a copy of each production request to a candidate model, discards the candidate's answer as far as the user is concerned, and records it for comparison. It answers one question well: given the traffic production actually sees, where and how often does the candidate answer differently, and when labels exist, who was right? It does not answer what would happen to users, revenue, or downstream systems if the candidate's answers were acted on. That gap is where shadow verdicts go wrong, and the lab below reproduces two such cases with numbers.
Tested versions: Python 3.14.7, NumPy 2.4.4. Nothing else was installed. The lab is examples/ai-shadow-testing in the companion repository; python3 -m unittest -v test_shadow_lab.py runs 9 tests and python3 shadow_lab.py --seed 0 --json results-seed0.json reproduces the tables.
The lab
Two traffic segments, A and B, share three features and disagree on three, and segment B inputs are shifted. The label is a Bernoulli draw from the segment's own logistic rule, so no linear model is right for both segments at once.
- Production model: logistic regression trained on 20,000 historical requests in which segment B was 5 percent of traffic.
- Candidate model: logistic regression trained on 20,000 requests in which segment B was 40 percent.
- Shadow window: 20,000 fresh requests, both models scored on the same rows.
The candidate is therefore better on B and worse on A by construction. Whether it is "better" depends entirely on the traffic mix, which is the point.
def make_stream(rng, n, share_b):
share = np.broadcast_to(np.asarray(share_b, dtype=float), (n,))
seg_b = rng.random(n) < share
x = rng.standard_normal((n, N_FEATURES))
x[seg_b, :3] += 0.7
w_a, w_b = _true_weights()
logit = np.where(seg_b, x @ w_b, x @ w_a)
y = (rng.random(n) < 1.0 / (1.0 + np.exp(-logit))).astype(int)
features = np.hstack([x, seg_b[:, None].astype(float)])
return features, y, seg_b
The report a shadow run should produce
Every request yields a pair of predictions, so the comparison is paired. The report the lab computes per window:
- Agreement rate and disagreement share per segment. Disagreements are where the candidate would change behavior; everything else is invisible to users.
- Paired accuracy delta with a 95 percent bootstrap interval (2,000 resamples of request indices, so each resample keeps both models' answers for the same row).
- Discordant pairs: b = production right and candidate wrong, c = production wrong and candidate right. Agreements cancel; only b and c carry information about which model is better.
- Exact McNemar test on (b, c): under the null each discordant pair is a fair coin, so the p-value is the two-sided binomial tail. Computed in log space because b + c runs into the thousands.
- Accuracy restricted to disagreements, which is what a reviewer looking at a sample of "the candidate did something different" would see.
- Mean confidence and ECE (15 equal-width bins) per model, and the mean absolute gap between the two probability outputs.
def mcnemar_exact(b, c):
n = b + c
if n == 0:
return 1.0
k = min(b, c)
log_terms = [math.lgamma(n + 1) - math.lgamma(i + 1) - math.lgamma(n - i + 1) - n * math.log(2.0)
for i in range(k + 1)]
peak = max(log_terms)
log_tail = peak + math.log(sum(math.exp(t - peak) for t in log_terms))
return float(min(1.0, 2.0 * math.exp(log_tail)))
Baseline: a fully labeled window, 20 percent segment B
| Seed | Agreement | Acc prod | Acc cand | Delta (cand minus prod) | 95% CI | b / c | McNemar p |
|---|---|---|---|---|---|---|---|
| 0 | 0.8568 | 0.7708 | 0.7468 | -0.0241 | [-0.0291, -0.0186] | 1673 / 1192 | 2.5e-19 |
| 1 | 0.8608 | 0.7716 | 0.7469 | -0.0247 | [-0.0301, -0.0197] | 1639 / 1145 | 7.4e-21 |
Per segment (seed 0): on A production 0.815 versus candidate 0.760; on B production 0.593 versus candidate 0.694. Disagreement share is 14.6 percent of A requests and 13.1 percent of B requests. On the disagreements alone production is right 58.4 percent of the time and the candidate 41.6 percent. Mean confidence 0.791 (production) versus 0.728 (candidate), ECE 0.028 versus 0.018, mean absolute probability gap 0.122. Seed 1 is within a point of every figure.
With 20 percent B traffic and full labels, the shadow verdict is unambiguous: reject the candidate. Both of the cases below start from this same pair of models.
Case 1: label delay plus a traffic-mix shift
Labels rarely arrive with the request. In this scenario the window's segment-B share ramps linearly from 5 percent to 80 percent over 20,000 requests, and at decision time only the first half of the window has labels. The decision is made on the labeled prefix; the "online" truth is the full window once every label has landed.
| Seed | Window | n | Share B | Delta | 95% CI | b / c | McNemar p |
|---|---|---|---|---|---|---|---|
| 0 | labeled prefix | 10,000 | 0.235 | -0.0124 | [-0.0200, -0.0051] | 771 / 647 | 1.1e-3 |
| 0 | full window | 20,000 | 0.427 | +0.0164 | [+0.0109, +0.0216] | 1267 / 1595 | 9.4e-10 |
| 1 | labeled prefix | 10,000 | 0.233 | -0.0171 | [-0.0245, -0.0103] | 770 / 599 | 4.2e-6 |
| 1 | full window | 20,000 | 0.425 | +0.0145 | [+0.0094, +0.0195] | 1238 / 1527 | 4.2e-8 |
The labeled prefix says the candidate is significantly worse. The full window says it is significantly better. Both are statistically clean; neither is a bug. The per-segment accuracies barely move between the two windows (seed 0, segment A: 0.807 versus 0.803 for production, 0.757 versus 0.752 for the candidate), so nothing about either model changed. Only the mix changed, and the labeled sample was drawn from the older, A-heavy part of it.
What to do about it: report the segment mix of the labeled sample next to the mix of the full request stream, and report per-segment deltas rather than one aggregate. A single aggregate delta on delayed labels is a verdict about last month's traffic.
Case 2: labels only exist for what production approved
In approve/decline, fraud, lending, or content-moderation settings, the outcome is observed only for requests production let through. The candidate gets graded on production's choices. The lab reproduces this by keeping labels only where production predicted the positive class.
| Seed | Rows graded | n | Share B | Delta | 95% CI |
|---|---|---|---|---|---|
| 0 | production positives only (what the shadow sees) | 10,680 | 0.246 | -0.0317 | [-0.0394, -0.0252] |
| 0 | production negatives only (what the shadow never sees) | 9,320 | 0.153 | -0.0186 | [-0.0262, -0.0104] |
| 0 | all requests | 20,000 | 0.202 | -0.0256 | [-0.0309, -0.0201] |
| 1 | production positives only | 10,342 | 0.243 | -0.0272 | [-0.0342, -0.0203] |
| 1 | production negatives only | 9,658 | 0.156 | -0.0139 | [-0.0218, -0.0064] |
| 1 | all requests | 20,000 | 0.201 | -0.0208 | [-0.0261, -0.0156] |
Here the verdict does not flip, but the size is wrong: the observable subset overstates the candidate's deficit (seed 0: -0.032 measured versus -0.026 true), and the half of traffic production declined, where the candidate is closest to production, is exactly the half no shadow log will ever have a label for. The observed subset is also a different population (24.6 percent segment B versus 15.3 percent among the declined). Whether this bias helps or hurts a given candidate depends on the models; what is general is that a shadow estimate on selectively labeled traffic is an estimate on production's decisions, not on the request stream.
What to do about it: either accept that shadow testing measures agreement rather than outcome in these settings, or budget a small randomized exploration slice where the candidate's decision is acted on. That is a canary, not a shadow, and it has user-facing risk, which is the trade the shadow was avoiding.
The serving side: the shadow must be allowed to fall behind
The earlier version of this article started a thread per request. The lab replaces that with a bounded queue and one worker: when the queue is full the request is dropped and counted, and any exception in the candidate is counted and swallowed.
class ShadowRunner:
def __init__(self, candidate, maxsize=64):
self.queue = queue.Queue(maxsize=maxsize)
...
def submit(self, request_id, features):
try:
self.queue.put_nowait((request_id, features))
except queue.Full:
with self._lock:
self.dropped += 1
Measured with production paced at 1 ms per request and a candidate that takes 2 ms, 2,000 requests, queue of 64:
| Seed | Production wall time | Shadow scored | Shadow dropped | Crashing candidate |
|---|---|---|---|---|
| 0 | 2.820 s | 1124 | 876 | 2,000 failures counted, production output identical |
| 1 | 2.566 s | 1090 | 910 | 2,000 failures counted, production output identical |
Production finished in roughly the time its own pacing implies; the candidate, at half the throughput, scored a little over half the stream and the rest was dropped rather than queued without limit. The two wall times differ because these are time.sleep loops on a laptop; they are not a latency benchmark of anything. The invariant that matters is asserted by a test: a candidate that raises on every call leaves the production outputs byte-for-byte identical.
Two consequences for the analysis: dropped requests must be random with respect to the traffic, or the shadow sample is biased in a third way (a queue that fills during peak hours under-samples peak traffic), and every record needs a request id so the production answer, the candidate answer, and the eventual label can be joined. The earlier article's log line contained the input and the candidate output only.
Where the mirroring usually lives
The lab does the duplication in-process because that is what can be run here. In practice it is usually done by infrastructure, and the following was checked against current documentation only, not executed:
- Istio 1.31
VirtualServicehasmirror(host and subset) andmirrorPercentage(avaluein percent; when absent, all traffic is mirrored). The docs state the mirrored requests are "fire and forget", responses are discarded, and the mirrored request's Host/Authority header gets a-shadowsuffix, which is how the candidate can tell it is a shadow copy. - Seldon Core 2 (v2.10.2) has an
Experimentresource with amirrorblock (name,percent) next to the weightedcandidates; the documentation states responses from the mirror model are not returned to the caller. - KServe (v0.20.0) documents canary rollouts by traffic percentage; a request for native shadow deployments (issue 2978, filed 2023) was closed as not planned. Shadowing an
InferenceServicemeans using the mesh's mirroring underneath it.
None of these give you the comparison. They give you two answer streams; joining them with a request id and with delayed labels, and reading the join with the two cases above in mind, is still your job.
What this lab is not
It is a six-feature linear model on Gaussian data. Real drift is not a linear ramp, real label delay is not a clean cutoff, and real selective labeling has its own feedback loops. It measures no real latency and touches no serving framework. Its value is that the three effects above, a verdict that flips with the labeled window, a biased estimate under selective labels, and a shadow that drops rather than blocks, are each reproduced by a script that runs in seconds and by tests that pin them.
Reproduce it
python3 -m unittest -v test_shadow_lab.py # 9 tests, OK
python3 shadow_lab.py --seed 0 --json results-seed0.json
python3 shadow_lab.py --seed 1 --json results-seed1.json
Sources
- Istio 1.31: Traffic mirroring task for
mirror,mirrorPercentage, the discarded responses, and the-shadowhost suffix. - Seldon Core 2 user guide: Experiments and the mirror sample manifest.
- KServe issue 2978: Shadow Deployments, closed as not planned.
- Python threading documentation: "Daemon threads are abruptly stopped at shutdown."
- Python release status: 3.8 and 3.9 are end-of-life.
- McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika.
- NumPy 2.4 documentation
