Temperature Scaling Under Data Drift: What Recalibration Fixes and What It Cannot

Lab run with Python 3.14.7 and NumPy 2.4.4 for seeds 0 and 1; all numbers in the article come from those runs

Revision note (2026-09-15). The earlier version of this article fitted a temperature on randomly generated logits and randomly generated labels, printed the temperature, and stopped. It reported no metric before or after, and it claimed calibration "improves accuracy", which a positive temperature cannot do. This version replaces that example with a small, seeded experiment whose every number is reproduced by the code shown.

The question

A classifier is deployed. Months later the inputs look different from the training data. Someone proposes "continuous calibration": keep re-fitting a temperature on recent data so the predicted probabilities stay honest. Is that a fix, a patch, or a distraction? The answer depends on what changed, and the lab below separates the cases.

Calibration is not accuracy

Temperature scaling divides every logit vector by one positive scalar T before the softmax. Because division by a positive constant preserves the order of the logits, the arg max of the probabilities does not move. Accuracy is a function of the arg max only, so accuracy before and after temperature scaling is identical, always. The lab asserts this in code rather than assuming it.

What temperature scaling changes is the confidence. T greater than 1 flattens the distribution, T less than 1 sharpens it. The right metrics are the ones that look at probabilities:

  • Negative log-likelihood (NLL): mean of the negative log probability assigned to the true class.
  • Brier score: mean squared error between the probability vector and the one-hot label.
  • Expected calibration error (ECE): bin predictions by top-class confidence (15 equal-width bins here) and average the gap between confidence and accuracy per bin, weighted by bin size.

Lab setup

Everything is synthetic and seeded, so the numbers reproduce exactly. The code is a single NumPy script (Python 3.14.7, NumPy 2.4.4, no other dependency).

  • Data: 3 classes, 20 features, each class a Gaussian around its own centroid, unit noise.
  • Model: softmax regression trained by full-batch gradient descent, 3,000 epochs, no regularization, on only 240 samples. This deliberately overfits so the model is over-confident, which is the situation temperature scaling exists for.
  • Splits: 2,000 validation samples to fit T, 4,000 in-distribution test samples to evaluate.
  • Fitting T: minimize validation NLL over T in [0.05, 50] with golden-section search on log T. The objective is convex in log T, so the 1-D search is sufficient and needs no optimizer library.
  • Drift scenario A, covariate drift: same centroids, noise raised from 1.0 to 1.8. Inputs are noisier, the true class boundary is the same.
  • Drift scenario B, concept drift: the centroids of classes 0 and 1 are swapped. The learned decision rule is now wrong for two of three classes.
  • For each drift scenario, T is also refitted on a small "recent" labeled batch of 500 drifted samples, which is what a continuous calibration job would have.

The essential code:

def softmax(logits):
    shifted = logits - logits.max(axis=1, keepdims=True)
    exp = np.exp(shifted)
    return exp / exp.sum(axis=1, keepdims=True)

def nll(probs, labels):
    picked = probs[np.arange(len(labels)), labels]
    return float(-np.mean(np.log(np.clip(picked, 1e-12, 1.0))))

def ece(probs, labels, bins=15):
    confidence = probs.max(axis=1)
    correct = (probs.argmax(axis=1) == labels).astype(float)
    edges = np.linspace(0.0, 1.0, bins + 1)
    total = 0.0
    for lo, hi in zip(edges[:-1], edges[1:]):
        mask = (confidence > lo) & (confidence <= hi)
        if mask.any():
            total += mask.mean() * abs(correct[mask].mean() - confidence[mask].mean())
    return float(total)

def fit_temperature(logits, labels, lo=0.05, hi=50.0):
    a, b = math.log(lo), math.log(hi)
    phi = (math.sqrt(5.0) - 1.0) / 2.0
    objective = lambda log_t: nll(softmax(logits / math.exp(log_t)), labels)
    c, d = b - phi * (b - a), a + phi * (b - a)
    fc, fd = objective(c), objective(d)
    for _ in range(80):
        if fc < fd:
            b, d, fd = d, c, fc
            c = b - phi * (b - a)
            fc = objective(c)
        else:
            a, c, fc = c, d, fd
            d = a + phi * (b - a)
            fd = objective(d)
    return math.exp((a + b) / 2.0)

And the assertion that anchors the article:

before = metrics(test_logits, y_test)
after = metrics(test_logits, y_test, t_clean)
assert before["accuracy"] == after["accuracy"], "temperature scaling must not change argmax"

Results, seed 0

ScenarioVariantTAccuracyNLLBrierECEMean confidence
In distributionbefore1.000.84850.84080.25630.11580.9643
In distributionT from clean validation4.380.84850.38200.21540.01980.8379
Covariate driftbefore1.000.64924.18320.65980.32000.9693
Covariate driftT from clean validation4.380.64921.18900.54510.21410.8632
Covariate driftT refitted on 500 recent13.970.64920.78310.46140.01810.6485
Concept driftbefore1.000.36057.70371.21720.60120.9617
Concept driftT from clean validation4.380.36051.95870.99140.47620.8367
Concept driftT refitted on 500 recent27.860.36051.03440.63130.14370.4747

Results, seed 1

ScenarioVariantTAccuracyNLLBrierECEMean confidence
In distributionbefore1.000.87550.38990.19610.05610.9304
In distributionT from clean validation1.930.87550.32220.18660.01300.8654
Covariate driftbefore1.000.67652.22230.57090.25820.9347
Covariate driftT from clean validation1.930.67651.26550.52010.20200.8784
Covariate driftT refitted on 500 recent6.800.67650.74760.44100.02490.6713
Concept driftbefore1.000.29789.36581.31520.63220.9296
Concept driftT from clean validation1.930.29785.01081.24680.56880.8665
Concept driftT refitted on 500 recent50.000.29781.15470.70300.08310.3809

What the tables say

Accuracy never moves. Every row within a scenario has the same accuracy. The earlier article's "improved accuracy" was not a measurement error; it was a category error.

In distribution, one temperature works well. With seed 0 the over-fitted model reports 96 percent mean confidence at 85 percent accuracy; T = 4.38 brings ECE from 0.116 to 0.020 and NLL from 0.84 to 0.38. Seed 1 over-fits less (T = 1.93) and gains less. How much calibration helps depends on how over-confident the model was to begin with.

Under covariate drift the old temperature is stale. Accuracy falls to about 65 percent while the model keeps reporting 97 percent confidence. The temperature fitted on clean validation data helps but leaves ECE at 0.21 (seed 0). Refitting on 500 recent labeled samples brings ECE to 0.018. This is the case where a continuous calibration job earns its keep: the model's ranking is still useful, its confidence is not, and recent labels fix the confidence.

Under concept drift calibration can only confess. Accuracy is 36 percent (seed 0) and 30 percent (seed 1), below the 33 percent a coin would get on three classes, because two classes swapped. No temperature changes that. What the refit does is push T toward the upper bound of the search (27.9, and 50.0 which is the bound itself for seed 1), flattening the probabilities toward uniform. Mean confidence drops from 0.96 to 0.47. That is the honest answer, and it is also a signal: when the best temperature runs away, recalibration is telling you the model needs retraining, not calibrating.

Reading the fitted temperature as a monitor

Because T is one number per refit, it is cheap to log. In this lab it moved from about 4 to 14 (seed 0) when noise increased and toward the bound when the concept changed. A rising temperature trend on recent data, alongside stable accuracy on the same batch, points to covariate drift; a runaway temperature with collapsing accuracy points to concept drift. This is an observation from two seeds of synthetic data, not a validated detector. Treat it as a hypothesis to test on your own model.

Failure modes of the continuous-calibration idea itself

  • Labels arrive late or not at all. Every refit above used 500 labeled recent samples. Without recent labels there is nothing to fit; unsupervised drift statistics on the inputs (for example a Kolmogorov-Smirnov test per feature, which the earlier article showed) can tell you inputs changed, not whether the probabilities are still honest.
  • Small recent batches are noisy. Refitting on 500 samples gave a well-calibrated result here, but the variance of T across batches was not measured. A production job should fit on a window and smooth or bound the temperature.
  • ECE depends on binning. Fifteen equal-width bins is a common choice, not the only one; report the binning with the number.
  • One temperature per model is a strong assumption. Temperature scaling cannot fix miscalibration that differs by class or by input region. Class-wise or vector scaling exist for that, and were not evaluated here.

What this lab is not

It is a 20-feature Gaussian toy with a linear model. It says nothing about neural networks, about how much real-world drift looks like added noise, or about any particular monitoring product. Its value is that each claim above can be re-run in a second and checked.

Reproduce it

python3 calibration_lab.py --seed 0 --json results-seed0.json
python3 calibration_lab.py --seed 1
python3 -m unittest -q test_calibration_lab.py   # argmax invariance, T recovery, ECE sanity, determinism

Sources