Compressing a Model for the Edge: What int8 and Pruning Actually Cost, Measured

Lab run with Python 3.14.7, NumPy 2.4.4, onnx 1.22.0 and onnxruntime 1.30.0 for seeds 0 and 1, 26 unit tests; TensorFlow Model Optimization Toolkit, LiteRT, TensorRT, OpenVINO and PyTorch claims checked against documentation and source only, not run

Revision note (2026-09-15). The earlier version of this article was a TensorFlow script that was never run, and it could not have run as written. It called pruned_model.fit(...) on a prune_low_magnitude model without the tfmot.sparsity.keras.UpdatePruningStep callback, which the toolkit's pruning wrapper asserts on at the first training step ("Prune() wrapper requires the UpdatePruningStep callback to be provided during training"). Had it run, its "fine-tuning" was two steps over 32 random images against a PolynomialDecay schedule spanning 1,000 steps, which reaches at most 0.3 percent sparsity, not the 50 percent the comment promised. It said converter.optimizations = [tf.lite.Optimize.DEFAULT] "converts weights and activations to 8-bit integers"; without a representative_dataset that flag performs dynamic-range quantization, in which only weights are stored as int8 and outputs stay in float, and its own validation snippet fed float32 input, which only works because the model was not full-integer. It uploaded the pruned .tflite as the deliverable without noting that a pruned model is the same size on disk until it is compressed, which the toolkit's guide calls a "common mistake". It pinned TensorFlow 2.8 (February 2022) and Python 3.7, used actions/upload-artifact@v3, which GitHub switched off on 2025-01-30, and asserted an accuracy loss "within ~2% of baseline for INT8" with nothing behind it. TensorFlow Lite has been called LiteRT since September 2024. This version replaces the untested pipeline with a lab in which every number below is produced by code you can run; the TensorFlow, LiteRT, TensorRT, OpenVINO and PyTorch statements were checked against documentation and source and are labelled as such.

What "compression" means here

Two techniques are measured: post-training int8 quantization of weights (and, separately, activations) and unstructured magnitude pruning without fine-tuning. Quantization-aware training, distillation, structured sparsity and low-rank factorization are not measured and are only mentioned where the measured results point at them. The model is a small dense network, not a convolutional or transformer model, and the numbers are about the mechanisms, not about any real edge workload.

Tested versions: Python 3.14.7, NumPy 2.4.4, onnx 1.22.0, onnxruntime 1.30.0 on an Apple M3 Pro laptop (macOS 26.5.1). The lab is examples/ai-model-compression in the companion repository; python3 -m unittest -v test_compression_lab.py runs 26 tests (5 need onnxruntime) and python3 compression_lab.py --seed 0 --json results-seed0.json reproduces the tables. Both onnx and onnxruntime installed into a fresh Python 3.14 virtual environment with a plain pip install; PyTorch and TensorFlow are not installed and nothing depends on them.

The lab

The model is a 32-512-512-10 ReLU MLP with 283,648 weights, trained in NumPy (Adam, 10 epochs) on 50,000 Gaussian inputs labelled by a fixed random teacher network. Held-out accuracy on 5,000 test samples is 0.8230 (seed 0) and 0.8290 (seed 1); the majority class is 0.22 and 0.19. The task is deliberately hard enough that the model has something to lose, and it overfits (train accuracy 0.96), which is noted because it is the kind of model a pipeline actually receives. A separate 512-sample slice is the calibration set. The float32 weights and biases are 1,138,728 bytes.

Quantization is symmetric int8 in the range -127 to 127 with scale = max|w| / 127, computed once per matrix (per-tensor) or once per output column (per-channel). Weight-only quantization dequantizes at load time and runs float kernels, which is what the earlier article's converter flag does to the weights. The integer path also quantizes each layer's input activations with a per-tensor scale taken from the calibration set, multiplies int8 by int8 with int32 accumulation and rescales; that is the shape of a full-integer model.

def quantize_symmetric(w, per_channel):
    amax = np.max(np.abs(w), axis=0, keepdims=True) if per_channel else np.max(np.abs(w), keepdims=True)
    scale = np.where(amax == 0, 1.0, amax / 127).astype(np.float32)
    q = np.clip(np.rint(w / scale), -127, 127).astype(np.int8)
    return q, scale

int8 weights: a quarter of the bytes for a fifth of a point

SeedVariantBytesAccuracyArgmax agreement with float
0float321,138,7280.82301.0000
0int8 weights per-tensor287,7960.82480.9932
0int8 weights per-channel291,9200.82320.9950
0int8 weights per-tensor + int8 activations (MinMax)287,8080.82420.9870
0int8 weights per-channel + int8 activations (MinMax)291,9320.82120.9850
1float321,138,7280.82901.0000
1int8 weights per-tensor287,7960.83000.9948
1int8 weights per-channel291,9200.82880.9964
1int8 weights per-tensor + int8 activations (MinMax)287,8080.82720.9878
1int8 weights per-channel + int8 activations (MinMax)291,9320.82800.9860

The bytes are 25.3 to 25.6 percent of float32; the remainder above one quarter is the float32 biases and the scales, and per-channel costs 4,124 bytes more than per-tensor for the 1,031 extra scales. On this model the accuracy change is within the noise of a 5,000-sample test set in every row, and the accuracy column hides what the agreement column shows: even the best variant changes the predicted class of 0.4 to 0.5 percent of inputs, and quantizing activations as well changes 1.2 to 1.5 percent. A test that only checks aggregate accuracy would pass a model that answers differently on one input in eighty. Report the disagreement rate with the float model, not just the accuracy.

Calibrating activations on the 99.99th percentile instead of the maximum gave the same accuracy to within a point here, and a larger maximum logit error (5.1 versus 2.6), because clipping the rare large activation costs more than the finer resolution buys on this model. That is one data point, not a rule.

The failure case: the same float model, 12 percent accuracy in int8

For ReLU networks, multiplying one hidden unit's incoming weights and bias by a factor and dividing its outgoing weights by the same factor changes nothing: relu(64 z) = 64 relu(z). The lab does exactly that to hidden unit 7 of layer 1, and a test asserts the float outputs are unchanged to 1e-3 (the recorded maximum difference is 0.0). Accuracy stays at 0.8230. The int8 story is different, because layer 1 now has one column whose magnitude is 16 to 1,024 times the others.

FactorLayer-1 max weightWeights per-tensor, weight-onlyWeights per-channel, weight-onlyPer-tensor + int8 activationsPer-channel + int8 activations
1 (original)1.160.82480.82320.82420.8212
1610.50.81920.82300.80780.8182
6442.20.79180.82320.71920.7442
256168.70.29020.81800.11640.2102
1,024674.60.22280.81660.11400.1140

Seed 0; seed 1 is in the lab README and shows the same shape (0.4120 and 0.1532 for per-tensor at 256 and 1,024). Two separate things happen in that table.

The first is the textbook one. With a per-tensor scale, the outlier column sets the step size for the whole matrix: at factor 256, 98.7 percent of layer 1's weights round to zero and a normal column has exactly one distinct value. Per-channel scales fix that completely, because each column gets its own step; the per-channel weight-only column loses at most 0.7 points at any factor. This is what the onnxruntime documentation means by "Per-channel quantization can improve the accuracy for models whose weight ranges are large", and it is why per-channel weights are the default you want and why a pipeline should print each layer's max-to-median weight ratio before quantizing.

The second is the one that per-channel weights do not fix. The outlier unit's activation is also 256 times larger than its neighbours, and the input to layer 2 is quantized with one per-tensor scale. That scale goes from 0.063 on the original model to 11.2, a step 177 times coarser for every other unit, and accuracy collapses to 0.21 even with per-channel weights. In the real onnxruntime run below, at factor 64, quantize_static with per_channel=True lands at 0.7260 against 0.8230 in float for the same reason. Weight outliers are a weight-quantization problem; activation outliers are an activation-quantization problem, and the standard remedies (per-channel or per-token activation scales where the runtime supports them, scaling the outlier out of the activations and into the weights, or keeping that layer in float) are different from the remedy for weights. A pipeline that only checks weight ranges will pass this model and ship a 12-percent classifier.

Pruning without fine-tuning: the bytes do not move

Magnitude pruning zeroes the smallest weights, either per matrix (layer-wise) or by one threshold across all matrices (global). No fine-tuning follows, which is what the earlier version's two-step fit amounted to.

SparsityLayer-wise accuracyGlobal accuracyDense bytesgzip of dense bytesCSR estimate
00.82300.82301,138,7281,056,9961,710,260
0.500.77700.80961,138,728610,398859,316
0.700.70380.75781,138,728394,117518,942
0.800.61180.69321,138,728273,788348,752
0.900.42140.53921,138,728146,240178,562
0.950.27940.35881,138,72877,57693,470
0.980.22760.16521,138,72833,31942,410

Seed 0; gzip and CSR columns are for the global variant. Three things to read off.

The dense byte count is identical in every row. A zero is a four-byte float like any other, and a .tflite, .onnx or .npy file that stores the matrix densely is the same size at 98 percent sparsity as at zero. The TensorFlow Model Optimization guide says it directly: "both strip_pruning and applying a standard compression algorithm (e.g. via gzip) are necessary to see the compression benefits of pruning". The earlier version's pipeline uploaded the dense file and measured nothing, so it would have reported a compression step that compressed nothing. Compressing the dense bytes with gzip brings 90 percent sparsity to 146,240 bytes; a sparse storage format brings it to 178,562 bytes, and note that CSR is larger than dense below about 33 percent sparsity, so the format has to be chosen after measuring, not before.

The latency does not move either. The pruned-90-percent row in the latency table below runs at the same 11.5 us as the dense model, because a dense kernel multiplies zeros at full price. A speedup needs a kernel that skips them, and such kernels have conditions: LiteRT's XNNPACK sparse inference, per its README, applies only to subgraphs beginning with a 3×3 stride-2 CONV_2D with 3 input channels, whose 1×1 CONV_2D weights are at least two-thirds zero and are stored with DENSIFY operators, and the toolkit exposes this through the PruneForLatencyOnXNNPack policy. None of that applies to a dense MLP. Pruning for size and pruning for speed are different projects with different constraints, and the earlier version treated them as one step.

Accuracy falls immediately and falls to chance. Fifty percent costs 1.3 points with a global threshold and 4.6 layer-wise; 90 percent costs 28 to 40 points; 98 percent is at the majority-class rate (0.2276 and 0.1652 against 0.2228). Global thresholds beat layer-wise at every level up to 95 percent because the last layer (5,120 weights) matters more per weight than the middle one (262,144), and a per-layer quota cannot express that. Every published pruning result that keeps accuracy at high sparsity does so with fine-tuning between pruning steps, which is what the PolynomialDecay schedule is for and what two steps over random data cannot do. Pruning to 50 percent and then quantizing per-channel kept the zeros (a test asserts it) and gave 0.7778 at 196,550 gzipped bytes, which is 17 percent of the float file for a 4.5-point loss; whether that trade is worth it is your model's decision.

Speed: int8 is faster only where an int8 kernel exists

PathBatch 1 p50Batch 256 p50
NumPy float3211.4 us224 us
NumPy int8 weights dequantized (float kernels)11.4 us
NumPy int8 x int8, int32 accumulate190.8 us41,800 us
NumPy pruned 90 percent, dense kernels11.5 us
onnxruntime float32, 1 thread19.2 us1,331 us
onnxruntime quantize_dynamic10.8 us449 us
onnxruntime quantize_static QDQ per-channel11.7 us454 us

Seed 0 on the laptop; seed 1 is within 15 percent on every row. The NumPy rows use the multi-threaded BLAS NumPy links against; the onnxruntime rows are pinned to one intra-op thread, so compare within a runtime, never across. Two lessons. NumPy has no int8 matrix-multiply kernel, so the integer path is 17 times slower than float at batch 1 and 186 times slower at batch 256: the same bytes that make a model small make it slow on a runtime that cannot multiply them natively. onnxruntime's CPU provider has such kernels, and there the int8 graphs run 1.8 times faster at batch 1 and 2.9 times faster at batch 256 than its own float graph. Whether your edge runtime is closer to the first row or the second is a fact about the runtime and the chip, not about the model, and it is the fact the earlier version told you to "benchmark on actual edge hardware" without ever producing one number itself.

The real tool path: onnxruntime static quantization

The lab exports the MLP to ONNX with onnx.helper (three MatMul, three Add, two Relu; onnxruntime's float output matches NumPy to 1e-4, asserted), runs the recommended quant_pre_process, then quantizes:

from onnxruntime.quantization import CalibrationMethod, QuantFormat, QuantType, quantize_static
from onnxruntime.quantization.shape_inference import quant_pre_process

class CalibReader:                      # onnxruntime CalibrationDataReader
    def __init__(self, x, batch=64):
        self.batches = [x[i:i + batch] for i in range(0, len(x), batch)]
        self.pos = 0
    def get_next(self):
        if self.pos >= len(self.batches):
            return None
        self.pos += 1
        return {"input": self.batches[self.pos - 1]}

quant_pre_process("mlp_float.onnx", "mlp_pre.onnx", skip_symbolic_shape=True)
quantize_static("mlp_pre.onnx", "mlp_int8.onnx", CalibReader(x_calib),
                quant_format=QuantFormat.QDQ, per_channel=True,
                activation_type=QuantType.QInt8, weight_type=QuantType.QInt8,
                calibrate_method=CalibrationMethod.MinMax)

Results, seed 0: the float file is 1,139,311 bytes; quantize_dynamic gives 289,897 bytes at 0.8224 with 3 MatMulInteger and 3 DynamicQuantizeLinear nodes; quantize_static in QDQ format gives 288,563 bytes per-tensor at 0.8226 and 293,781 bytes per-channel at 0.8198, with 7 QuantizeLinear and 13 DequantizeLinear nodes around the original MatMuls. On the factor-64 outlier model the same quantize_static call gives 0.7106 per-tensor and 0.7260 per-channel against 0.8230 in float, the activation-outlier effect from the table above reproduced by a shipping tool. A test asserts that at factor 256 static quantization loses more than ten points whichever way per_channel is set.

One thing the documentation does not tell you: quant_pre_process raises ImportError: sympy is required for symbolic shape inference on a fresh install, because sympy is not a dependency of the onnxruntime wheel. Pass skip_symbolic_shape=True or install sympy; a test pins the error.

What the earlier version's TensorFlow pipeline needed

None of this was run here; each item is a check against the toolkit's source or the current documentation.

  • The prune_low_magnitude wrapper requires tfmot.sparsity.keras.UpdatePruningStep() in the callbacks of fit; without it, PruneLowMagnitude.call fails a tf.debugging.assert_greater_equal on the pruning step with the message quoted in the revision note.
  • PolynomialDecay(initial_sparsity, final_sparsity, begin_step, end_step, power=3, frequency=100) computes sparsity = (initial - final) * (1 - p)^power + final with p = (step - begin_step) / (end_step - begin_step), and recomputes the mask only when (step - begin_step) % frequency == 0. Thirty-two samples at the default batch size is one step per epoch; two epochs is two steps of a 1,000-step schedule; the target was never reachable.
  • converter.optimizations = [tf.lite.Optimize.DEFAULT] alone is dynamic-range quantization: per the LiteRT post-training quantization page it "statically quantizes only the weights from floating point to integer at conversion time" and "the outputs are still stored using floating point". Full-integer quantization needs converter.representative_dataset (the page suggests around 100 to 500 samples), converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8], and inference_input_type / inference_output_type set to tf.int8 or tf.uint8, after which the validation snippet's float32 input would have to change too.
  • After strip_pruning, compress before you measure, or measure the sparse-format size the target runtime will actually load; the dense .tflite is not smaller.
  • The toolkit's last release is 0.8.1 (2024-05-12), its guides import tf_keras, and its 0.8.0 notes say it forces Keras 2. TensorFlow 2.16 made Keras 3 the default, so a current stack needs pip install tf-keras and TF_USE_LEGACY_KERAS=1 before tensorflow_model_optimization works with tf.keras models. The article's TensorFlow 2.8 pin predates that problem and is four years old; the current release is 2.21.0.
  • TensorFlow Lite was renamed LiteRT on 2024-09-04; the Python package is ai-edge-litert, the .tflite format is unchanged, and tensorflow.org/lite URLs redirect to developers.google.com/edge/litert.
  • actions/upload-artifact@v3 stopped working on 2025-01-30; v4 is the supported line.

For the other tools the earlier version listed as alternatives, checked against documentation only: TensorRT's current guide states that implicit quantization (calibration-driven INT8) is deprecated in favour of explicit Q/DQ quantization and supports per-tensor, per-channel and block scales for INT8, FP8 and FP4; OpenVINO's compression path is NNCF with a post-training quantization flow and a separate "quantizing with accuracy control" flow; PyTorch's quantization page for 2.14 says development is centralized in torchao, directing eager-mode users to torchao's quantize_ and FX users to its PT2E flow. If you are writing a pipeline around torch.ao.quantization today, read that page first.

What to gate on, given the above

The earlier version's checklist was "accuracy within ~2%", "no exceptions", "benchmark on hardware". The lab suggests a more specific set, each of which it demonstrates a failure of:

  1. Disagreement rate with the float model on held-out data, not just aggregate accuracy: the two can move independently, and the disagreement rate is what a downstream consumer experiences.
  2. Per-layer weight range report (max over median) and per-layer activation range from the calibration set, before quantizing; a ratio in the hundreds on either predicts the collapse above, and only the weight ratio is fixed by per-channel scales.
  3. Size measured as the bytes the device will load, after whatever compression or sparse format the runtime supports, not len(model_bytes) of a dense file.
  4. Latency measured on the target runtime with its own kernels; a number from a different runtime, or from a dense kernel over a sparse matrix, is not a proxy.
  5. A hard failure on any of the above outside a threshold you set from the float model's own variance, so that a model like the factor-256 one cannot be uploaded as an artifact.

What this lab is not

It is a dense MLP on synthetic data on one laptop. It measures no convolutional or transformer model, no real dataset, no edge device or NPU, no TensorFlow or LiteRT conversion, no TensorRT or OpenVINO run, no quantization-aware training, no distillation, and no sparse kernel (the CSR column is a byte count, not a runtime). Its value is that four effects, a byte count that pruning does not change, a latency that pruning does not change without a matching kernel, a function-preserving rewrite that destroys int8 accuracy, and a runtime without int8 kernels making the "faster" model 17 times slower, are each produced by a script that runs in about 25 seconds and pinned by tests.

Reproduce it

python3 -m venv .venv && .venv/bin/pip install -r requirements.txt   # numpy, onnx, onnxruntime, pinned
.venv/bin/python -m unittest -v test_compression_lab.py               # 26 tests (5 skipped without onnxruntime)
.venv/bin/python compression_lab.py --seed 0 --json results-seed0.json
.venv/bin/python compression_lab.py --seed 1 --json results-seed1.json
python3 compression_lab.py --seed 0 --skip-onnx                       # NumPy only

Sources