Quantizing Models for Wasm Inference

This guide answers one task: convert a float32 model into an int8 model that downloads four times faster, uses a quarter of the memory, and runs faster on the WebAssembly backend — while proving how much accuracy you gave up rather than hoping it was fine.

Prerequisites

  • [ ] onnxruntime and onnxruntime-tools installed in a Python environment.
  • [ ] The float32 model, plus an evaluation set with labels — a few hundred examples is enough.
  • [ ] A calibration set of 100–500 representative inputs if you use static quantization.
  • [ ] A way to run the quantized model in the browser and compare against the reference.

What quantization actually changes

A float32 weight takes four bytes and covers an enormous dynamic range. An int8 weight takes one byte and covers 256 values, so the conversion stores a scale factor — and sometimes a zero point — alongside each group of weights and reconstructs an approximation at runtime.

The saving is not only size. Integer kernels move a quarter of the bytes through the cache, and on the WebAssembly backend the i8x16 SIMD lanes process sixteen values per instruction where f32x4 processes four. That is why an int8 model on this backend is commonly 1.5–3× faster as well as 4× smaller — a rare case where the cheap option is also the fast one.

Four bytes become one, plus a scale Each group of float weights is divided by a scale and rounded to an 8-bit integer. The scale is stored once per tensor or once per output channel, and the kernel reconstructs approximate float values as it multiplies. float32 tensor −0.418 0.221 −0.097 0.630 … 4 bytes each divide by scale, round int8 tensor + scale −53 28 −12 80 … 1 byte each scale 0.00787 one scale per tensor is simple and lossy; one per output channel costs almost nothing and recovers most of the lost accuracy. The rounding error is bounded by half a step, so the damage depends entirely on how wide the range within each group is.

Dynamic quantization: start here

Dynamic quantization converts weights offline and computes activation scales at runtime. It needs no calibration data, it is one command, and for transformer-style models it is usually within a fraction of a percent of the float32 model.

python -m onnxruntime.quantization.preprocess --input model.onnx --output model-prep.onnx

python - <<'PY'
from onnxruntime.quantization import quantize_dynamic, QuantType
quantize_dynamic(
    "model-prep.onnx", "model-int8.onnx",
    weight_type=QuantType.QInt8,
    per_channel=True,          # almost always worth it
    reduce_range=False,        # set True only for old x86 targets
)
PY

The preprocessing step matters more than its obscurity suggests: it runs shape inference and folds constants, and skipping it leaves nodes that the quantizer cannot handle, producing a model that is “quantized” but still mostly float. Check the result rather than trusting the exit code — the file size tells you immediately whether it worked.

Static quantization when activations matter

Static quantization also quantizes activations, using ranges measured from calibration data. It is faster at runtime because nothing is computed per inference, and it is more fragile because the calibration set has to resemble real input.

from onnxruntime.quantization import quantize_static, CalibrationDataReader, QuantFormat

class Reader(CalibrationDataReader):
    def __init__(self, samples): self.it = iter([{"pixel_values": s} for s in samples])
    def get_next(self): return next(self.it, None)

quantize_static(
    "model-prep.onnx", "model-int8-static.onnx",
    Reader(calibration_samples),
    quant_format=QuantFormat.QDQ,
    per_channel=True,
)

Use between 100 and 500 calibration samples drawn from the same distribution as production input. Too few and the measured ranges are noisy; too many and you spend an hour for no further benefit. Include the awkward cases — very dark images, very short texts, silence — because a range measured only on typical input clips the tails and produces exactly the failures users report.

Measuring what you lost

Quantization without an evaluation is a guess. Run both models over the same held-out set and compare the metric you actually care about, not a proxy.

import numpy as np, onnxruntime as ort
ref = ort.InferenceSession("model.onnx")
q   = ort.InferenceSession("model-int8.onnx")

agree = 0
for x, y in eval_set:
    a = ref.run(None, {"pixel_values": x})[0].argmax()
    b = q.run(None,   {"pixel_values": x})[0].argmax()
    agree += (a == b)
print(f"top-1 agreement: {agree / len(eval_set):.3%}")

Agreement with the float model is a more useful number than absolute accuracy, because it isolates the damage quantization did from whatever the model’s baseline errors already were. Below about 99% agreement on a classifier, investigate rather than ship: it usually means one layer has an unusually wide weight distribution and should be excluded from quantization.

Excluding the layers that hurt

Not every node benefits. The first convolution and the final classifier layer are frequently sensitive, and excluding them costs a few percent of the size saving while recovering most of the accuracy.

quantize_dynamic(
    "model-prep.onnx", "model-int8.onnx",
    weight_type=QuantType.QInt8,
    per_channel=True,
    nodes_to_exclude=["/classifier/Gemm", "/features/0/Conv"],
)

Finding the culprits is a bisection: quantize half the nodes, evaluate, and narrow down. In practice two or three passes find them, and the node names come straight from the graph — Netron or onnx.load().graph.node will list them.

What each precision costs and saves Moving from float32 to int8 cuts the file to a quarter with a small accuracy cost. Moving to int4 halves it again but the accuracy cost grows sharply for most models, so it suits only the largest ones. float32 92 MB · baseline accuracy f32x4 lanes reference only int8, per-channel 23 MB · −0.4% top-1 i8x16 lanes, 1.5–3× faster ship this int4, grouped 12 MB · −2 to −6% unpacking costs time back large models only Figures are typical for a vision classifier; your model's sensitivity is the only number that matters, and it takes ten minutes to measure.

Per-tensor versus per-channel scales

The choice of how finely you store scales is the single largest quality lever in the whole process, and it costs almost nothing to get right. A per-tensor scale stores one number for an entire weight tensor, so every output channel shares the same quantisation step. A per-channel scale stores one number per output channel, which is a few hundred extra floats in a file that just lost tens of megabytes.

Why it matters so much: convolution and linear layers routinely have channels whose weights span very different magnitudes. One channel might range over ±0.03 while its neighbour ranges over ±1.4. A single shared scale has to cover the widest channel, which means the narrow channel’s weights all collapse into a handful of the 256 available integer values — sometimes into two or three. That channel’s contribution becomes nearly random, and because it is one channel among hundreds, the damage shows up as a small accuracy drop rather than an obvious failure.

Per-channel quantisation gives each channel its own step size, so the narrow one gets fine resolution and the wide one gets coarse resolution, exactly as each needs. The measured difference on typical vision models is between 0.5 and 3 percentage points of top-1 accuracy — the difference between shipping and not shipping. There is no meaningful runtime cost, because the kernel folds the scale into the accumulation it was already doing.

Set per_channel=True unless you have measured that it is unsupported by your target backend. A handful of older mobile execution paths only handle per-tensor, and the symptom is a model that loads and then produces uniformly wrong results.

Keeping the pipeline reproducible

Quantization is a build step, and treating it as one prevents the situation where nobody can reproduce the model that is currently in production. Check the float model, the calibration set and the quantizer invocation into version control, and produce the quantized artifact from CI rather than from someone’s laptop.

# quantize.sh — the only supported way to produce a shipping model
set -euo pipefail
python -m onnxruntime.quantization.preprocess --input src/model.onnx --output build/model-prep.onnx
python tools/quantize.py --in build/model-prep.onnx --out build/model-int8.onnx --per-channel
python tools/evaluate.py --ref src/model.onnx --quant build/model-int8.onnx --min-agreement 0.99
sha256sum build/model-int8.onnx | tee build/model-int8.sha256

The evaluation step with a threshold is what turns this from a script into a gate: if a model change drops agreement below the bar, the build fails rather than shipping a quietly worse model. The hash goes into the filename you deploy, so the browser’s cache key changes exactly when the weights do — the same discipline as versioning any other artifact.

When quantization makes things slower

Three situations reliably produce a quantized model that is bigger, slower or both, and all three surprise people who expected a free win.

The first is a graph full of quantize and dequantize pairs. If the quantizer could not fuse them, every operator converts back and forth, and the conversion costs more than the integer arithmetic saves. Look at the node count before and after — a large increase is the symptom.

The second is int4 on a backend without native int4 kernels. The weights unpack to int8 or float before every multiply, so you pay unpacking on every inference to save download size once. That trade is correct for a model where the download dominates and wrong for one that runs continuously.

The third is a model dominated by operators that have no integer implementation at all — some attention variants, unusual normalisations. The quantized weights are stored small and converted to float on load, giving the download saving with none of the speed.

What each precision costs Quantization trades a little accuracy for a lot of size and speed. The step from 32-bit to 8-bit is usually worth it; the step below that rarely is. float32 102 MB · 140 ms · baseline accuracy float16 51 MB · 96 ms accuracy within noise int8 26 MB · 52 ms 0.4 points below baseline on the evaluation set Always evaluate on your own data: published accuracy deltas are for the benchmark set, not your task. Not every operator has a quantized kernel — an unsupported one falls back and can end up slower.

Gotchas

  • The file barely shrank. The preprocessing step was skipped, so most nodes were not quantizable. Run quantize_dynamic on the preprocessed model, not the raw export.
  • Quantization parameters are not specified at runtime. A QDQ model was produced without calibration for some tensor. Either supply calibration data or use dynamic quantization.
  • Accuracy collapses on a specific input class. Calibration data did not cover it. Widen the calibration set rather than tuning the quantizer.
  • The browser refuses to load the model. Some quantized formats are not supported by every execution provider — notably on the GPU path. Test the quantized model on every backend in your fallback chain.
  • Results differ between the Python check and the browser. Expected, within tolerance: the browser’s kernels accumulate differently. Compare argmax agreement rather than raw values.

Performance note

For a 92 MB vision model, per-channel dynamic int8 produced a 23 MB file, cut browser session creation from 620 ms to 190 ms, and reduced p50 inference on the threaded WebAssembly backend from 71 ms to 34 ms, with 99.6% top-1 agreement against the float model. The download saving is what users feel first; the inference saving is what makes the feature usable on a phone at all.

Frequently Asked Questions

Should I quantize before or after other graph optimisations? After. Run shape inference and constant folding first — that is what the preprocess step does — so the quantizer sees the simplest possible graph and can fuse more.

Is float16 a middle ground? On the WebAssembly backend, not usefully: there are no native float16 kernels, so values are widened to float32 on the fly. It halves the download, which is real, but you get none of the speed benefit that int8 brings.

Can I quantize a model I did not train? Yes, that is the normal case. You need an evaluation set representative of your use, not the original training data, and the agreement metric above tells you whether the result is acceptable.

← Back to Machine Learning Inference in the Browser