Understanding Relaxed SIMD

This page answers one question: WebAssembly’s standard SIMD is fully deterministic — what does the relaxed SIMD proposal add, why are its results allowed to differ between machines, and when should you use it?

Prerequisites

  • [ ] Familiarity with Wasm SIMD, as in writing SIMD from C with wasm_simd128.h.
  • [ ] A workload that is numeric and tolerant of small differences — machine learning inference, graphics, audio effects.
  • [ ] A recent clang (-mrelaxed-simd) or Rust nightly, and an engine with relaxed SIMD support.

Determinism has a price

WebAssembly specifies the result of every instruction exactly, down to the bit, on every machine. That determinism is a deliberate feature: the same module produces the same output on an ARM phone and an x86 server, which matters for simulations, lockstep multiplayer games, and anything cryptographic. The standard SIMD proposal follows the same rule.

The cost is that some fast native instructions cannot be used directly, because they behave slightly differently on different CPUs. Fused multiply-add computes a * b + c with one rounding instead of two — faster and more accurate, but a different result from separate multiply and add, and not available on every CPU. Native float minimum and maximum handle NaN and negative zero differently on x86 and ARM. Native byte shuffles treat out-of-range indices differently. To keep results identical everywhere, standard SIMD emulates the strict semantics with extra instructions on the platforms whose native behaviour differs.

Relaxed SIMD adds a small set of instructions whose results are allowed to vary within a defined set of outcomes, so engines can map each to the fastest native instruction. Results are deterministic on a given machine — the same input gives the same output every time — but may differ between machines.

Standard SIMD versus relaxed SIMD semantics Standard SIMD instructions produce bit-identical results on every machine, sometimes by emulating strict semantics with extra instructions. Relaxed SIMD instructions may produce slightly different results between machines but map to the fastest native instruction. standard SIMD identical bits on every CPU strict NaN, zero and rounding rules emulation where CPUs differ reproducible everywhere relaxed SIMD results may vary between CPUs fixed per machine, not random one native instruction where possible faster, tolerant code only

Step 1 — know the instructions

The proposal is small. Its main groups:

f32x4.relaxed_madd / nmadd     a*b+c, fused or unfused         (FMA where available)
f64x2.relaxed_madd / nmadd
f32x4.relaxed_min / max        NaN and ±0 handling may differ   (native minps/maxps)
f64x2.relaxed_min / max
i8x16.relaxed_swizzle          out-of-range indices may differ  (native pshufb / tbl)
i8x16/i16x8/i32x4/i64x2.relaxed_laneselect   bit-select or lane-select semantics
i32x4.relaxed_trunc_f32x4_s/u  out-of-range float→int may differ
i16x8.relaxed_dot_i8x16_i7x16_s           8-bit dot products for quantized ML
i32x4.relaxed_dot_i8x16_i7x16_add_s
f32x4.relaxed_dot_bf16x8_add_f32          bfloat16 dot product

The pattern is the same in each: on inputs where all platforms agree — finite floats, in-range indices, in-range conversions — results are identical. Only edge cases may differ, and they differ in documented ways.

Step 2 — see where it pays off

The biggest wins are fused multiply-add in floating-point kernels and the 8-bit dot products in quantized neural network inference. A matrix multiply inner loop with standard SIMD needs a multiply and an add per step; with relaxed_madd on hardware with FMA, it needs one instruction. The dot-product instructions map to dedicated integer dot-product hardware on modern CPUs (VNNI on x86, SDOT on ARM), which is why machine learning runtimes were the main motivation for the proposal — see quantizing models for Wasm inference.

#include <wasm_simd128.h>
// accumulate a 4-wide dot product with relaxed FMA
v128_t acc = wasm_f32x4_splat(0.0f);
for (size_t i = 0; i + 4 <= n; i += 4) {
  acc = wasm_f32x4_relaxed_madd(wasm_v128_load(a + i), wasm_v128_load(b + i), acc);
}
clang --target=wasm32 -O3 -msimd128 -mrelaxed-simd dot.c -c -o dot.o

Step 3 — decide whether your code can tolerate it

Use relaxed SIMD only where small, machine-dependent differences are acceptable. Neural network inference is the canonical fit: models are trained with noise and quantization, and a last-bit difference in an activation changes nothing meaningful. Graphics and audio effects usually fit. Physics for a single-player visual effect fits.

Do not use it where results must match exactly across machines: deterministic simulations, replays, lockstep networking, financial calculations, test suites that compare against golden files bit for bit, and anything that hashes or signs numeric output. And avoid it where inputs routinely include the edge cases — NaN-heavy data in min/max, out-of-range indices in swizzles — because those are exactly where results differ.

Should this kernel use relaxed SIMD? A decision tree. If results must match bit for bit across machines, use standard SIMD. If small cross-machine differences are acceptable and the kernel is dominated by multiply-add or 8-bit dot products, relaxed SIMD pays off; otherwise the gain is small. Must results be identical on every machine? yes (replays, lockstep, golden tests) Standard SIMD only determinism is the requirement no, FMA or int8 dot heavy Use relaxed SIMD largest speedups no, other operations Measure first gains are often small

Step 4 — detect support and ship a fallback

Relaxed SIMD is newer than standard SIMD, so detect it before loading a module that uses it. Validate a tiny module containing one relaxed instruction:

const RELAXED_PROBE = new Uint8Array([
  0x00, 0x61, 0x73, 0x6d, 0x01, 0x00, 0x00, 0x00, 0x01, 0x05, 0x01, 0x60, 0x00, 0x01, 0x7b,
  0x03, 0x02, 0x01, 0x00, 0x0a, 0x0f, 0x01, 0x0d, 0x00, 0x41, 0x00, 0xfd, 0x0f, 0x41, 0x00,
  0xfd, 0x0f, 0xfd, 0x80, 0x02, 0x0b,                       // i8x16.relaxed_swizzle
]);
const hasRelaxedSimd = WebAssembly.validate(RELAXED_PROBE);

Ship three builds if you need the widest reach — relaxed, standard SIMD and scalar — or two, since standard SIMD is universal in current browsers. Machine-learning libraries such as ONNX Runtime Web and TensorFlow.js already do this selection internally. The general approach is in shipping SIMD and baseline builds together.

Step 5 — test on more than one architecture

Because relaxed results may differ between machines, test on both x86-64 and ARM64 — a laptop and a phone, or CI runners of both kinds — and compare outputs with a tolerance rather than exact equality. A test that passes on your x86 development machine and fails on an ARM phone is the characteristic relaxed-SIMD bug, and the cure is a tolerance chosen from measured differences, not disabling the test.

Choosing tolerances for tests

The one practical change relaxed SIMD forces on a project is in its tests. Bit-exact comparisons that worked with standard SIMD will fail on some machines, so tests of relaxed code paths need tolerances, and those tolerances should be chosen deliberately rather than loosened until tests pass. For fused multiply-add, the difference between fused and unfused results is bounded by roughly one rounding error per operation; a long dot product accumulates those, so a relative tolerance of around 1e-5 for f32 accumulations over a few thousand terms is a reasonable start. For inference, compare final outputs — the predicted class, the top-k tokens — exactly, and intermediate activations with a tolerance, since the application cares about the former.

Measure the actual differences on the machines you have — an x86 laptop, an ARM phone or CI runner — and set the tolerance a small factor above the largest observed difference. Record those measurements next to the test, so the next person knows the tolerance came from data. And keep at least one test of the standard SIMD path with exact comparisons, so a change that accidentally introduces relaxed behaviour into code that must be deterministic is caught.

Why the proposal is designed this way

Relaxed SIMD is a careful compromise rather than a loosening of WebAssembly’s principles. Each instruction’s possible results are enumerated in the specification — “either the fused result or the unfused result”, “either of these two NaN behaviours” — so the nondeterminism is bounded and documented, not undefined behaviour. Results are fixed per machine, so a given device is internally consistent and reproducible in its own tests. And the instructions are opt-in, so modules that need strict determinism are unaffected. The design lets the code that most needs raw speed — inference kernels — use the hardware fully, while everything else keeps WebAssembly’s portability guarantees.

Expected output

On a machine with FMA, the relaxed dot-product kernel produces results within a few ULPs of the standard version and runs faster; on a machine without FMA, it produces exactly the standard result at the standard speed.

Gotchas

  • Golden-file tests fail on another CPU. Relaxed results vary by machine. Compare with tolerances in tests for relaxed code paths.
  • Assuming relaxed means random. It does not. Results are fixed per machine; nondeterminism is between machines.
  • Using relaxed min/max on NaN-heavy data. NaN handling is exactly what varies. Filter NaNs or use the standard instructions.
  • No fallback. Older engines reject the module. Detect and load a standard SIMD build where needed.

Performance note

For a quantized 8-bit matrix multiply in Chrome on a laptop with VNNI support, the relaxed dot-product path was 2.4× faster than standard SIMD emulating the same arithmetic with widening multiplies and adds. An f32 matrix multiply using relaxed_madd gained about 1.3× on the same machine. On an ARM phone with dot-product support the int8 gain was similar.

Relaxed SIMD speedup over standard SIMD for two kernels Speedup of relaxed SIMD over standard SIMD for an int8 quantized matrix multiply and an f32 matrix multiply, in Chrome on a laptop with VNNI and FMA support. speedup over standard SIMD (×) standard SIMD baseline 1 × f32 matmul with relaxed_madd 1.3 × int8 matmul with relaxed dot 2.4 ×

Frequently Asked Questions

Is relaxed SIMD supported everywhere? It is supported in current Chromium-based browsers and Firefox and progressing elsewhere; detect it rather than assuming.

Does relaxed SIMD break Wasm’s security guarantees? No. Results vary only in values, never in memory safety or control flow.

Can I use it from Rust? Yes — core::arch::wasm32 exposes relaxed intrinsics behind the relaxed-simd target feature, currently on nightly or recent stable releases.

Do autovectorizers emit relaxed instructions? Generally not by default, because they would change results. Use intrinsics or a library that opts in explicitly.

← Back to Wasm SIMD & Vectorized Computation