Understanding Relaxed SIMD
This page answers one question: WebAssembly’s standard SIMD is fully deterministic — what does the relaxed SIMD proposal add, why are its results allowed to differ between machines, and when should you use it?
Prerequisites
- [ ] Familiarity with Wasm SIMD, as in writing SIMD from C with wasm_simd128.h.
- [ ] A workload that is numeric and tolerant of small differences — machine learning inference, graphics, audio effects.
- [ ] A recent clang (
-mrelaxed-simd) or Rust nightly, and an engine with relaxed SIMD support.
Determinism has a price
WebAssembly specifies the result of every instruction exactly, down to the bit, on every machine. That determinism is a deliberate feature: the same module produces the same output on an ARM phone and an x86 server, which matters for simulations, lockstep multiplayer games, and anything cryptographic. The standard SIMD proposal follows the same rule.
The cost is that some fast native instructions cannot be used directly, because they behave slightly differently on different CPUs. Fused
multiply-add computes a * b + c with one rounding instead of two — faster and more accurate, but a different result from separate multiply and
add, and not available on every CPU. Native float minimum and maximum handle NaN and negative zero differently on x86 and ARM. Native byte
shuffles treat out-of-range indices differently. To keep results identical everywhere, standard SIMD emulates the strict semantics with extra
instructions on the platforms whose native behaviour differs.
Relaxed SIMD adds a small set of instructions whose results are allowed to vary within a defined set of outcomes, so engines can map each to the fastest native instruction. Results are deterministic on a given machine — the same input gives the same output every time — but may differ between machines.
Step 1 — know the instructions
The proposal is small. Its main groups:
f32x4.relaxed_madd / nmadd a*b+c, fused or unfused (FMA where available)
f64x2.relaxed_madd / nmadd
f32x4.relaxed_min / max NaN and ±0 handling may differ (native minps/maxps)
f64x2.relaxed_min / max
i8x16.relaxed_swizzle out-of-range indices may differ (native pshufb / tbl)
i8x16/i16x8/i32x4/i64x2.relaxed_laneselect bit-select or lane-select semantics
i32x4.relaxed_trunc_f32x4_s/u out-of-range float→int may differ
i16x8.relaxed_dot_i8x16_i7x16_s 8-bit dot products for quantized ML
i32x4.relaxed_dot_i8x16_i7x16_add_s
f32x4.relaxed_dot_bf16x8_add_f32 bfloat16 dot product
The pattern is the same in each: on inputs where all platforms agree — finite floats, in-range indices, in-range conversions — results are identical. Only edge cases may differ, and they differ in documented ways.
Step 2 — see where it pays off
The biggest wins are fused multiply-add in floating-point kernels and the 8-bit dot products in quantized neural network inference. A matrix
multiply inner loop with standard SIMD needs a multiply and an add per step; with relaxed_madd on hardware with FMA, it needs one instruction.
The dot-product instructions map to dedicated integer dot-product hardware on modern CPUs (VNNI on x86, SDOT on ARM), which is why machine
learning runtimes were the main motivation for the proposal — see
quantizing models for Wasm inference.
#include <wasm_simd128.h>
// accumulate a 4-wide dot product with relaxed FMA
v128_t acc = wasm_f32x4_splat(0.0f);
for (size_t i = 0; i + 4 <= n; i += 4) {
acc = wasm_f32x4_relaxed_madd(wasm_v128_load(a + i), wasm_v128_load(b + i), acc);
}
clang --target=wasm32 -O3 -msimd128 -mrelaxed-simd dot.c -c -o dot.o
Step 3 — decide whether your code can tolerate it
Use relaxed SIMD only where small, machine-dependent differences are acceptable. Neural network inference is the canonical fit: models are trained with noise and quantization, and a last-bit difference in an activation changes nothing meaningful. Graphics and audio effects usually fit. Physics for a single-player visual effect fits.
Do not use it where results must match exactly across machines: deterministic simulations, replays, lockstep networking, financial calculations, test suites that compare against golden files bit for bit, and anything that hashes or signs numeric output. And avoid it where inputs routinely include the edge cases — NaN-heavy data in min/max, out-of-range indices in swizzles — because those are exactly where results differ.
Step 4 — detect support and ship a fallback
Relaxed SIMD is newer than standard SIMD, so detect it before loading a module that uses it. Validate a tiny module containing one relaxed instruction:
const RELAXED_PROBE = new Uint8Array([
0x00, 0x61, 0x73, 0x6d, 0x01, 0x00, 0x00, 0x00, 0x01, 0x05, 0x01, 0x60, 0x00, 0x01, 0x7b,
0x03, 0x02, 0x01, 0x00, 0x0a, 0x0f, 0x01, 0x0d, 0x00, 0x41, 0x00, 0xfd, 0x0f, 0x41, 0x00,
0xfd, 0x0f, 0xfd, 0x80, 0x02, 0x0b, // i8x16.relaxed_swizzle
]);
const hasRelaxedSimd = WebAssembly.validate(RELAXED_PROBE);
Ship three builds if you need the widest reach — relaxed, standard SIMD and scalar — or two, since standard SIMD is universal in current browsers. Machine-learning libraries such as ONNX Runtime Web and TensorFlow.js already do this selection internally. The general approach is in shipping SIMD and baseline builds together.
Step 5 — test on more than one architecture
Because relaxed results may differ between machines, test on both x86-64 and ARM64 — a laptop and a phone, or CI runners of both kinds — and compare outputs with a tolerance rather than exact equality. A test that passes on your x86 development machine and fails on an ARM phone is the characteristic relaxed-SIMD bug, and the cure is a tolerance chosen from measured differences, not disabling the test.
Choosing tolerances for tests
The one practical change relaxed SIMD forces on a project is in its tests. Bit-exact comparisons that worked with standard SIMD will fail on some
machines, so tests of relaxed code paths need tolerances, and those tolerances should be chosen deliberately rather than loosened until tests pass.
For fused multiply-add, the difference between fused and unfused results is bounded by roughly one rounding error per operation; a long dot
product accumulates those, so a relative tolerance of around 1e-5 for f32 accumulations over a few thousand terms is a reasonable start. For
inference, compare final outputs — the predicted class, the top-k tokens — exactly, and intermediate activations with a tolerance, since the
application cares about the former.
Measure the actual differences on the machines you have — an x86 laptop, an ARM phone or CI runner — and set the tolerance a small factor above the largest observed difference. Record those measurements next to the test, so the next person knows the tolerance came from data. And keep at least one test of the standard SIMD path with exact comparisons, so a change that accidentally introduces relaxed behaviour into code that must be deterministic is caught.
Why the proposal is designed this way
Relaxed SIMD is a careful compromise rather than a loosening of WebAssembly’s principles. Each instruction’s possible results are enumerated in the specification — “either the fused result or the unfused result”, “either of these two NaN behaviours” — so the nondeterminism is bounded and documented, not undefined behaviour. Results are fixed per machine, so a given device is internally consistent and reproducible in its own tests. And the instructions are opt-in, so modules that need strict determinism are unaffected. The design lets the code that most needs raw speed — inference kernels — use the hardware fully, while everything else keeps WebAssembly’s portability guarantees.
Expected output
On a machine with FMA, the relaxed dot-product kernel produces results within a few ULPs of the standard version and runs faster; on a machine without FMA, it produces exactly the standard result at the standard speed.
Gotchas
- Golden-file tests fail on another CPU. Relaxed results vary by machine. Compare with tolerances in tests for relaxed code paths.
- Assuming relaxed means random. It does not. Results are fixed per machine; nondeterminism is between machines.
- Using relaxed min/max on NaN-heavy data. NaN handling is exactly what varies. Filter NaNs or use the standard instructions.
- No fallback. Older engines reject the module. Detect and load a standard SIMD build where needed.
Performance note
For a quantized 8-bit matrix multiply in Chrome on a laptop with VNNI support, the relaxed dot-product path was 2.4× faster than standard SIMD
emulating the same arithmetic with widening multiplies and adds. An f32 matrix multiply using relaxed_madd gained about 1.3× on the same
machine. On an ARM phone with dot-product support the int8 gain was similar.
Frequently Asked Questions
Is relaxed SIMD supported everywhere? It is supported in current Chromium-based browsers and Firefox and progressing elsewhere; detect it rather than assuming.
Does relaxed SIMD break Wasm’s security guarantees? No. Results vary only in values, never in memory safety or control flow.
Can I use it from Rust?
Yes — core::arch::wasm32 exposes relaxed intrinsics behind the relaxed-simd target feature, currently on nightly or recent stable releases.
Do autovectorizers emit relaxed instructions? Generally not by default, because they would change results. Use intrinsics or a library that opts in explicitly.
Related
- Wasm SIMD & vectorized computation — the standard proposal this extends.
- Detecting proposal support at runtime — the detection technique used above.
- Choosing between the Wasm and WebGPU backends — where relaxed SIMD matters for inference.
- Benchmarking SIMD vs scalar Wasm kernels — measuring the difference fairly.
← Back to Wasm SIMD & Vectorized Computation