Writing SIMD from C with wasm_simd128.h
This page answers one task: write explicitly vectorized C code for WebAssembly using 128-bit SIMD intrinsics, for loops the compiler will not vectorize on its own.
Prerequisites
- [ ] clang 13+ with the
wasm32target, from the WASI SDK, emsdk or an LLVM release. - [ ] A hot loop over arrays — a filter, a sum, a conversion — that you have profiled.
- [ ] A browser or runtime with SIMD support (every current engine).
What the header gives you
WebAssembly’s SIMD proposal, now standard, adds a single 128-bit vector type, v128, and a few hundred instructions that treat it as lanes:
sixteen 8-bit integers, eight 16-bit, four 32-bit, two 64-bit, four f32 or two f64. Engines map these onto the host CPU’s vector units —
SSE and AVX on x86-64, NEON on ARM — so one Wasm SIMD instruction becomes one or a few native instructions processing several values at once.
Clang exposes the instruction set to C through wasm_simd128.h. It defines v128_t and an intrinsic function for nearly every instruction,
named by lane type: wasm_f32x4_add, wasm_i16x8_mul, wasm_u8x16_avgr. Unlike autovectorization — covered in
autovectorizing loops for Wasm SIMD —
intrinsics give you exact control over which instructions run, at the cost of writing the vector code yourself.
Step 1 — compile with SIMD enabled
clang --target=wasm32-wasip1 -O3 -msimd128 -c blend.c -o blend.o
# or with Emscripten
emcc -O3 -msimd128 blend.c -o blend.mjs -sMODULARIZE -sEXPORT_ES6
-msimd128 enables the feature; without it, including wasm_simd128.h produces errors about unavailable builtins. Remember that a module built
with SIMD will not compile in an engine without SIMD support — every current browser has it, but if you need a fallback, see
shipping SIMD and baseline builds together.
Step 2 — write a vectorized loop with a scalar tail
The standard shape of a SIMD loop processes four (or eight, or sixteen) elements per iteration, then handles the remainder one at a time:
#include <wasm_simd128.h>
#include <stddef.h>
// out[i] = a[i] * t + b[i] * (1 - t)
__attribute__((export_name("blend")))
void blend(float *out, const float *a, const float *b, float t, size_t n) {
v128_t vt = wasm_f32x4_splat(t);
v128_t v1t = wasm_f32x4_splat(1.0f - t);
size_t i = 0;
for (; i + 4 <= n; i += 4) {
v128_t va = wasm_v128_load(a + i);
v128_t vb = wasm_v128_load(b + i);
v128_t r = wasm_f32x4_add(wasm_f32x4_mul(va, vt), wasm_f32x4_mul(vb, v1t));
wasm_v128_store(out + i, r);
}
for (; i < n; i++) out[i] = a[i] * t + b[i] * (1.0f - t); // scalar tail
}
wasm_v128_load and wasm_v128_store accept unaligned addresses — WebAssembly memory accesses are allowed to be unaligned — so there is no
alignment prologue as in some native SIMD code. Alignment can still affect speed on some hardware; allocating buffers on 16-byte boundaries is a
cheap habit.
Step 3 — work with integer lanes and saturation
Image and audio code mostly works on narrow integers, where SIMD helps most because sixteen bytes fit in one vector. Saturating and averaging operations exist for exactly these cases:
// brighten 8-bit pixels by `delta`, clamping at 255 — sixteen bytes per iteration
__attribute__((export_name("brighten")))
void brighten(uint8_t *px, size_t n, uint8_t delta) {
v128_t d = wasm_u8x16_splat(delta);
size_t i = 0;
for (; i + 16 <= n; i += 16) {
v128_t v = wasm_v128_load(px + i);
wasm_v128_store(px + i, wasm_u8x16_add_sat(v, d));
}
for (; i < n; i++) { unsigned s = px[i] + delta; px[i] = s > 255 ? 255 : s; }
}
The scalar version needs a comparison per byte; the SIMD version processes sixteen bytes with one saturating add. Lane-widening operations —
wasm_u16x8_extend_low_u8x16 and friends — convert narrow lanes to wider ones when intermediate results need more bits, and narrowing
operations with saturation convert back.
Step 4 — port SSE code with the compatibility headers
Existing x86 SIMD code using SSE intrinsics can often be compiled for WebAssembly without rewriting, through compatibility headers that Emscripten provides:
emcc -O3 -msimd128 -msse4.1 legacy_filter.c -o filter.mjs
#include <smmintrin.h> // SSE4.1, mapped onto Wasm SIMD by Emscripten's headers
__m128 r = _mm_add_ps(_mm_mul_ps(a, b), c);
Most SSE operations map one-to-one onto Wasm SIMD instructions; some — certain shuffles, horizontal operations, _mm_movemask variants — have no
direct equivalent and are emulated with several instructions, which can be slower than the native original. The compatibility path is a quick
way to get existing code running; profile afterwards and rewrite the hottest emulated operations with native wasm_* intrinsics.
Step 5 — check the output and measure
Confirm the compiler emitted SIMD instructions where you expected:
wasm2wat blend.wasm | grep -oE 'f32x4\.[a-z_]+|v128\.(load|store)' | sort | uniq -c
1 f32x4.add
2 f32x4.mul
2 v128.load
1 v128.store
Then benchmark against the scalar version with real data sizes, warm-up and medians, as in benchmarking SIMD vs scalar Wasm kernels. SIMD helps compute-bound loops most; a loop already limited by memory bandwidth may barely move.
Structuring data for SIMD
The speed of SIMD code depends as much on data layout as on the instructions. Vector loads read sixteen contiguous bytes, so SIMD works best
when the values a loop processes together sit next to each other in memory. An array of structures — { x, y, z, w } repeated — puts one
point’s coordinates together but scatters each coordinate across the array, so a loop that scales every x must load whole structures and pick
lanes out. A structure of arrays — separate arrays for all x, all y, all z — puts like values side by side, and the same loop becomes a
straight run of loads, multiplies and stores.
Image data is a common case. Interleaved RGBA pixels suit operations that treat all four channels the same way — brightness, blending — because a
16-byte load holds four whole pixels. Operations that treat channels differently — converting to grayscale with different weights per channel —
either need shuffles to separate channels or benefit from planar storage. When designing the memory layout of a Wasm module that will be
vectorized, think about the inner loop first and choose the layout that gives it contiguous, same-typed data. Shuffles are possible with
wasm_i8x16_shuffle, but every shuffle is an instruction that does no arithmetic.
Finally, keep vector loops free of function calls and unpredictable branches. A call inside the loop forces vector values to be spilled and
reloaded; a data-dependent branch per element undoes the point of processing lanes together. Replace per-lane conditions with comparisons and
wasm_v128_bitselect, which computes both outcomes and chooses lane by lane.
Choosing between intrinsics and letting the compiler vectorize
Intrinsics are the right tool when the compiler cannot or will not vectorize a loop: when the loop needs saturating or averaging arithmetic the
compiler does not infer, when lanes must be shuffled or blended in specific ways, when the loop has a structure — gathers, conditional updates —
that defeats autovectorization, or when performance must be predictable across compiler versions. They are the wrong tool when a plain loop
already vectorizes: the compiler then handles tails, alignment and unrolling for you, and the code stays portable to other targets. A practical
workflow is to write the scalar loop first, check whether -O3 -msimd128 vectorizes it, and reach for intrinsics only where it does not and the
profile says the loop matters.
Expected output
For a 1920×1080 RGBA frame, brighten processes 8.3 million bytes; with SIMD the result matches the scalar version byte for byte, and the console
timing shows the vector version several times faster.
Gotchas
'__builtin_wasm_…' needs target feature simd128. Compile with-msimd128.- Missing scalar tail. Loops that step by 4 or 16 skip the remainder when
nis not a multiple. Always handle the tail. - Float results differ slightly. Vectorized code may evaluate in a different order; floating-point addition is not associative. Compare with a tolerance.
- Emulated SSE operations are slow. Some intrinsics have no direct Wasm equivalent. Profile and replace them.
Performance note
On a laptop in Chrome, the SIMD brighten processed a full HD frame in 0.31 ms against 1.9 ms for the scalar loop — about 6×. The blend kernel on
f32 data gained about 3.4×, closer to the four-lane theoretical maximum because each element needs more arithmetic per byte loaded.
Frequently Asked Questions
Is 128 bits the maximum vector width? Yes, for the standard SIMD proposal. Wider vectors are not part of WebAssembly; engines use 128-bit native instructions even on CPUs with AVX-512.
Can I use these intrinsics from C++? Yes — the header works in C++ too, and some libraries wrap it in C++ types.
What about relaxed SIMD? Relaxed SIMD adds faster operations with implementation-defined results, such as fused multiply-add; see understanding relaxed SIMD.
How do I write the same thing in Rust?
Rust’s core::arch::wasm32 module provides equivalent intrinsics; see
writing v128 SIMD intrinsics in Rust.
Related
- Detecting SIMD support at runtime — choosing the build to load.
- Vectorizing byte scanning with SIMD — a harder vectorization pattern.
- Building a Wasm image filter pipeline — where these kernels live.
- Benchmarking memory bandwidth in Wasm — knowing when SIMD cannot help.
← Back to Wasm SIMD & Vectorized Computation