Writing SIMD from C with wasm_simd128.h

This page answers one task: write explicitly vectorized C code for WebAssembly using 128-bit SIMD intrinsics, for loops the compiler will not vectorize on its own.

Prerequisites

  • [ ] clang 13+ with the wasm32 target, from the WASI SDK, emsdk or an LLVM release.
  • [ ] A hot loop over arrays — a filter, a sum, a conversion — that you have profiled.
  • [ ] A browser or runtime with SIMD support (every current engine).

What the header gives you

WebAssembly’s SIMD proposal, now standard, adds a single 128-bit vector type, v128, and a few hundred instructions that treat it as lanes: sixteen 8-bit integers, eight 16-bit, four 32-bit, two 64-bit, four f32 or two f64. Engines map these onto the host CPU’s vector units — SSE and AVX on x86-64, NEON on ARM — so one Wasm SIMD instruction becomes one or a few native instructions processing several values at once.

Clang exposes the instruction set to C through wasm_simd128.h. It defines v128_t and an intrinsic function for nearly every instruction, named by lane type: wasm_f32x4_add, wasm_i16x8_mul, wasm_u8x16_avgr. Unlike autovectorization — covered in autovectorizing loops for Wasm SIMD — intrinsics give you exact control over which instructions run, at the cost of writing the vector code yourself.

One v128 value viewed as different lane shapes A single 128-bit WebAssembly SIMD register can be interpreted as sixteen 8-bit lanes, eight 16-bit lanes, four 32-bit lanes for i32 or f32, or two 64-bit lanes. The intrinsic name chooses the interpretation for each operation. one v128 (128 bits) — f32x4 view lane 0 (f32) lane 1 (f32) lane 2 (f32) lane 3 (f32) bit 0 bit 64 bit 128

Step 1 — compile with SIMD enabled

clang --target=wasm32-wasip1 -O3 -msimd128 -c blend.c -o blend.o
# or with Emscripten
emcc -O3 -msimd128 blend.c -o blend.mjs -sMODULARIZE -sEXPORT_ES6

-msimd128 enables the feature; without it, including wasm_simd128.h produces errors about unavailable builtins. Remember that a module built with SIMD will not compile in an engine without SIMD support — every current browser has it, but if you need a fallback, see shipping SIMD and baseline builds together.

Step 2 — write a vectorized loop with a scalar tail

The standard shape of a SIMD loop processes four (or eight, or sixteen) elements per iteration, then handles the remainder one at a time:

#include <wasm_simd128.h>
#include <stddef.h>

// out[i] = a[i] * t + b[i] * (1 - t)
__attribute__((export_name("blend")))
void blend(float *out, const float *a, const float *b, float t, size_t n) {
  v128_t vt  = wasm_f32x4_splat(t);
  v128_t v1t = wasm_f32x4_splat(1.0f - t);
  size_t i = 0;
  for (; i + 4 <= n; i += 4) {
    v128_t va = wasm_v128_load(a + i);
    v128_t vb = wasm_v128_load(b + i);
    v128_t r  = wasm_f32x4_add(wasm_f32x4_mul(va, vt), wasm_f32x4_mul(vb, v1t));
    wasm_v128_store(out + i, r);
  }
  for (; i < n; i++) out[i] = a[i] * t + b[i] * (1.0f - t);   // scalar tail
}

wasm_v128_load and wasm_v128_store accept unaligned addresses — WebAssembly memory accesses are allowed to be unaligned — so there is no alignment prologue as in some native SIMD code. Alignment can still affect speed on some hardware; allocating buffers on 16-byte boundaries is a cheap habit.

Step 3 — work with integer lanes and saturation

Image and audio code mostly works on narrow integers, where SIMD helps most because sixteen bytes fit in one vector. Saturating and averaging operations exist for exactly these cases:

// brighten 8-bit pixels by `delta`, clamping at 255 — sixteen bytes per iteration
__attribute__((export_name("brighten")))
void brighten(uint8_t *px, size_t n, uint8_t delta) {
  v128_t d = wasm_u8x16_splat(delta);
  size_t i = 0;
  for (; i + 16 <= n; i += 16) {
    v128_t v = wasm_v128_load(px + i);
    wasm_v128_store(px + i, wasm_u8x16_add_sat(v, d));
  }
  for (; i < n; i++) { unsigned s = px[i] + delta; px[i] = s > 255 ? 255 : s; }
}

The scalar version needs a comparison per byte; the SIMD version processes sixteen bytes with one saturating add. Lane-widening operations — wasm_u16x8_extend_low_u8x16 and friends — convert narrow lanes to wider ones when intermediate results need more bits, and narrowing operations with saturation convert back.

The anatomy of a SIMD loop iteration An annotated iteration of the brighten loop — splat once outside the loop, then load sixteen bytes, add with saturation, and store, followed by a scalar tail for leftover bytes. v128_t d = wasm_u8x16_splat(delta); broadcast once, outside the loop v128_t v = wasm_v128_load(px + i); 16 bytes in one load v = wasm_u8x16_add_sat(v, d); 16 clamped adds in one op wasm_v128_store(px + i, v); 16 bytes in one store for (; i < n; i++) { … } scalar tail for the remainder

Step 4 — port SSE code with the compatibility headers

Existing x86 SIMD code using SSE intrinsics can often be compiled for WebAssembly without rewriting, through compatibility headers that Emscripten provides:

emcc -O3 -msimd128 -msse4.1 legacy_filter.c -o filter.mjs
#include <smmintrin.h>      // SSE4.1, mapped onto Wasm SIMD by Emscripten's headers
__m128 r = _mm_add_ps(_mm_mul_ps(a, b), c);

Most SSE operations map one-to-one onto Wasm SIMD instructions; some — certain shuffles, horizontal operations, _mm_movemask variants — have no direct equivalent and are emulated with several instructions, which can be slower than the native original. The compatibility path is a quick way to get existing code running; profile afterwards and rewrite the hottest emulated operations with native wasm_* intrinsics.

Step 5 — check the output and measure

Confirm the compiler emitted SIMD instructions where you expected:

wasm2wat blend.wasm | grep -oE 'f32x4\.[a-z_]+|v128\.(load|store)' | sort | uniq -c
      1 f32x4.add
      2 f32x4.mul
      2 v128.load
      1 v128.store

Then benchmark against the scalar version with real data sizes, warm-up and medians, as in benchmarking SIMD vs scalar Wasm kernels. SIMD helps compute-bound loops most; a loop already limited by memory bandwidth may barely move.

Structuring data for SIMD

The speed of SIMD code depends as much on data layout as on the instructions. Vector loads read sixteen contiguous bytes, so SIMD works best when the values a loop processes together sit next to each other in memory. An array of structures — { x, y, z, w } repeated — puts one point’s coordinates together but scatters each coordinate across the array, so a loop that scales every x must load whole structures and pick lanes out. A structure of arrays — separate arrays for all x, all y, all z — puts like values side by side, and the same loop becomes a straight run of loads, multiplies and stores.

Image data is a common case. Interleaved RGBA pixels suit operations that treat all four channels the same way — brightness, blending — because a 16-byte load holds four whole pixels. Operations that treat channels differently — converting to grayscale with different weights per channel — either need shuffles to separate channels or benefit from planar storage. When designing the memory layout of a Wasm module that will be vectorized, think about the inner loop first and choose the layout that gives it contiguous, same-typed data. Shuffles are possible with wasm_i8x16_shuffle, but every shuffle is an instruction that does no arithmetic.

Finally, keep vector loops free of function calls and unpredictable branches. A call inside the loop forces vector values to be spilled and reloaded; a data-dependent branch per element undoes the point of processing lanes together. Replace per-lane conditions with comparisons and wasm_v128_bitselect, which computes both outcomes and chooses lane by lane.

Choosing between intrinsics and letting the compiler vectorize

Intrinsics are the right tool when the compiler cannot or will not vectorize a loop: when the loop needs saturating or averaging arithmetic the compiler does not infer, when lanes must be shuffled or blended in specific ways, when the loop has a structure — gathers, conditional updates — that defeats autovectorization, or when performance must be predictable across compiler versions. They are the wrong tool when a plain loop already vectorizes: the compiler then handles tails, alignment and unrolling for you, and the code stays portable to other targets. A practical workflow is to write the scalar loop first, check whether -O3 -msimd128 vectorizes it, and reach for intrinsics only where it does not and the profile says the loop matters.

Expected output

For a 1920×1080 RGBA frame, brighten processes 8.3 million bytes; with SIMD the result matches the scalar version byte for byte, and the console timing shows the vector version several times faster.

Gotchas

  • '__builtin_wasm_…' needs target feature simd128. Compile with -msimd128.
  • Missing scalar tail. Loops that step by 4 or 16 skip the remainder when n is not a multiple. Always handle the tail.
  • Float results differ slightly. Vectorized code may evaluate in a different order; floating-point addition is not associative. Compare with a tolerance.
  • Emulated SSE operations are slow. Some intrinsics have no direct Wasm equivalent. Profile and replace them.

Performance note

On a laptop in Chrome, the SIMD brighten processed a full HD frame in 0.31 ms against 1.9 ms for the scalar loop — about 6×. The blend kernel on f32 data gained about 3.4×, closer to the four-lane theoretical maximum because each element needs more arithmetic per byte loaded.

Scalar versus SIMD for two kernels Run time for a full HD frame or equivalent float array, scalar C versus wasm_simd128.h intrinsics, in Chrome on a laptop. ms per frame (lower is better) brighten, scalar 1.9 ms brighten, SIMD u8x16 0.3 ms blend, scalar f32 2.7 ms blend, SIMD f32x4 0.8 ms

Frequently Asked Questions

Is 128 bits the maximum vector width? Yes, for the standard SIMD proposal. Wider vectors are not part of WebAssembly; engines use 128-bit native instructions even on CPUs with AVX-512.

Can I use these intrinsics from C++? Yes — the header works in C++ too, and some libraries wrap it in C++ types.

What about relaxed SIMD? Relaxed SIMD adds faster operations with implementation-defined results, such as fused multiply-add; see understanding relaxed SIMD.

How do I write the same thing in Rust? Rust’s core::arch::wasm32 module provides equivalent intrinsics; see writing v128 SIMD intrinsics in Rust.

← Back to Wasm SIMD & Vectorized Computation