Reading SIMD Instructions in WAT
This page answers one task: you are looking at disassembled WebAssembly — your own build, or a library’s — and it is full of v128.load, i32x4.add and
i8x16.shuffle. You want to read what the code does, check that a loop really vectorized, and write small SIMD functions by hand in WAT to test ideas.
Prerequisites
- [ ]
wasm2wat,wasm-tools printorwasm-objdump -dto disassemble modules. - [ ]
wat2wasmorwasm-tools parseto assemble examples. - [ ] Basic WAT reading skills (locals, blocks, loops, memory access).
The v128 type and lane shapes
SIMD adds one value type, v128: 128 bits with no fixed interpretation. Each instruction decides how to view those bits through its prefix, called the lane
shape: i8x16 (sixteen 8-bit lanes), i16x8, i32x4, i64x2, f32x4 and f64x2. The same v128 can be loaded as bytes, added as i32x4, then compared as
f32x4 — no conversion instructions are needed to reinterpret bits, which is why reading SIMD WAT means tracking what each value means in your head.
Instructions that do not care about lanes use the v128 prefix: v128.load, v128.store, v128.and, v128.or, v128.xor, v128.not, v128.bitselect,
v128.any_true.
Step 1 — a first SIMD function
(module
(memory (export "memory") 1)
(func (export "add4") (param $a i32) (param $b i32) (param $out i32)
(v128.store (local.get $out)
(f32x4.add (v128.load (local.get $a)) (v128.load (local.get $b))))))
add4 adds four pairs of floats at once: load 16 bytes from each address, view both as four f32 lanes, add lane by lane, store 16 bytes. In the disassembly,
v128.load shows two immediates like any memory instruction — the alignment hint (4, meaning 2⁴ = 16 bytes) and the offset — but SIMD loads work at any
address; alignment is only a hint.
Step 2 — read a vectorized loop
(func (export "sum") (param $p i32) (param $n i32) (result i32)
(local $acc v128)
(block $done
(loop $next
(br_if $done (i32.lt_u (local.get $n) (i32.const 4)))
(local.set $acc (i32x4.add (local.get $acc) (v128.load (local.get $p))))
(local.set $p (i32.add (local.get $p) (i32.const 16)))
(local.set $n (i32.sub (local.get $n) (i32.const 4)))
(br $next)))
(i32.add
(i32.add (i32x4.extract_lane 0 (local.get $acc)) (i32x4.extract_lane 1 (local.get $acc)))
(i32.add (i32x4.extract_lane 2 (local.get $acc)) (i32x4.extract_lane 3 (local.get $acc)))))
This is the shape compilers emit for a vectorized reduction: a v128 accumulator, a loop that loads 16 bytes and adds four lanes per iteration with the pointer
advancing by 16, and a horizontal reduction after the loop that extracts and adds the four lanes. (A real compiler also emits a scalar loop for the leftover
elements when $n is not a multiple of 4.) Summing the integers 1–8 returns 36. When you look for vectorization in compiler output, the signs are a v128
local, v128.load in the loop body, a pointer step of 16, and extract_lane after the loop.
Step 3 — learn the lane-moving instructions
(i8x16.splat (i32.const 44)) ;; sixteen copies of ',' (44)
(i32x4.extract_lane 2 (local.get $v)) ;; lane 2 as an i32
(i32x4.replace_lane 0 (local.get $v) (i32.const 7)) ;; copy of $v with lane 0 = 7
(i8x16.shuffle 12 13 14 15 8 9 10 11 4 5 6 7 0 1 2 3
(local.get $v) (local.get $v)) ;; reverse the four 32-bit lanes
(i8x16.swizzle (local.get $table) (local.get $idx)) ;; byte lookup with indices in a vector
(v128.const i32x4 1 2 3 4) ;; a constant vector
shuffle takes sixteen constant byte indices selecting from the 32 bytes of two inputs (0–15 from the first, 16–31 from the second); swizzle takes the indices
from a run-time vector and yields 0 for out-of-range indices. Lane indices in extract_lane and replace_lane are immediates, never run-time values.
Step 4 — follow comparisons into branches
Comparisons produce masks: each lane becomes all-ones (true) or all-zeros (false). Code typically combines masks with v128.and/v128.or, selects values with
v128.bitselect, and leaves the vector world with i8x16.bitmask (one bit per lane, into an i32) or v128.any_true:
(i8x16.bitmask (i8x16.eq (v128.load (local.get $p)) (i8x16.splat (i32.const 44))))
For the text a,b,,c, commas sit at bytes 1, 3 and 4, so this returns 0b11010 — exactly the pattern used by fast CSV parsers and string searches.
Step 5 — map WAT to binary
All SIMD instructions use the 0xFD prefix followed by a LEB128 sub-opcode: v128.load is fd 00, v128.store is fd 0b, i32x4.add is fd ae 01, and
f32x4.add is fd e4 01. In wasm-objdump -d output, a run of fd bytes in a function marks SIMD code; their absence in a hot function you expected to be
vectorized means the compiler fell back to scalar code.
Signed, unsigned and saturating variants
Many integer instructions come in variants: _s and _u (signed and unsigned interpretation, for comparisons, extensions, narrowing, min/max), and _sat
(saturating add and subtract for i8x16 and i16x8, which clamp instead of wrapping — essential for pixel arithmetic). Reading a variant wrong is the most common
mistake when reviewing SIMD code: i8x16.lt_s and i8x16.lt_u give different answers for bytes ≥ 128.
Running the examples from JavaScript
Because v128 values cannot cross into JavaScript, every test goes through memory: write inputs with a typed array, call the exported function with addresses,
and read results back. The examples above, assembled with wat2wasm into one module, behave like this in Node.js:
const { instance } = await WebAssembly.instantiate(bytes);
const { memory, sum, add4 } = instance.exports;
const ints = new Int32Array(memory.buffer);
for (let k = 0; k < 8; k++) ints[k] = k + 1;
sum(0, 8); // 36
const f = new Float32Array(memory.buffer);
f.set([1, 2, 3, 4], 16); f.set([10, 20, 30, 40], 20); // a at byte 64, b at byte 80
add4(64, 80, 96);
f.slice(24, 28); // Float32Array [11, 22, 33, 44]
This small harness is worth keeping next to any hand-written SIMD WAT: it turns each instruction you are unsure about into a one-line experiment, which is faster and more reliable than reasoning about lane orders from the specification.
Lane order and endianness
Lanes are numbered from the lowest memory address: lane 0 of an i32x4 loaded from address p is the 32-bit little-endian integer at p..p+3. The same holds
for v128.const, whose first listed value is lane 0. Shuffle byte indices follow the same order, so reversing the four 32-bit lanes means selecting bytes 12–15
first, as in the example above. Keeping “lane 0 = lowest address” in mind resolves most confusion when reading shuffles and bitmasks, where bit 0 of the result
corresponds to lane 0.
Reading relaxed SIMD
Modules built with relaxed SIMD contain instructions with a relaxed_ prefix, such as f32x4.relaxed_madd or i8x16.relaxed_swizzle. They behave like their
non-relaxed counterparts for ordinary inputs but may differ across hardware for edge cases (rounding of fused multiply-add, out-of-range indices). When you see
them in a disassembly, the module requires an engine with relaxed SIMD support, and results may vary in the last bit between machines — something to account for
in tests that compare floats exactly.
Expected output
You can assemble and run small SIMD functions in WAT; read a compiler’s vectorized loop and identify the accumulator, the 16-byte pointer step and the horizontal
reduction; explain what splat, extract_lane, shuffle, swizzle, comparisons and bitmask do; and spot SIMD in a binary dump by its 0xFD prefix.
Gotchas
- Assuming lane types are stored. v128 has no lane type; each instruction picks one.
- Run-time lane indices. extract_lane and replace_lane need immediates; use swizzle for run-time selection.
- Signed versus unsigned variants. Bytes ≥ 128 compare differently. Check suffixes.
- Shuffle index ranges. 16–31 select from the second operand.
- Expecting alignment to matter for correctness. It is a hint only.
- Testing SIMD WAT only by reading it. Lane order is easy to get backwards. Run each example through memory.
Performance note
On the sum example over a million i32 values, the i32x4 loop ran about 3.5× faster than the scalar loop in Chrome; adding a second accumulator to break the
dependency chain pushed it to about 4.5×.
Frequently Asked Questions
Can JavaScript receive a v128 value? No — functions with v128 parameters or results cannot be called from JavaScript; pass data through memory.
What does load8_splat do? It loads one byte and copies it into all sixteen lanes.
Is v128.const expensive? No — engines materialise constants efficiently, often hoisting them out of loops.
How do I see SIMD in browser DevTools? The Sources panel disassembles Wasm to WAT, showing the same instruction names.
Which lane is lane 0? The one at the lowest memory address; bit 0 of a bitmask also corresponds to lane 0.
What do relaxed_ instructions mean in a disassembly? The module needs relaxed SIMD support, and some edge-case results may vary by hardware.
Related
- Writing SIMD from C with wasm_simd128.h — intrinsics in C.
- Autovectorizing loops for Wasm SIMD — compiler output.
- Writing loops and branches in WAT — WAT control flow.
- Decoding Wasm opcodes for debugging — opcode bytes.
← Back to Wasm SIMD & Vectorized Computation