How Wasm Locals Map to Machine Registers

This page answers one question: a WebAssembly function declares locals and pushes values on an operand stack — what do those become on a real CPU, and does the shape of the WAT affect how fast the compiled code runs?

Prerequisites

  • [ ] Familiarity with the operand stack, as in how the Wasm operand stack works.
  • [ ] Node 20+ for the flag experiments, or wasmtime with Cranelift.
  • [ ] A small numeric function to compile and compare.

Locals are not memory — they are candidates for registers

A WebAssembly local is a typed, unaddressable slot that only local.get, local.set and local.tee can touch, and only within its own function. That makes it exactly what a compiler wants as a register candidate: nothing else can observe or modify it, so the engine is free to keep its value wherever is fastest. The operand stack is the same — its values never leak out of the function except as arguments and results.

So engines treat both as virtual registers. The job of the engine’s compiler is to map an unlimited number of virtual registers onto the CPU’s small set of physical ones — sixteen general-purpose and sixteen vector registers on x86-64, thirty-one on ARM64 — and to put the overflow, when there is any, in spill slots on the engine’s native stack. How well it does that depends on which compiler tier is running.

From WAT locals to machine registers in two tiers WAT locals and operand-stack values are virtual registers. The baseline tier keeps a simple cache of values in registers and spills often. The optimizing tier builds an SSA graph, allocates registers globally, and spills only under real pressure. WAT: (local $a i32) (local $b i32) … virtual registers, unlimited, typed baseline tier (Liftoff, BBQ) one pass; values cached in registers, spilled at calls and merges optimizing tier (TurboFan, Ion, Cranelift) SSA graph; global register allocation CPU registers + spill slots 16-31 physical registers; spills on the native stack

Step 1 — see the baseline tier’s approach

A baseline compiler compiles each function in one forward pass, as fast as possible. It cannot see ahead, so it uses simple rules: keep recently produced values in registers, reload locals from their stack slots when the cache misses, and spill everything to the native stack at points where control flow merges or a call happens, so all paths agree on where values live. The result is correct and quickly generated, with more memory traffic than necessary.

You can see the effect by forcing the baseline tier in Node and timing a hot loop:

node --liftoff --no-wasm-tier-up bench.mjs      # baseline only
node --no-liftoff bench.mjs                     # optimizing only

For a tight numeric loop, the baseline-only run is typically two to four times slower. That is not the cost of “stack machine” bytecode — it is the cost of compiling without global analysis. In normal operation, engines run baseline code first and replace hot functions with optimized code within moments, as described in how V8 compiles Wasm with Liftoff and TurboFan.

Step 2 — see the optimizing tier’s approach

An optimizing compiler first converts the function into static single assignment (SSA) form: every computed value gets a unique name, and locals disappear as such — local.set $x simply means “from here on, $x refers to this value”. Two different values written to the same local at different points become two different SSA values, possibly in two different registers. Then it runs ordinary optimizations — constant folding, common subexpression elimination, loop-invariant code motion — and finally a register allocator that assigns physical registers across the whole function, spilling only the values with the longest lives under the most pressure.

The consequence is that WAT-level choices about locals barely matter. Reusing one local for two unrelated purposes, or declaring fifty locals where five would do, produces the same SSA graph and the same machine code. The compiler sees through the local declarations to the data flow.

;; two shapes of the same computation — the optimizing tier emits identical code
(func $a (param $x i32) (result i32)
  (local $t i32)
  (local.set $t (i32.mul (local.get $x) (i32.const 3)))
  (i32.add (local.get $t) (i32.const 1)))

(func $b (param $x i32) (result i32)
  (i32.add (i32.mul (local.get $x) (i32.const 3)) (i32.const 1)))

Step 3 — know when registers really run out

Register pressure is real when a loop keeps many values alive at once: a SIMD kernel with a dozen accumulators, an unrolled loop holding many partial results, a function with many simultaneously live pointers. Past the number of physical registers, values must be spilled to the native stack and reloaded, and the loop slows. This is the same trade-off native compilers face, and the remedies are the same: fewer simultaneously live values, smaller unroll factors, splitting a large loop body into stages.

One WebAssembly-specific detail matters: the engine reserves a few registers for its own purposes — the base address of linear memory, the instance pointer, sometimes a bounds-check limit — so slightly fewer registers are available to your code than to the same code compiled natively. On x86-64 that leaves around twelve general-purpose registers for user values, which is why heavily unrolled kernels sometimes do better with a smaller unroll factor in Wasm than natively.

Effect of unroll factor on a SIMD kernel in Wasm Throughput of a vectorized dot-product kernel with different numbers of independent accumulators in Chrome on x86-64. Throughput improves up to four accumulators and falls at sixteen, when register pressure forces spills. GFLOP/s (higher is better) 1 accumulator 4.1 GFLOP/s 4 accumulators 11.8 GFLOP/s 8 accumulators 12.3 GFLOP/s 16 accumulators 7.9 GFLOP/s

Step 4 — keep values unaddressed

The one thing at the source level that reliably defeats register allocation is taking a variable’s address. An address-taken variable cannot be a WebAssembly local at all; the compiler spills it to the shadow stack in linear memory, and every access becomes a load or store — the subject of why Wasm cannot take the address of a local. That spill happens before the engine ever sees the code, so no amount of engine optimization undoes it. Passing values by value, returning results instead of writing through pointers, and letting the source compiler inline small helpers keep hot values in locals and therefore in registers.

Step 5 — measure instead of reading WAT

Because the optimizing tier rewrites everything, reading WAT is a poor guide to performance. Count instructions in a profile, not in the disassembly; compare optimized-tier timings, not baseline ones; and when register pressure is suspected, look at the engine’s generated machine code. V8 can print it:

node --no-liftoff --print-wasm-code bench.mjs 2>&1 | grep -A40 'kind: wasm function.*dot_kernel' | head -60

Lots of mov instructions to and from [rbp - N] inside the loop are spills. Few or none means the allocator found registers for everything.

Why the design leaves this to the engine

WebAssembly could have exposed registers directly — a register-based bytecode with a fixed register count — and some argued for it. The designers chose not to, because the right number of registers differs by CPU, and a bytecode tuned for x86-64’s sixteen would waste ARM64’s thirty-one or overflow on smaller architectures. Leaving allocation to the engine means each device gets code fitted to its own CPU. The cost is that the engine must do the work at load time, which is why tiering exists: a cheap allocation first, a thorough one for code that turns out to be hot. For you, the practical upshot is simple — write clear code with values passed by value, and let the compilers do their job.

Expected output

Timing the dot-product kernel with forced tiers in Node shows the gap and confirms that, in normal operation, the optimized timing is what users get once the function is hot:

--liftoff --no-wasm-tier-up   dot_kernel: 41.2 ms
--no-liftoff                  dot_kernel: 12.9 ms
default (after warm-up)       dot_kernel: 13.1 ms

Gotchas

  • Optimizing WAT by hand for register use. The optimizing tier rewrites locals entirely; hand-tuning local reuse changes nothing.
  • Benchmarking the baseline tier by accident. Short benchmarks run before tier-up. Warm up first.
  • Huge unroll factors. More accumulators than registers cause spills and slow the loop. Measure a few factors.
  • Address-taken scalars in hot code. These become memory accesses before the engine sees them. Keep them unaddressed.

Performance note

Across a set of numeric kernels, the optimizing tier ran 2.1–3.8 times faster than the baseline tier in V8, and within 10–30% of native code compiled with Clang at -O2 on the same machine. Most of the remaining gap came from bounds-checking and the registers reserved for memory and instance pointers, not from the stack-machine encoding.

The same kernel under each compiler tier and natively Run time of a numeric kernel compiled by V8's baseline tier only, by its optimizing tier only, and natively with Clang at -O2 on the same laptop. ms per run (lower is better) V8 baseline tier only 41.2 ms V8 optimizing tier 12.9 ms native clang -O2 10.6 ms

Frequently Asked Questions

Do more locals make a module bigger? Slightly — each local declaration is a few bytes. They do not make it slower in the optimizing tier.

Does Cranelift behave the same way? Yes. wasmtime’s Cranelift compiles to SSA and allocates registers globally; it has no separate baseline tier, so all code is optimized from the start.

Are SIMD values in registers too? Yes. v128 locals map to vector registers (XMM/YMM on x86-64, NEON registers on ARM64) under the same rules.

Can I see Firefox’s generated code? SpiderMonkey has similar debugging switches in developer builds; the Firefox Profiler’s assembly view shows hot code for profiled functions.

← Back to Stack vs Heap Execution Model