How Wasm Locals Map to Machine Registers
This page answers one question: a WebAssembly function declares locals and pushes values on an operand stack — what do those become on a real CPU, and does the shape of the WAT affect how fast the compiled code runs?
Prerequisites
- [ ] Familiarity with the operand stack, as in how the Wasm operand stack works.
- [ ] Node 20+ for the flag experiments, or
wasmtimewith Cranelift. - [ ] A small numeric function to compile and compare.
Locals are not memory — they are candidates for registers
A WebAssembly local is a typed, unaddressable slot that only local.get, local.set and local.tee can touch, and only within its own
function. That makes it exactly what a compiler wants as a register candidate: nothing else can observe or modify it, so the engine is free to
keep its value wherever is fastest. The operand stack is the same — its values never leak out of the function except as arguments and results.
So engines treat both as virtual registers. The job of the engine’s compiler is to map an unlimited number of virtual registers onto the CPU’s small set of physical ones — sixteen general-purpose and sixteen vector registers on x86-64, thirty-one on ARM64 — and to put the overflow, when there is any, in spill slots on the engine’s native stack. How well it does that depends on which compiler tier is running.
Step 1 — see the baseline tier’s approach
A baseline compiler compiles each function in one forward pass, as fast as possible. It cannot see ahead, so it uses simple rules: keep recently produced values in registers, reload locals from their stack slots when the cache misses, and spill everything to the native stack at points where control flow merges or a call happens, so all paths agree on where values live. The result is correct and quickly generated, with more memory traffic than necessary.
You can see the effect by forcing the baseline tier in Node and timing a hot loop:
node --liftoff --no-wasm-tier-up bench.mjs # baseline only
node --no-liftoff bench.mjs # optimizing only
For a tight numeric loop, the baseline-only run is typically two to four times slower. That is not the cost of “stack machine” bytecode — it is the cost of compiling without global analysis. In normal operation, engines run baseline code first and replace hot functions with optimized code within moments, as described in how V8 compiles Wasm with Liftoff and TurboFan.
Step 2 — see the optimizing tier’s approach
An optimizing compiler first converts the function into static single assignment (SSA) form: every computed value gets a unique name, and
locals disappear as such — local.set $x simply means “from here on, $x refers to this value”. Two different values written to the same local
at different points become two different SSA values, possibly in two different registers. Then it runs ordinary optimizations — constant
folding, common subexpression elimination, loop-invariant code motion — and finally a register allocator that assigns physical registers across
the whole function, spilling only the values with the longest lives under the most pressure.
The consequence is that WAT-level choices about locals barely matter. Reusing one local for two unrelated purposes, or declaring fifty locals where five would do, produces the same SSA graph and the same machine code. The compiler sees through the local declarations to the data flow.
;; two shapes of the same computation — the optimizing tier emits identical code
(func $a (param $x i32) (result i32)
(local $t i32)
(local.set $t (i32.mul (local.get $x) (i32.const 3)))
(i32.add (local.get $t) (i32.const 1)))
(func $b (param $x i32) (result i32)
(i32.add (i32.mul (local.get $x) (i32.const 3)) (i32.const 1)))
Step 3 — know when registers really run out
Register pressure is real when a loop keeps many values alive at once: a SIMD kernel with a dozen accumulators, an unrolled loop holding many partial results, a function with many simultaneously live pointers. Past the number of physical registers, values must be spilled to the native stack and reloaded, and the loop slows. This is the same trade-off native compilers face, and the remedies are the same: fewer simultaneously live values, smaller unroll factors, splitting a large loop body into stages.
One WebAssembly-specific detail matters: the engine reserves a few registers for its own purposes — the base address of linear memory, the instance pointer, sometimes a bounds-check limit — so slightly fewer registers are available to your code than to the same code compiled natively. On x86-64 that leaves around twelve general-purpose registers for user values, which is why heavily unrolled kernels sometimes do better with a smaller unroll factor in Wasm than natively.
Step 4 — keep values unaddressed
The one thing at the source level that reliably defeats register allocation is taking a variable’s address. An address-taken variable cannot be a WebAssembly local at all; the compiler spills it to the shadow stack in linear memory, and every access becomes a load or store — the subject of why Wasm cannot take the address of a local. That spill happens before the engine ever sees the code, so no amount of engine optimization undoes it. Passing values by value, returning results instead of writing through pointers, and letting the source compiler inline small helpers keep hot values in locals and therefore in registers.
Step 5 — measure instead of reading WAT
Because the optimizing tier rewrites everything, reading WAT is a poor guide to performance. Count instructions in a profile, not in the disassembly; compare optimized-tier timings, not baseline ones; and when register pressure is suspected, look at the engine’s generated machine code. V8 can print it:
node --no-liftoff --print-wasm-code bench.mjs 2>&1 | grep -A40 'kind: wasm function.*dot_kernel' | head -60
Lots of mov instructions to and from [rbp - N] inside the loop are spills. Few or none means the allocator found registers for everything.
Why the design leaves this to the engine
WebAssembly could have exposed registers directly — a register-based bytecode with a fixed register count — and some argued for it. The designers chose not to, because the right number of registers differs by CPU, and a bytecode tuned for x86-64’s sixteen would waste ARM64’s thirty-one or overflow on smaller architectures. Leaving allocation to the engine means each device gets code fitted to its own CPU. The cost is that the engine must do the work at load time, which is why tiering exists: a cheap allocation first, a thorough one for code that turns out to be hot. For you, the practical upshot is simple — write clear code with values passed by value, and let the compilers do their job.
Expected output
Timing the dot-product kernel with forced tiers in Node shows the gap and confirms that, in normal operation, the optimized timing is what users get once the function is hot:
--liftoff --no-wasm-tier-up dot_kernel: 41.2 ms
--no-liftoff dot_kernel: 12.9 ms
default (after warm-up) dot_kernel: 13.1 ms
Gotchas
- Optimizing WAT by hand for register use. The optimizing tier rewrites locals entirely; hand-tuning local reuse changes nothing.
- Benchmarking the baseline tier by accident. Short benchmarks run before tier-up. Warm up first.
- Huge unroll factors. More accumulators than registers cause spills and slow the loop. Measure a few factors.
- Address-taken scalars in hot code. These become memory accesses before the engine sees them. Keep them unaddressed.
Performance note
Across a set of numeric kernels, the optimizing tier ran 2.1–3.8 times faster than the baseline tier in V8, and within 10–30% of native code
compiled with Clang at -O2 on the same machine. Most of the remaining gap came from bounds-checking and the registers reserved for memory and
instance pointers, not from the stack-machine encoding.
Frequently Asked Questions
Do more locals make a module bigger? Slightly — each local declaration is a few bytes. They do not make it slower in the optimizing tier.
Does Cranelift behave the same way? Yes. wasmtime’s Cranelift compiles to SSA and allocates registers globally; it has no separate baseline tier, so all code is optimized from the start.
Are SIMD values in registers too?
Yes. v128 locals map to vector registers (XMM/YMM on x86-64, NEON registers on ARM64) under the same rules.
Can I see Firefox’s generated code? SpiderMonkey has similar debugging switches in developer builds; the Firefox Profiler’s assembly view shows hot code for profiled functions.
Related
- Why the first call into Wasm is slow — baseline code in practice.
- Writing v128 SIMD intrinsics in Rust — kernels where register pressure matters.
- Avoiding JIT warm-up errors in Wasm benchmarks — measuring the right tier.
- Stack vs heap execution model — the wider memory model.
← Back to Stack vs Heap Execution Model