Why Wasm Cannot Take the Address of a Local
This page answers one question: in C, Rust or C++ you can take a pointer to any local variable, but WebAssembly locals have no addresses — why was it designed that way, what do compilers do instead, and does it affect how you should write performance-sensitive code?
Prerequisites
- [ ] A C or Rust function to experiment with, and a compiler targeting
wasm32. - [ ]
wasm2watto read the compiled output. - [ ] Background on the shadow stack from understanding the shadow stack in linear memory.
Locals are slots, not memory
A WebAssembly local is declared with a type — (local $count i32) — and accessed with local.get and local.set by index. That is all.
There is no instruction that produces the address of a local, no way to load or store through a local’s location, and no way for another
function to reach it. The only addressable storage in a module is linear memory.
The restriction is deliberate and buys several things at once. It keeps locals entirely under the engine’s control, so it can place them in
registers without worrying that some pointer might alias them — the optimizer knows every read and write of a local is an explicit
local.get or local.set. It keeps the engine’s native stack, where spilled locals and return addresses live, out of reach of the module,
which eliminates the classic stack-smashing attack: no buffer overflow in linear memory can overwrite a return address. And it makes
validation and compilation simpler and faster, because locals are just typed registers in a function-sized namespace.
Step 1 — see what the compiler does with &x
When source code takes the address of a local, the compiler cannot keep that variable in a WebAssembly local. It spills it to the shadow stack: allocates space in the function’s frame in linear memory, stores the value there, and uses the memory address as the pointer.
// spill.c
void read_value(int *out);
int doubled(void) {
int v; // address taken below → must live in memory
read_value(&v);
return v * 2;
}
int tripled(int x) {
int y = x * 3; // never addressed → stays a Wasm local
return y;
}
clang --target=wasm32 -O2 -c spill.c -o spill.o && wasm2wat spill.o | grep -A16 'func $doubled'
(func $doubled (result i32)
(local i32)
global.get $__stack_pointer
i32.const 16
i32.sub
local.tee 0 ;; frame pointer
global.set $__stack_pointer
local.get 0
i32.const 12
i32.add ;; &v = frame + 12
call $read_value
local.get 0
i32.load offset=12 ;; v loaded back from memory
...
tripled compiles to two instructions and no memory traffic at all. doubled gets a stack frame, a store through read_value, and a load —
because of one &.
Step 2 — recognise patterns that force spills
Some spills are obvious, others are not. Common causes:
- Passing
&xor a reference to a function the compiler cannot inline. - Arrays and structs declared as locals, which need addresses for indexing and field access unless the optimizer can break them apart.
- Variables captured by reference in closures or lambdas.
- Rust values borrowed mutably across a call (
foo(&mut state)), and large values that the optimizer cannot keep in a few locals. - Variadic calls, whose arguments are laid out in memory.
The optimizer removes many spills by inlining the callee — once read_value is inlined, v can return to a local. So spills in -O0 builds
are far more common than in optimized builds, and judging performance from debug builds is misleading.
Step 3 — keep hot values in locals
In performance-sensitive inner loops, spills turn register arithmetic into memory traffic. A few habits keep hot values unaddressed:
// spills `acc` if `update` is not inlined: it takes &mut acc
fn sum_slow(xs: &[u32]) -> u32 {
let mut acc = 0;
for &x in xs { update(&mut acc, x); }
acc
}
// keeps `acc` in a local: values in, value out
fn sum_fast(xs: &[u32]) -> u32 {
let mut acc = 0;
for &x in xs { acc = updated(acc, x); }
acc
}
Return values instead of writing through out-pointers; mark small helpers #[inline] (or static inline in C) so the optimizer can see
through them; and avoid passing references to loop state into functions that live in other crates or compilation units, where inlining
needs LTO to happen. Multi-value returns help here too, since a function can hand back several values without an out-pointer — see
returning multiple values from Wasm functions.
Step 4 — find spills in real code
Search the disassembly for functions that adjust __stack_pointer — those have shadow frames — and look inside hot ones for loads and stores
relative to the frame pointer:
wasm2wat app.wasm | awk '/\(func /{name=$2} /global.set \$__stack_pointer/{print name}' | sort | uniq -c | sort -rn | head
Cross-reference with a profile: a hot function with a shadow frame and frame-relative loads inside its loop is a candidate for restructuring. Functions with frames that run once per call are not worth touching.
Step 5 — accept spills where they belong
Spills are not bugs. Arrays, buffers, structs passed by reference and anything genuinely shared across functions should live in memory. The goal is only to avoid accidental spills of scalars in hot loops. A function that processes an image buffer will always read and write memory — that is its job — but the loop counter, the accumulator and the current pixel should be in locals.
What this means for language design
The no-address rule ripples upward into how languages target WebAssembly. Languages with garbage collectors traditionally scanned the native stack for pointers; in WebAssembly they cannot read the native stack, so they must keep pointers in the shadow stack or in memory where the collector can find them — one reason the GC proposal, which lets the engine manage objects directly, matters so much for them, as covered in using Wasm GC for managed languages. Languages with coroutines or stack switching cannot capture the native stack either, which is why features like Asyncify and JSPI exist. In each case, the inability to address locals is a constraint that pushes runtime machinery into explicit, inspectable forms — less convenient for implementers, and much safer for everyone running the result.
Expected output
Comparing the two Rust functions above in a disassembly, sum_fast contains no __stack_pointer adjustment and a loop of register-only
arithmetic; sum_slow, if update is not inlined, contains a frame, a store before each call and a load after it.
Gotchas
- Judging spills from a debug build.
-O0spills nearly everything. Look at optimized output. - Out-pointers in hot APIs.
void f(int *out)style interfaces force callers to spill. Return values instead where possible. - Closures capturing by reference. A closure that captures a loop variable by reference forces it into memory. Capture by value where the closure does not need to modify it.
- Assuming inlining across crates. Without LTO, functions in other crates are often not inlined, so their reference parameters force
spills. Enable LTO or mark small functions
#[inline]. - Trying to emulate addresses with tables or globals. Globals are not per-call and tables hold references, not data. Use memory.
Performance note
In a microbenchmark summing ten million integers, the version that passed &mut acc to a non-inlined helper ran in 23 ms; the version that
returned the new value ran in 6 ms with the same helper logic. With LTO enabling inlining, both compiled to the same 6 ms loop.
Frequently Asked Questions
Can I get the address of a WebAssembly global? No — globals are also unaddressable. Compiled languages put their mutable global variables in linear memory instead, and use Wasm globals only for special values such as the stack pointer.
Does this make WebAssembly slower than native? Rarely in practice. Optimizing compilers keep most values in locals, and engines map locals to registers; spills cost about what they cost natively.
Do GC proposal references count as addresses? GC references point to engine-managed objects, not linear memory, and cannot be converted to integers. They do not change the rule for locals.
Is the restriction ever relaxed by a proposal? No proposal adds addressable locals; the design is considered a feature. Proposals instead reduce the need for spills — multi-value returns, GC references and stack switching each remove a reason compilers had to put values in memory.
Can tools show which variables were spilled? Not directly. Reading the WAT for frame-relative loads and stores is the practical method; DWARF can map those addresses back to variable names in a debugger.
Related
- How Wasm locals map to machine registers — what happens to locals that are not spilled.
- How the Wasm operand stack works — the other unaddressable stack.
- Tuning LTO and codegen-units for Wasm — enabling the inlining that removes spills.
- Browser sandbox & security boundaries — the security model this design supports.
← Back to Stack vs Heap Execution Model