Understanding the Shadow Stack in Linear Memory

This page answers one question: WebAssembly already has an operand stack and the engine has a call stack — why do compiled C, C++ and Rust programs maintain a third stack inside linear memory, and what do you need to know about it?

Prerequisites

Why a second stack is needed

In native code, local variables live on the machine stack, and any of them can have its address taken: &x is just an address in the stack region. WebAssembly deliberately does not allow that. Locals are typed slots with no address, and the engine’s call stack is invisible to the module — a security property, since it stops code from reading return addresses or overwriting them. That leaves C, C++ and Rust with a problem: their semantics allow pointers to locals, arrays declared inside functions, and structs passed by address.

Compilers solve it by keeping a shadow stack in linear memory, which is addressable. On entry, a function that needs addressable storage subtracts its frame size from a global named __stack_pointer and uses the memory between the old and new values for its address-taken variables. On exit, it restores the pointer. Values that never have their address taken stay in WebAssembly locals, which the engine keeps in registers — so the shadow stack holds only what genuinely needs an address.

A function's frame on the shadow stack The shadow stack grows downward in linear memory. On entry a function subtracts its frame size from __stack_pointer and places an array and an address-taken struct in the new frame; on exit it restores the pointer. Scalars that are never addressed stay in Wasm locals and never touch memory. shadow stack region (grows downward ←) free stack space frame: char buf[256] Point p caller frames stack limit SP in callee SP on entry stack top

Step 1 — see it in compiled code

Compile a function with an address-taken local and look at the result:

// shadow.c
void fill(char *dst, int n);

int checksum_line(int n) {
  char buf[256];            // array: must have an address → shadow stack
  fill(buf, n);
  int sum = 0;              // scalar, never addressed → Wasm local
  for (int i = 0; i < n; i++) sum += buf[i];
  return sum;
}
clang --target=wasm32 -O2 -c shadow.c -o shadow.o && wasm2wat shadow.o | sed -n '/func $checksum_line/,/^  )/p' | head -14
(func $checksum_line (param i32) (result i32)
  (local i32 i32 i32)
  global.get $__stack_pointer      ;; load SP
  i32.const 256
  i32.sub                          ;; SP - 256
  local.tee 1
  global.set $__stack_pointer      ;; store new SP: frame allocated
  local.get 1                      ;; &buf
  local.get 0
  call $fill
  ...
  local.get 1
  i32.const 256
  i32.add
  global.set $__stack_pointer)     ;; restore SP: frame released

The prologue and epilogue are ordinary global reads and writes. Functions that do not need addressable storage have no such code at all.

Step 2 — know what lands on it

Typical residents of the shadow stack are local arrays and buffers, structs and objects whose address is passed to another function, variables captured by reference, va_list arguments for variadic functions, and large structs returned by value (through a hidden pointer to space in the caller’s frame). In Rust, the same applies to anything borrowed with &mut and passed out of the function, to large values the optimizer could not keep in locals, and to arrays.

What does not land there is equally important: scalar locals, loop counters, intermediate results and most small values stay in WebAssembly locals and become registers. A function that works mostly on scalars may never touch the shadow stack, and optimization levels change the picture considerably — -O0 builds put far more on the shadow stack than -O2 builds.

Step 3 — know how big it is

The shadow stack has a fixed size chosen at link time: 64 KiB by default for C with clang and for Emscripten, 1 MiB for Rust’s wasm32-unknown-unknown. It does not grow. You can read its bounds from the globals the linker emits:

wasm-objdump -x app.wasm | grep -E '__stack_pointer|__data_end|__heap_base'
 - global[0] i32 mutable=1 <__stack_pointer> - init i32=1114112
 - global[1] i32 mutable=0 <__data_end> - init i32=1048576
 - global[2] i32 mutable=0 <__heap_base> - init i32=1114112

Here the stack runs from 1,114,112 (the initial pointer, its top) down towards 1,048,576 — 64 KiB. Raising it is a linker flag, as shown in setting stack size and memory limits at link time.

Step 4 — recognise an overflow

Nothing checks the stack pointer against the stack’s lower bound. If recursion or a large local array pushes it below the bottom, the next frame is written into whatever lies below — static data by default — and the program continues with corrupted globals, string constants or allocator metadata. The symptoms appear later and elsewhere: a string literal that prints garbage, a lookup table with wrong entries, a crash inside malloc.

Two defences turn silent corruption into a clean failure. Linking with --stack-first puts the stack below static data, so an overflow runs past address zero and traps on the next access. Emscripten’s -sSTACK_OVERFLOW_CHECK=2 adds an explicit check in every prologue. The broader diagnosis is covered in fixing stack overflow in Wasm.

What a shadow-stack overflow does under each layout With the default layout the stack grows down into static data, so an overflow silently corrupts globals. With stack-first the stack sits at the bottom of memory, so an overflow wraps below address zero and traps immediately. default layout stack sits above static data overflow writes into globals program keeps running, corrupted silent corruption --stack-first stack sits at the bottom of memory overflow goes below address 0 next access traps with a stack trace loud, early failure

Step 5 — account for threads

In a threaded build every thread needs its own shadow stack, because each runs its own calls. The threading runtime allocates a stack for each new thread from the heap and sets that thread’s __stack_pointer — which is a per-thread global in threaded builds — to point at it. The main thread’s stack size comes from the linker flag; worker stacks come from a runtime setting such as Emscripten’s -sDEFAULT_PTHREAD_STACK_SIZE or the stack size passed when spawning a Rust thread. A program that runs fine single-threaded and crashes only in workers often has a worker stack that is too small.

Why this design is a good trade

The shadow stack can look like a workaround, and in a sense it is: it reconstructs, in memory the module controls, the addressable stack that native code takes for granted. But it keeps something valuable intact. Because return addresses and the engine’s frames live on a stack the module cannot address, a buffer overflow in WebAssembly cannot overwrite a return address and redirect control flow — the attack that made native stack overflows so dangerous for decades. The worst a shadow-stack overflow can do is corrupt the module’s own data. Combined with WebAssembly’s structured control flow, that removes an entire class of exploits by construction, at the cost of a few global accesses in functions that need addressable locals.

Expected output

Instrumenting the prologue — or reading __stack_pointer from an exported debug function — shows the pointer dropping by each function’s frame size on entry and returning to the same value on exit:

enter checksum_line: SP 1114112 → 1113856  (frame 256 bytes)
exit  checksum_line: SP 1113856 → 1114112

Gotchas

  • Large local arrays. A 1 MB array in a function exceeds a 64 KiB stack immediately. Allocate it on the heap instead.
  • Deep recursion on input. Recursive parsers overflow on deeply nested input. Cap the depth or convert to iteration.
  • Variadic functions in hot paths. Each call to a printf-style function builds its argument list on the shadow stack. Keep them out of inner loops.
  • Small worker stacks. Threads get a separate, often smaller, stack. Size it explicitly.
  • Inspecting __stack_pointer in release builds. It may be renamed or internalized. Export it explicitly in debug builds if you need it.

Performance note

Each shadow-stack frame costs a global read, a subtraction and a global write on entry and a write on exit — a handful of instructions. In a benchmark of a recursive function with an address-taken local, removing the address-taking (so the value stayed in a local) made it 9% faster; for functions that do real work, the prologue cost is negligible.

Cost of shadow-stack frames in a call-heavy benchmark Run time of a recursive benchmark whose function keeps a small struct in a Wasm local versus taking its address, which forces a shadow-stack frame in every call. ms for ten million calls value kept in Wasm locals 61 ms address taken → shadow frame 67 ms

Frequently Asked Questions

Do all WebAssembly modules have a shadow stack? No — only those compiled from languages that need addressable locals. Hand-written WAT, AssemblyScript in some configurations, and GC languages may not have one.

Can I read the shadow stack from JavaScript? It is ordinary linear memory, so yes, given its address. That is occasionally useful for debugging and never for anything else.

Does Rust’s safety prevent shadow-stack overflow? No. Safe Rust prevents memory unsafety in your code, but stack exhaustion is a resource limit, not a type error. Rust traps cleanly because its target uses --stack-first.

Can the shadow stack be moved or resized at runtime? Not safely in general — compiled code assumes the region between the stack’s limits belongs to it. Size it correctly at link time and keep large data on the heap instead.

Is the shadow stack the same as the Emscripten “stack”? Yes — Emscripten’s STACK_SIZE setting sizes exactly this region.

← Back to Stack vs Heap Execution Model