Fixing Stack Overflow in Wasm

This page answers one task: a WebAssembly module crashes with a stack-related error, or corrupts its own data under deep input — work out which stack ran out and fix it.

Prerequisites

  • [ ] The failing module and a way to reproduce the crash.
  • [ ] The ability to relink with different flags (wasm-ld, rustc link args, or emcc settings).
  • [ ] A debug or names-preserving build, so stack traces are readable.

Two stacks, two different overflows

WebAssembly code compiled from C, C++ or Rust uses two stacks that can each run out, and they fail in very different ways.

The engine’s call stack holds the native frames for every active WebAssembly function call. Its depth is limited by the engine — V8, SpiderMonkey and JavaScriptCore each allow a certain amount of native stack per thread — and exceeding it throws a RangeError: Maximum call stack size exceeded in browsers, or a trap reporting call stack exhaustion in standalone runtimes. This limit is about the number and size of native frames: deep recursion hits it.

The shadow stack lives in linear memory and holds the module’s address-taken locals, arrays and large structs, as described in understanding the shadow stack in linear memory. Its size is fixed at link time, and nothing checks it by default. Exceeding it either corrupts memory silently or, with a protective layout, traps with an out-of-bounds memory access. This limit is about bytes of addressable local data: large local arrays hit it, and so does moderately deep recursion with big frames.

The two stack overflows side by side Engine call-stack exhaustion comes from too many nested calls and surfaces as RangeError or a call-stack trap. Shadow-stack overflow comes from too many bytes of address-taken locals and surfaces as silent corruption or an out-of-bounds trap. engine call stack limited by native frames per thread caused by deep recursion RangeError: Maximum call stack size exceeded fix by reducing depth shadow stack (linear memory) fixed size chosen at link time caused by big local arrays and deep frames silent corruption or out-of-bounds trap fix by size, layout, heap allocation

Step 1 — identify which one you hit

Read the error, then the stack trace. An engine overflow names itself:

RangeError: Maximum call stack size exceeded
    at parse_value (wasm://wasm/0002a9be:wasm-function[37]:0x1c4f)
    at parse_value (wasm://wasm/0002a9be:wasm-function[37]:0x1d02)
    ... (thousands of identical frames)

A shadow-stack overflow with --stack-first looks like an ordinary out-of-bounds access, often in a function prologue or at the first store into a local array:

RuntimeError: memory access out of bounds
    at render_tile (wasm://wasm/7c11f0a2:wasm-function[212]:0x9a31)
    at render_row (wasm://wasm/7c11f0a2:wasm-function[211]:0x98f0)

And a shadow-stack overflow without protection produces no error at the overflow at all — just wrong results or a crash later, frequently in malloc or when printing a string constant. If the symptom changes with input depth or with the size of a local buffer, suspect the shadow stack even without an error.

Step 2 — fix deep recursion

Recursion that follows input structure — nested JSON, an expression tree, a directory hierarchy — can be driven arbitrarily deep by the input. Raising limits only moves the failure. Make the depth bounded, or make the algorithm iterative:

// before: recursion depth equals nesting depth of the input
fn depth_sum(v: &Value) -> i64 {
    match v { Value::Array(xs) => xs.iter().map(depth_sum).sum(), Value::Int(n) => *n }
}

// after: explicit stack on the heap; depth limited only by memory
fn depth_sum(root: &Value) -> i64 {
    let mut stack = vec![root];
    let mut total = 0;
    while let Some(v) = stack.pop() {
        match v { Value::Array(xs) => stack.extend(xs.iter()), Value::Int(n) => total += n }
    }
    total
}

Where recursion is the clearest expression of the algorithm, cap it instead: track depth and return an error past a sensible limit, such as a few hundred levels for a document format. That turns a crash on hostile input into a validation error, which is what fuzzing a Wasm module will otherwise find for you. Tail calls, where the language supports them, remove the frame growth for one important pattern; see tail calls and deep recursion.

Step 3 — move large frames to the heap

A local array of a megabyte overflows a 64 KiB shadow stack on the first call, no recursion needed. Large or variable-size buffers belong on the heap:

// before: 1 MB on the shadow stack
void render_tile(Tile *t) {
  uint32_t scratch[256 * 1024];
  ...
}

// after: allocated once, reused
static uint32_t *scratch;
void render_tile(Tile *t) {
  if (!scratch) scratch = malloc(256 * 1024 * sizeof *scratch);
  ...
}

In Rust, a large array declared as a local — let buf = [0u8; 1 << 20]; — lands on the shadow stack too; use vec![0u8; 1 << 20] or a Box<[u8]> instead. Keep stack frames to kilobytes, not megabytes.

Step 4 — size the shadow stack and protect it

After removing unbounded recursion and large frames, set the shadow stack to what the program needs plus a margin, and use a layout where overflow traps:

# C / C++ with clang
clang --target=wasm32 -O2 ... -Wl,-z,stack-size=1048576 -Wl,--stack-first

# Emscripten
emcc -O2 ... -sSTACK_SIZE=1MB -sSTACK_OVERFLOW_CHECK=1
# Rust (.cargo/config.toml) — Rust already uses --stack-first on wasm32-unknown-unknown
[target.wasm32-unknown-unknown]
rustflags = ["-C", "link-arg=-zstack-size=2097152"]

Measure the high-water mark under your deepest realistic workload rather than guessing; the method is described in setting stack size and memory limits at link time.

Choosing the fix for a stack overflow A decision tree. If the engine reports call-stack exhaustion, reduce recursion depth. If the shadow stack overflows because of large locals, move them to the heap. If it overflows from legitimate depth, raise the shadow stack size and use stack-first so future overflows trap. Which stack overflowed? engine (RangeError) Reduce call depth iterate, cap depth, tail calls shadow, large locals Move buffers to heap Vec, Box, malloc shadow, legitimate depth Raise size, --stack-first measure the high-water mark

Step 5 — check worker threads separately

Threads get their own shadow stacks, allocated by the threading runtime, and often smaller ones than the main thread. A program that passes every test on the main thread can overflow in a worker. Set the worker stack size explicitly — -sDEFAULT_PTHREAD_STACK_SIZE=1MB in Emscripten, std::thread::Builder::new().stack_size(…) in Rust — and run the deepest test cases in a worker too. Engine stack limits also differ between the main thread and workers in some browsers, which matters for deeply recursive code moved into a worker for responsiveness.

Why the limits exist at all

It is reasonable to ask why engines do not simply let the stack grow. For the engine’s own stack, the answer is safety and predictability: native stacks are fixed reservations per thread, and a WebAssembly function that could consume unlimited native stack could crash the browser’s process rather than just the page’s script. The engine therefore checks the remaining space on entry to each function and throws before it runs out — a deliberate, catchable failure. For the shadow stack, the limit is simply the region the linker reserved; it could be larger, but every byte reserved is a byte of linear memory committed in every instance and every thread. Both limits are reasonable defaults, and the right response to hitting them is usually to change the code’s shape rather than to raise them indefinitely.

Keep the deepest test inputs in the regular test suite once the fix is in, so a later refactor that reintroduces recursion is caught.

Expected output

After converting the recursive parser to an explicit stack and capping nesting, deeply nested input returns a clean error rather than crashing:

parse error: nesting deeper than 512 levels

and the render path, with its scratch buffer moved to the heap, runs with the default 64 KiB shadow stack.

Gotchas

  • Raising the engine limit. Browsers do not let pages change their native stack size. Reduce depth instead.
  • Raising the shadow stack to hide a large local. Works until someone calls the function from a worker with a smaller stack. Move the buffer.
  • Catching RangeError and continuing. After an engine overflow, the module may be in an inconsistent state — locks held, data half updated. Re-instantiate, as in recovering a module after a trap.
  • Recursion hidden behind callbacks. A visitor pattern or a recursive closure is still recursion. Count depth there too.
  • Debug builds overflow, release builds do not. Unoptimized code uses much larger frames. Test the stack limits with optimized builds.

Performance note

The iterative parser was 11% faster than the recursive one on typical input — fewer function calls and better locality — and handled nesting depths of 100,000 without trouble. Moving the 1 MB scratch buffer to the heap had no measurable cost because it was allocated once.

Maximum nesting depth handled before failure The deepest nested input the parser handled in Chrome with the recursive implementation and default stacks, with a larger shadow stack, and with the iterative rewrite. maximum nesting depth handled recursive, default stacks 2,900 levels recursive, 4 MB shadow stack 9,800 levels iterative with heap stack 100,000 levels The recursive version's limit with a larger shadow stack came from the engine's own call-depth limit, which a page cannot raise.

Frequently Asked Questions

Why does the same code overflow in Safari and not in Chrome? Engines reserve different amounts of native stack and use different frame sizes. Design for the smallest limit you support, or remove the deep recursion.

How do I know how deep my recursion actually goes? Count it: pass a depth parameter through the recursive calls and record the maximum seen in tests and in production telemetry. The number tells you how close real inputs come to the limit, and whether a cap would ever affect legitimate data.

Do WASI runtimes have the same limits? wasmtime and others have configurable native stack limits (--max-wasm-stack) and the same shadow-stack behaviour, since that is part of the compiled module.

Can a shadow-stack overflow be exploited? It cannot overwrite return addresses, which live on the engine’s stack. It can corrupt the module’s own data, which for multi-tenant hosts is still a security concern.

Does Rust detect stack overflow? On wasm32-unknown-unknown the --stack-first layout makes overflow trap. Rust’s guard-page mechanism used natively does not exist in Wasm.

← Back to Stack vs Heap Execution Model