Fixing Stack Overflow in Wasm
This page answers one task: a WebAssembly module crashes with a stack-related error, or corrupts its own data under deep input — work out which stack ran out and fix it.
Prerequisites
- [ ] The failing module and a way to reproduce the crash.
- [ ] The ability to relink with different flags (
wasm-ld, rustc link args, or emcc settings). - [ ] A debug or names-preserving build, so stack traces are readable.
Two stacks, two different overflows
WebAssembly code compiled from C, C++ or Rust uses two stacks that can each run out, and they fail in very different ways.
The engine’s call stack holds the native frames for every active WebAssembly function call. Its depth is limited by the engine — V8,
SpiderMonkey and JavaScriptCore each allow a certain amount of native stack per thread — and exceeding it throws a RangeError: Maximum call stack size exceeded in browsers, or a trap reporting call stack exhaustion in standalone runtimes. This limit is about the number and
size of native frames: deep recursion hits it.
The shadow stack lives in linear memory and holds the module’s address-taken locals, arrays and large structs, as described in understanding the shadow stack in linear memory. Its size is fixed at link time, and nothing checks it by default. Exceeding it either corrupts memory silently or, with a protective layout, traps with an out-of-bounds memory access. This limit is about bytes of addressable local data: large local arrays hit it, and so does moderately deep recursion with big frames.
Step 1 — identify which one you hit
Read the error, then the stack trace. An engine overflow names itself:
RangeError: Maximum call stack size exceeded
at parse_value (wasm://wasm/0002a9be:wasm-function[37]:0x1c4f)
at parse_value (wasm://wasm/0002a9be:wasm-function[37]:0x1d02)
... (thousands of identical frames)
A shadow-stack overflow with --stack-first looks like an ordinary out-of-bounds access, often in a function prologue or at the first
store into a local array:
RuntimeError: memory access out of bounds
at render_tile (wasm://wasm/7c11f0a2:wasm-function[212]:0x9a31)
at render_row (wasm://wasm/7c11f0a2:wasm-function[211]:0x98f0)
And a shadow-stack overflow without protection produces no error at the overflow at all — just wrong results or a crash later, frequently in
malloc or when printing a string constant. If the symptom changes with input depth or with the size of a local buffer, suspect the shadow
stack even without an error.
Step 2 — fix deep recursion
Recursion that follows input structure — nested JSON, an expression tree, a directory hierarchy — can be driven arbitrarily deep by the input. Raising limits only moves the failure. Make the depth bounded, or make the algorithm iterative:
// before: recursion depth equals nesting depth of the input
fn depth_sum(v: &Value) -> i64 {
match v { Value::Array(xs) => xs.iter().map(depth_sum).sum(), Value::Int(n) => *n }
}
// after: explicit stack on the heap; depth limited only by memory
fn depth_sum(root: &Value) -> i64 {
let mut stack = vec![root];
let mut total = 0;
while let Some(v) = stack.pop() {
match v { Value::Array(xs) => stack.extend(xs.iter()), Value::Int(n) => total += n }
}
total
}
Where recursion is the clearest expression of the algorithm, cap it instead: track depth and return an error past a sensible limit, such as a few hundred levels for a document format. That turns a crash on hostile input into a validation error, which is what fuzzing a Wasm module will otherwise find for you. Tail calls, where the language supports them, remove the frame growth for one important pattern; see tail calls and deep recursion.
Step 3 — move large frames to the heap
A local array of a megabyte overflows a 64 KiB shadow stack on the first call, no recursion needed. Large or variable-size buffers belong on the heap:
// before: 1 MB on the shadow stack
void render_tile(Tile *t) {
uint32_t scratch[256 * 1024];
...
}
// after: allocated once, reused
static uint32_t *scratch;
void render_tile(Tile *t) {
if (!scratch) scratch = malloc(256 * 1024 * sizeof *scratch);
...
}
In Rust, a large array declared as a local — let buf = [0u8; 1 << 20]; — lands on the shadow stack too; use vec![0u8; 1 << 20] or a
Box<[u8]> instead. Keep stack frames to kilobytes, not megabytes.
Step 4 — size the shadow stack and protect it
After removing unbounded recursion and large frames, set the shadow stack to what the program needs plus a margin, and use a layout where overflow traps:
# C / C++ with clang
clang --target=wasm32 -O2 ... -Wl,-z,stack-size=1048576 -Wl,--stack-first
# Emscripten
emcc -O2 ... -sSTACK_SIZE=1MB -sSTACK_OVERFLOW_CHECK=1
# Rust (.cargo/config.toml) — Rust already uses --stack-first on wasm32-unknown-unknown
[target.wasm32-unknown-unknown]
rustflags = ["-C", "link-arg=-zstack-size=2097152"]
Measure the high-water mark under your deepest realistic workload rather than guessing; the method is described in setting stack size and memory limits at link time.
Step 5 — check worker threads separately
Threads get their own shadow stacks, allocated by the threading runtime, and often smaller ones than the main thread. A program that passes
every test on the main thread can overflow in a worker. Set the worker stack size explicitly — -sDEFAULT_PTHREAD_STACK_SIZE=1MB in
Emscripten, std::thread::Builder::new().stack_size(…) in Rust — and run the deepest test cases in a worker too. Engine stack limits also
differ between the main thread and workers in some browsers, which matters for deeply recursive code moved into a worker for responsiveness.
Why the limits exist at all
It is reasonable to ask why engines do not simply let the stack grow. For the engine’s own stack, the answer is safety and predictability: native stacks are fixed reservations per thread, and a WebAssembly function that could consume unlimited native stack could crash the browser’s process rather than just the page’s script. The engine therefore checks the remaining space on entry to each function and throws before it runs out — a deliberate, catchable failure. For the shadow stack, the limit is simply the region the linker reserved; it could be larger, but every byte reserved is a byte of linear memory committed in every instance and every thread. Both limits are reasonable defaults, and the right response to hitting them is usually to change the code’s shape rather than to raise them indefinitely.
Keep the deepest test inputs in the regular test suite once the fix is in, so a later refactor that reintroduces recursion is caught.
Expected output
After converting the recursive parser to an explicit stack and capping nesting, deeply nested input returns a clean error rather than crashing:
parse error: nesting deeper than 512 levels
and the render path, with its scratch buffer moved to the heap, runs with the default 64 KiB shadow stack.
Gotchas
- Raising the engine limit. Browsers do not let pages change their native stack size. Reduce depth instead.
- Raising the shadow stack to hide a large local. Works until someone calls the function from a worker with a smaller stack. Move the buffer.
- Catching
RangeErrorand continuing. After an engine overflow, the module may be in an inconsistent state — locks held, data half updated. Re-instantiate, as in recovering a module after a trap. - Recursion hidden behind callbacks. A visitor pattern or a recursive closure is still recursion. Count depth there too.
- Debug builds overflow, release builds do not. Unoptimized code uses much larger frames. Test the stack limits with optimized builds.
Performance note
The iterative parser was 11% faster than the recursive one on typical input — fewer function calls and better locality — and handled nesting depths of 100,000 without trouble. Moving the 1 MB scratch buffer to the heap had no measurable cost because it was allocated once.
Frequently Asked Questions
Why does the same code overflow in Safari and not in Chrome? Engines reserve different amounts of native stack and use different frame sizes. Design for the smallest limit you support, or remove the deep recursion.
How do I know how deep my recursion actually goes? Count it: pass a depth parameter through the recursive calls and record the maximum seen in tests and in production telemetry. The number tells you how close real inputs come to the limit, and whether a cap would ever affect legitimate data.
Do WASI runtimes have the same limits?
wasmtime and others have configurable native stack limits (--max-wasm-stack) and the same shadow-stack behaviour, since that is part of the
compiled module.
Can a shadow-stack overflow be exploited? It cannot overwrite return addresses, which live on the engine’s stack. It can corrupt the module’s own data, which for multi-tenant hosts is still a security concern.
Does Rust detect stack overflow?
On wasm32-unknown-unknown the --stack-first layout makes overflow trap. Rust’s guard-page mechanism used natively does not exist in Wasm.
Related
- Understanding Wasm linear memory limits — the memory the shadow stack lives in.
- Reading Wasm stack traces — interpreting the traces above.
- Catching Wasm traps in JavaScript — handling the failure at the boundary.
- Why Wasm cannot take the address of a local — why the second stack exists.
← Back to Stack vs Heap Execution Model