Avoiding Unnecessary Stack Allocation in Rust Wasm
This page answers one task: a Rust WebAssembly module traps with “memory access out of bounds” or corrupts data on certain inputs, and investigation shows the shadow stack overflowed — usually because of a large array, a big struct passed by value, or deep recursion. You want to keep large data on the heap, write code that does not silently copy big values through the stack, and size the stack sensibly.
Prerequisites
- [ ] A Rust crate compiled for
wasm32-unknown-unknownor WASI. - [ ] A reproduction that overflows, or a function suspected of large frames.
- [ ] Optionally, nightly Rust for stack-size reporting.
The stack is small and silent
Rust code compiled to Wasm uses a shadow stack in linear memory for anything that needs an address or does not fit in Wasm locals: arrays, larger structs,
values passed by reference. Rust’s wasm32-unknown-unknown target reserves 1 MiB for it by default — generous for typical code, small compared with native
thread stacks of 2–8 MiB. A function with let buf = [0u8; 512 * 1024]; uses half the stack in one frame; two nested calls like that, or a recursive function
with a modest array, overflow.
Overflow behaviour depends on memory layout. With the stack placed first (Rust’s default), running off the bottom wraps the stack pointer to a huge address and the next access traps. Without stack-first layout, the stack grows into static data and corrupts it silently. Either way, the fix is the same: keep large data off the stack.
Step 1 — move large buffers to the heap
Replace large local arrays with heap allocations:
// before: 256 KiB on the shadow stack
fn process(input: &[u8]) -> u32 {
let mut table = [0u32; 65536];
build(&mut table, input);
score(&table)
}
// after: heap-allocated, reused if called often
fn process(input: &[u8]) -> u32 {
let mut table = vec![0u32; 65536];
build(&mut table, input);
score(&table)
}
For functions called frequently, keep the buffer in a struct or a thread_local! and reuse it, so you pay the allocation once.
Step 2 — beware Box::new of large values
Box::new([0u8; 1 << 20]) looks like a heap allocation, but in unoptimised builds (and sometimes optimised ones) the array is first constructed on the stack and
then copied into the box — overflowing the stack before the heap is involved. Allocate directly on the heap instead:
let big: Box<[u8]> = vec![0u8; 1 << 20].into_boxed_slice(); // no stack temporary
let state: Box<BigState> = Box::new_uninit().write(BigState::new()); // where BigState::new is cheap; or build fields in place
For large structs, construct them field by field behind a pointer, or design them to hold their large parts in Vecs from the start, so the struct itself stays
small.
Step 3 — avoid large by-value moves and returns
Passing or returning large structs by value can create temporaries in frames — for the return slot, for the argument copy — even when the optimiser removes some
of them. Pass large data by reference (&BigState, &mut BigState) and return large results in a caller-provided buffer or a Box. Iterator chains that
collect into arrays ([T; N]) also build on the stack; collect into Vec instead.
Step 4 — tame recursion
Recursive algorithms on input-controlled depth (parsers for nested structures, tree traversals) can overflow even with small frames. Convert deep recursion to
iteration with an explicit stack in a Vec (heap-allocated, growable), or enforce a maximum nesting depth and return an error beyond it. Explicit stacks also make
it possible to process inputs of arbitrary depth safely, which matters for untrusted input.
Step 5 — measure, then size the stack
Find the largest frames: on nightly, -Z emit-stack-sizes produces a section a tool can read (for example cargo-call-stack analyses call graphs and stack
usage); otherwise read the prologue constants in WAT for suspect functions. If after moving data to the heap the program still needs more stack — legitimately
deep but bounded recursion — raise the stack size at link time:
# .cargo/config.toml
[target.wasm32-unknown-unknown]
rustflags = ["-C", "link-arg=-zstack-size=4194304"] # 4 MiB
Every instance reserves the stack in linear memory, and threads each get their own, so prefer moving data off the stack over raising the size.
Debug builds use far more stack
Unoptimised builds give every local its own frame slot and keep temporaries the optimiser would remove, so debug builds can need several times more stack than release builds. A test suite that overflows only in debug mode is common; it is usually a sign that some function is close to the limit in release too, and worth fixing rather than masking by raising the stack only for debug builds.
Async code and large futures
Async Rust has its own version of the problem. An async fn compiles into a state machine whose size includes every local that lives across an .await, so a
large array held across an await point makes the future itself large. Futures are moved when spawned and sometimes when polled through combinators, and if they
live on the stack at those moments, a multi-hundred-kilobyte future can overflow it. With wasm-bindgen-futures, spawn_local boxes the future, which puts it on
the heap, but constructing it before boxing still happens on the stack. Keep large data in heap allocations inside async functions (a Vec rather than an
array), avoid holding big buffers across .await, and use Box::pin for large sub-futures you compose. Clippy’s large_futures lint warns about futures above a
configurable size, which catches most of these cases at compile time.
Lints and checks that catch it early
Several lints help: Clippy’s large_stack_arrays flags large local arrays, large_types_passed_by_value flags big structs passed by value, large_futures
flags oversized futures, and large_stack_frames (where available) flags functions whose frames exceed a threshold. Configure their thresholds in clippy.toml
to suit a 1 MiB stack — for example warning at 16 KiB rather than the native-oriented defaults — and run Clippy for the Wasm target in CI. A test that runs the
module’s deepest realistic workload in the Wasm build, rather than natively, confirms there is headroom; native tests have larger stacks and miss these overflows.
Interaction with the heap
Moving data to the heap trades stack overflow risk for allocation cost and memory growth. For hot paths, reuse buffers; for large one-off structures, allocate once and drop promptly, since Wasm memory grows but never shrinks.
Expected output
The parser’s 256 KiB lookup table moves to a reused Vec; a Box::new of a 1 MiB state becomes a boxed slice allocation; recursive descent over nested input is
replaced with an explicit stack and a depth limit; the largest remaining frame is 2 KiB; deep-input tests pass in debug and release builds with the default
1 MiB stack.
Gotchas
- Large local arrays. Overflow quickly at 1 MiB. Use
vec!. Box::newof huge values. Built on the stack first. Allocate on the heap directly.- Unbounded recursion on input. Attackers control depth. Iterate or limit depth.
- Raising the stack instead of fixing frames. Costs memory per instance and thread. Prefer the heap.
- Ignoring debug-only overflows. They signal tight release margins. Investigate.
- Native-only tests. Native stacks are larger. Run the deepest workloads in the Wasm build.
Performance note
Reusing a heap buffer across calls instead of a fresh 256 KiB stack array made the function slightly faster (no 256 KiB zeroing per call) and removed the overflow;
allocating a new Vec per call cost about 2% compared with the stack version.
Frequently Asked Questions
What is Rust’s default Wasm stack size?
1 MiB for wasm32-unknown-unknown; check your toolchain version and target.
Does wasm-bindgen change the stack size?
No — it is set at link time; wasm-bindgen uses the module’s stack.
Can I detect overflow instead of trapping? Stack-first layout makes overflow trap reliably; checking depth in recursive code prevents it.
Do threads share the stack? No — each thread has its own stack of the configured size, allocated by the runtime.
Can async functions overflow the stack? Yes — futures holding large locals across awaits are large; keep big data in heap allocations and box large sub-futures.
Which Clippy lints help?
large_stack_arrays, large_types_passed_by_value, large_futures and, where available, large_stack_frames, with thresholds tuned for a 1 MiB stack.
Does moving data to the heap cost performance? A little per allocation; reuse buffers in hot paths to avoid it.
Should I set Clippy thresholds lower for Wasm? Yes — native defaults assume megabytes of stack; warn at a few kilobytes to suit Wasm’s 1 MiB default.
Related
- Fixing stack overflow in Wasm — diagnosing overflows.
- Reading a compiled function’s frame in WAT — frame sizes.
- Laying out a Wasm module’s memory map — stack placement.
- Understanding the shadow stack in linear memory — the mechanism.
← Back to Stack vs Heap Execution Model