Avoiding Unnecessary Stack Allocation in Rust Wasm

This page answers one task: a Rust WebAssembly module traps with “memory access out of bounds” or corrupts data on certain inputs, and investigation shows the shadow stack overflowed — usually because of a large array, a big struct passed by value, or deep recursion. You want to keep large data on the heap, write code that does not silently copy big values through the stack, and size the stack sensibly.

Prerequisites

  • [ ] A Rust crate compiled for wasm32-unknown-unknown or WASI.
  • [ ] A reproduction that overflows, or a function suspected of large frames.
  • [ ] Optionally, nightly Rust for stack-size reporting.

The stack is small and silent

Rust code compiled to Wasm uses a shadow stack in linear memory for anything that needs an address or does not fit in Wasm locals: arrays, larger structs, values passed by reference. Rust’s wasm32-unknown-unknown target reserves 1 MiB for it by default — generous for typical code, small compared with native thread stacks of 2–8 MiB. A function with let buf = [0u8; 512 * 1024]; uses half the stack in one frame; two nested calls like that, or a recursive function with a modest array, overflow.

Overflow behaviour depends on memory layout. With the stack placed first (Rust’s default), running off the bottom wraps the stack pointer to a huge address and the next access traps. Without stack-first layout, the stack grows into static data and corrupts it silently. Either way, the fix is the same: keep large data off the stack.

Large data on the stack versus on the heap A large array declared as a local lives in the function's frame on the 1 MiB shadow stack, so a few such frames or recursion overflow it. Allocating the same buffer with vec! or a boxed slice puts it on the heap, leaving only a pointer and length on the stack, so stack usage stays small regardless of buffer size. let buf = [0u8; 512K] 512 KiB in one frame half the default stack overflow with recursion avoid in Wasm let buf = vec![0u8; 512K] heap allocation 12 bytes on the stack grows memory as needed preferred

Step 1 — move large buffers to the heap

Replace large local arrays with heap allocations:

// before: 256 KiB on the shadow stack
fn process(input: &[u8]) -> u32 {
    let mut table = [0u32; 65536];
    build(&mut table, input);
    score(&table)
}

// after: heap-allocated, reused if called often
fn process(input: &[u8]) -> u32 {
    let mut table = vec![0u32; 65536];
    build(&mut table, input);
    score(&table)
}

For functions called frequently, keep the buffer in a struct or a thread_local! and reuse it, so you pay the allocation once.

Step 2 — beware Box::new of large values

Box::new([0u8; 1 << 20]) looks like a heap allocation, but in unoptimised builds (and sometimes optimised ones) the array is first constructed on the stack and then copied into the box — overflowing the stack before the heap is involved. Allocate directly on the heap instead:

let big: Box<[u8]> = vec![0u8; 1 << 20].into_boxed_slice();          // no stack temporary
let state: Box<BigState> = Box::new_uninit().write(BigState::new());  // where BigState::new is cheap; or build fields in place

For large structs, construct them field by field behind a pointer, or design them to hold their large parts in Vecs from the start, so the struct itself stays small.

Step 3 — avoid large by-value moves and returns

Passing or returning large structs by value can create temporaries in frames — for the return slot, for the argument copy — even when the optimiser removes some of them. Pass large data by reference (&BigState, &mut BigState) and return large results in a caller-provided buffer or a Box. Iterator chains that collect into arrays ([T; N]) also build on the stack; collect into Vec instead.

Finding a stack overflow in Rust Wasm A trap or corruption appears on some inputs. Building with stack-size reporting or reading the WAT prologues shows a function with a very large frame. The source reveals a large local array or a large by-value struct. Moving it to the heap or passing it by reference shrinks the frame, and a test with deep inputs confirms the fix. trap on large inputs or corrupted data check frame sizes stack sizes / WAT find big local array or struct move to heap vec!, Box<[T]>, &mut deep-input test no overflow

Step 4 — tame recursion

Recursive algorithms on input-controlled depth (parsers for nested structures, tree traversals) can overflow even with small frames. Convert deep recursion to iteration with an explicit stack in a Vec (heap-allocated, growable), or enforce a maximum nesting depth and return an error beyond it. Explicit stacks also make it possible to process inputs of arbitrary depth safely, which matters for untrusted input.

Step 5 — measure, then size the stack

Find the largest frames: on nightly, -Z emit-stack-sizes produces a section a tool can read (for example cargo-call-stack analyses call graphs and stack usage); otherwise read the prologue constants in WAT for suspect functions. If after moving data to the heap the program still needs more stack — legitimately deep but bounded recursion — raise the stack size at link time:

# .cargo/config.toml
[target.wasm32-unknown-unknown]
rustflags = ["-C", "link-arg=-zstack-size=4194304"]   # 4 MiB

Every instance reserves the stack in linear memory, and threads each get their own, so prefer moving data off the stack over raising the size.

Debug builds use far more stack

Unoptimised builds give every local its own frame slot and keep temporaries the optimiser would remove, so debug builds can need several times more stack than release builds. A test suite that overflows only in debug mode is common; it is usually a sign that some function is close to the limit in release too, and worth fixing rather than masking by raising the stack only for debug builds.

Async code and large futures

Async Rust has its own version of the problem. An async fn compiles into a state machine whose size includes every local that lives across an .await, so a large array held across an await point makes the future itself large. Futures are moved when spawned and sometimes when polled through combinators, and if they live on the stack at those moments, a multi-hundred-kilobyte future can overflow it. With wasm-bindgen-futures, spawn_local boxes the future, which puts it on the heap, but constructing it before boxing still happens on the stack. Keep large data in heap allocations inside async functions (a Vec rather than an array), avoid holding big buffers across .await, and use Box::pin for large sub-futures you compose. Clippy’s large_futures lint warns about futures above a configurable size, which catches most of these cases at compile time.

Lints and checks that catch it early

Several lints help: Clippy’s large_stack_arrays flags large local arrays, large_types_passed_by_value flags big structs passed by value, large_futures flags oversized futures, and large_stack_frames (where available) flags functions whose frames exceed a threshold. Configure their thresholds in clippy.toml to suit a 1 MiB stack — for example warning at 16 KiB rather than the native-oriented defaults — and run Clippy for the Wasm target in CI. A test that runs the module’s deepest realistic workload in the Wasm build, rather than natively, confirms there is headroom; native tests have larger stacks and miss these overflows.

Interaction with the heap

Moving data to the heap trades stack overflow risk for allocation cost and memory growth. For hot paths, reuse buffers; for large one-off structures, allocate once and drop promptly, since Wasm memory grows but never shrinks.

Expected output

The parser’s 256 KiB lookup table moves to a reused Vec; a Box::new of a 1 MiB state becomes a boxed slice allocation; recursive descent over nested input is replaced with an explicit stack and a depth limit; the largest remaining frame is 2 KiB; deep-input tests pass in debug and release builds with the default 1 MiB stack.

Gotchas

  • Large local arrays. Overflow quickly at 1 MiB. Use vec!.
  • Box::new of huge values. Built on the stack first. Allocate on the heap directly.
  • Unbounded recursion on input. Attackers control depth. Iterate or limit depth.
  • Raising the stack instead of fixing frames. Costs memory per instance and thread. Prefer the heap.
  • Ignoring debug-only overflows. They signal tight release margins. Investigate.
  • Native-only tests. Native stacks are larger. Run the deepest workloads in the Wasm build.

Performance note

Reusing a heap buffer across calls instead of a fresh 256 KiB stack array made the function slightly faster (no 256 KiB zeroing per call) and removed the overflow; allocating a new Vec per call cost about 2% compared with the stack version.

Time per call for a function with a 256 KiB table Relative time per call for a function using a 256 KiB local array on the stack, a new Vec allocated per call, and a heap buffer reused across calls without re-zeroing. relative time per call stack array (overflow risk) 1 × new Vec per call 1.0 × reused heap buffer 0.8 ×

Frequently Asked Questions

What is Rust’s default Wasm stack size? 1 MiB for wasm32-unknown-unknown; check your toolchain version and target.

Does wasm-bindgen change the stack size? No — it is set at link time; wasm-bindgen uses the module’s stack.

Can I detect overflow instead of trapping? Stack-first layout makes overflow trap reliably; checking depth in recursive code prevents it.

Do threads share the stack? No — each thread has its own stack of the configured size, allocated by the runtime.

Can async functions overflow the stack? Yes — futures holding large locals across awaits are large; keep big data in heap allocations and box large sub-futures.

Which Clippy lints help? large_stack_arrays, large_types_passed_by_value, large_futures and, where available, large_stack_frames, with thresholds tuned for a 1 MiB stack.

Does moving data to the heap cost performance? A little per allocation; reuse buffers in hot paths to avoid it.

Should I set Clippy thresholds lower for Wasm? Yes — native defaults assume megabytes of stack; warn at a few kilobytes to suit Wasm’s 1 MiB default.

← Back to Stack vs Heap Execution Model