Setting Stack Size and Memory Limits at Link Time

This page answers one question: how does wasm-ld decide where static data, the stack and the heap live in linear memory and how big each may be — and which flags fix the crashes and failures that come from getting those sizes wrong?

Prerequisites

  • [ ] A C, C++ or Rust module linked with wasm-ld (directly or through clang, rustc or emcc).
  • [ ] wasm-objdump to read the memory section and globals.
  • [ ] A symptom, ideally: a stack overflow, an out-of-memory error, or corrupted static data.

The layout wasm-ld produces

Compiled C, C++ and Rust need three regions in linear memory. Static data holds string literals, initialised globals and zero-initialised statics. The shadow stack holds stack frames for anything that needs an address — arrays, structs passed by pointer, variables whose address is taken — because WebAssembly’s own operand stack cannot be addressed; the background is in understanding the shadow stack in linear memory. And the heap is everything above, managed by malloc.

wasm-ld lays these out at link time. By default static data starts at --global-base (1024), the stack comes next with a fixed size and grows downward towards the data, and the heap starts at the symbol __heap_base, after the stack. A global named __stack_pointer holds the current stack top. Nothing at runtime checks whether the stack pointer has moved past the bottom of the stack region — if it does, the next frame silently overwrites static data.

Default linear memory layout from wasm-ld Address zero to 1024 is left unused, static data follows at global-base, then the shadow stack of 64 KiB growing downward towards the data, then the heap starting at __heap_base and growing upward to the end of memory. linear memory, default layout (not to scale) unused static data ← shadow stack (64 KiB) heap (malloc) → 0 1024 __data_end __heap_base memory size

Step 1 — read the current layout

The linker exports the boundaries as globals in some configurations, and they are always visible in the name section and the global section:

wasm-objdump -x module.wasm | grep -A12 '^Global\['
wasm-objdump -x -j Memory module.wasm
Global[3]:
 - global[0] i32 mutable=1 <__stack_pointer> - init i32=1101824
 - global[1] i32 mutable=0 <__data_end> - init i32=1036288
 - global[2] i32 mutable=0 <__heap_base> - init i32=1101824
Memory[1]:
 - memory[0] pages: initial=17 max=16384

Here static data ends at about 1,012 KB, the stack occupies the 64 KiB above it (the initial stack pointer is the top), and the heap starts immediately after.

Step 2 — raise the stack size for deep recursion

The default stack is 64 KiB for most toolchains — Rust uses 1 MiB for wasm32-unknown-unknown, Emscripten 64 KiB in recent releases. Recursive parsers, large local arrays and deep call chains overflow it. With C and clang:

clang --target=wasm32 -O2 parser.c -o parser.wasm -Wl,--no-entry -Wl,-z,stack-size=1048576

From Rust:

# .cargo/config.toml
[target.wasm32-unknown-unknown]
rustflags = ["-C", "link-arg=-zstack-size=2097152"]

With Emscripten, use -sSTACK_SIZE=1MB. The stack is carved out of linear memory, so a larger stack raises the module’s minimum memory by the same amount.

Step 3 — put the stack first so overflows trap

Because the stack grows downward towards static data, an overflow corrupts data silently. --stack-first reorders the layout so the stack sits at the bottom of memory, below the data. An overflow then runs the stack pointer below address zero, which wraps to a huge unsigned address beyond the end of memory, and the next access traps instead of corrupting anything:

clang --target=wasm32 -O2 parser.c -o parser.wasm \
  -Wl,--no-entry -Wl,-z,stack-size=1048576 -Wl,--stack-first
Layout with --stack-first With stack-first, the shadow stack occupies the lowest addresses and grows downward towards zero, static data follows above it, and the heap comes last. A stack overflow runs below address zero and traps instead of overwriting data. linear memory with --stack-first (not to scale) ← shadow stack (1 MiB) static data heap (malloc) → 0 (overflow traps) stack top __heap_base memory size

Rust’s wasm32-unknown-unknown target already uses --stack-first; for C builds it is opt-in and almost always worth enabling. Emscripten offers -sSTACK_OVERFLOW_CHECK=2 for explicit checking instead.

Step 4 — set initial and maximum memory

--initial-memory sets how much memory the module starts with; it must cover static data plus the stack, and the linker rounds up to whole 64 KiB pages. --max-memory sets the ceiling memory.grow can reach. Without a maximum, the module may grow to the engine’s limit — 4 GiB for 32-bit memories in most browsers — which is usually more than a page should ever use.

clang --target=wasm32 -O2 app.c -o app.wasm -Wl,--no-entry \
  -Wl,--initial-memory=4194304 \
  -Wl,--max-memory=134217728

A sensible initial size avoids repeated growth during startup, each of which costs a reallocation; a sensible maximum turns a runaway allocation into a clean out-of-memory error instead of the tab being killed. How to choose the numbers is covered in sizing initial and maximum memory.

Layout flags, what they change, and the symptom when wrong The wasm-ld options that control memory layout, their defaults, and the failure each produces when set badly. flag default symptom when wrong -z stack-size 64 KiB (C), 1 MiB (Rust) corrupted statics, odd crashes --stack-first off (C), on (Rust) overflows corrupt instead of trap --initial-memory data + stack slow startup from many grows --max-memory engine limit tab killed instead of clean OOM --global-base 1024 rarely needs changing

Choosing a stack size

There is no universally right stack size, but there is a reliable way to find yours. The stack needs to hold the deepest chain of frames your program reaches, where each frame’s size is whatever the compiler spilled to the shadow stack for that function: address-taken locals, arrays, structs passed by pointer, and saved values across calls. Two things make that larger than people expect. Recursion multiplies a frame’s size by the depth, so a parser that recurses once per nesting level of its input needs a stack proportional to the deepest input it will accept. And large local arrays — a 16 KB scratch buffer declared inside a function — land on the stack in full.

Measure rather than guess. Fill the stack region with a known pattern at startup, run the deepest realistic workload, then scan for how much of the pattern was overwritten; that high-water mark plus a margin is your size. Emscripten’s STACK_OVERFLOW_CHECK and sanitizer builds do a version of this for you. Then do two things that matter more than the exact number: cap recursion depth in the code that recurses on input, so a hostile input cannot demand an unbounded stack, and use --stack-first, so the day the estimate is wrong the program traps instead of corrupting itself.

Remember the cost runs the other way too. A stack sized at 8 MB “to be safe” commits 8 MB of linear memory in every instance and every thread, which on a phone running several workers is a meaningful share of the memory budget.

Step 5 — reserve low memory when you need it

--global-base moves where static data begins. The 1,024 bytes below it are left unused by default, partly so that a null pointer dereference reads zeros rather than live data. Some hosts want a reserved region at a fixed low address — a scratch buffer the host writes into before each call, or a header shared by convention between modules:

clang ... -Wl,--global-base=65536      # reserve the first 64 KiB for the host

Document such conventions next to the build flags; a reserved region that one module honours and another does not is a memory corruption bug that only appears when they share memory.

Expected output

After raising the stack and adding --stack-first, a recursion that used to corrupt data now fails loudly:

RuntimeError: memory access out of bounds
    at parse_expr (wasm://wasm/0002a9be:wasm-function[37]:0x1c4f)
    at parse_expr (wasm://wasm/0002a9be:wasm-function[37]:0x1d02)
    ...

A trap at a consistent place, with a deep stack of the same function, identifies the overflow immediately.

Gotchas

  • Raising --initial-memory below the data plus stack. The linker reports initial memory too small; the initial size must cover everything placed statically.
  • --max-memory smaller than what the program needs. malloc returns null once growth fails. That is the point of a maximum, but make sure the program handles it — see handling out-of-memory in Wasm.
  • Stack size set at compile time, not link time. -z stack-size is a linker option; passed to the compiler without -Wl, it is ignored with a warning or not at all.
  • Imported memory smaller than the layout needs. With --import-memory, the host’s memory must be at least the module’s initial size, or instantiation fails.

Performance note

Layout choices cost nothing at runtime — the addresses are constants baked in at link time. The only performance effect is on startup: a module whose initial memory was 1 MB grew 14 times to reach its working set of 15 MB, adding about 3 ms on a phone; setting --initial-memory to 16 MB removed those grows, at the cost of committing the memory immediately.

Frequently Asked Questions

Does the stack grow at runtime? No. The shadow stack has the fixed size chosen at link time. Only the heap grows, through memory.grow.

Do threads each get their own stack? Yes — each thread’s stack is allocated from the heap when the thread starts, sized by the threading runtime (for example -sDEFAULT_PTHREAD_STACK_SIZE in Emscripten). The link-time stack size applies to the main thread.

What is the WebAssembly engine’s own stack? Separate. Engines limit the depth of native calls too; very deep recursion can exhaust that even with a large shadow stack. See fixing stack overflow in Wasm.

Can I see the layout from JavaScript? Export __heap_base and __data_end with --export=__heap_base --export=__data_end; they appear as WebAssembly.Global exports holding the addresses.

← Back to Linking Wasm Objects with wasm-ld