Setting Stack Size and Memory Limits at Link Time
This page answers one question: how does wasm-ld decide where static data, the stack and the heap live in linear memory and
how big each may be — and which flags fix the crashes and failures that come from getting those sizes wrong?
Prerequisites
- [ ] A C, C++ or Rust module linked with
wasm-ld(directly or through clang, rustc or emcc). - [ ]
wasm-objdumpto read the memory section and globals. - [ ] A symptom, ideally: a stack overflow, an out-of-memory error, or corrupted static data.
The layout wasm-ld produces
Compiled C, C++ and Rust need three regions in linear memory. Static data holds string literals, initialised globals and
zero-initialised statics. The shadow stack holds stack frames for anything that needs an address — arrays, structs passed by
pointer, variables whose address is taken — because WebAssembly’s own operand stack cannot be addressed; the background is in
understanding the shadow stack in linear memory.
And the heap is everything above, managed by malloc.
wasm-ld lays these out at link time. By default static data starts at --global-base (1024), the stack comes next with a fixed
size and grows downward towards the data, and the heap starts at the symbol __heap_base, after the stack. A global named
__stack_pointer holds the current stack top. Nothing at runtime checks whether the stack pointer has moved past the bottom of
the stack region — if it does, the next frame silently overwrites static data.
Step 1 — read the current layout
The linker exports the boundaries as globals in some configurations, and they are always visible in the name section and the global section:
wasm-objdump -x module.wasm | grep -A12 '^Global\['
wasm-objdump -x -j Memory module.wasm
Global[3]:
- global[0] i32 mutable=1 <__stack_pointer> - init i32=1101824
- global[1] i32 mutable=0 <__data_end> - init i32=1036288
- global[2] i32 mutable=0 <__heap_base> - init i32=1101824
Memory[1]:
- memory[0] pages: initial=17 max=16384
Here static data ends at about 1,012 KB, the stack occupies the 64 KiB above it (the initial stack pointer is the top), and the heap starts immediately after.
Step 2 — raise the stack size for deep recursion
The default stack is 64 KiB for most toolchains — Rust uses 1 MiB for wasm32-unknown-unknown, Emscripten 64 KiB in recent
releases. Recursive parsers, large local arrays and deep call chains overflow it. With C and clang:
clang --target=wasm32 -O2 parser.c -o parser.wasm -Wl,--no-entry -Wl,-z,stack-size=1048576
From Rust:
# .cargo/config.toml
[target.wasm32-unknown-unknown]
rustflags = ["-C", "link-arg=-zstack-size=2097152"]
With Emscripten, use -sSTACK_SIZE=1MB. The stack is carved out of linear memory, so a larger stack raises the module’s minimum
memory by the same amount.
Step 3 — put the stack first so overflows trap
Because the stack grows downward towards static data, an overflow corrupts data silently. --stack-first reorders the layout so
the stack sits at the bottom of memory, below the data. An overflow then runs the stack pointer below address zero, which
wraps to a huge unsigned address beyond the end of memory, and the next access traps instead of corrupting anything:
clang --target=wasm32 -O2 parser.c -o parser.wasm \
-Wl,--no-entry -Wl,-z,stack-size=1048576 -Wl,--stack-first
Rust’s wasm32-unknown-unknown target already uses --stack-first; for C builds it is opt-in and almost always worth enabling.
Emscripten offers -sSTACK_OVERFLOW_CHECK=2 for explicit checking instead.
Step 4 — set initial and maximum memory
--initial-memory sets how much memory the module starts with; it must cover static data plus the stack, and the linker rounds
up to whole 64 KiB pages. --max-memory sets the ceiling memory.grow can reach. Without a maximum, the module may grow to the
engine’s limit — 4 GiB for 32-bit memories in most browsers — which is usually more than a page should ever use.
clang --target=wasm32 -O2 app.c -o app.wasm -Wl,--no-entry \
-Wl,--initial-memory=4194304 \
-Wl,--max-memory=134217728
A sensible initial size avoids repeated growth during startup, each of which costs a reallocation; a sensible maximum turns a runaway allocation into a clean out-of-memory error instead of the tab being killed. How to choose the numbers is covered in sizing initial and maximum memory.
Choosing a stack size
There is no universally right stack size, but there is a reliable way to find yours. The stack needs to hold the deepest chain of frames your program reaches, where each frame’s size is whatever the compiler spilled to the shadow stack for that function: address-taken locals, arrays, structs passed by pointer, and saved values across calls. Two things make that larger than people expect. Recursion multiplies a frame’s size by the depth, so a parser that recurses once per nesting level of its input needs a stack proportional to the deepest input it will accept. And large local arrays — a 16 KB scratch buffer declared inside a function — land on the stack in full.
Measure rather than guess. Fill the stack region with a known pattern at startup, run the deepest realistic workload, then scan for
how much of the pattern was overwritten; that high-water mark plus a margin is your size. Emscripten’s STACK_OVERFLOW_CHECK and
sanitizer builds do a version of this for you. Then do two things that matter more than the exact number: cap recursion depth in the
code that recurses on input, so a hostile input cannot demand an unbounded stack, and use --stack-first, so the day the estimate
is wrong the program traps instead of corrupting itself.
Remember the cost runs the other way too. A stack sized at 8 MB “to be safe” commits 8 MB of linear memory in every instance and every thread, which on a phone running several workers is a meaningful share of the memory budget.
Step 5 — reserve low memory when you need it
--global-base moves where static data begins. The 1,024 bytes below it are left unused by default, partly so that a null pointer
dereference reads zeros rather than live data. Some hosts want a reserved region at a fixed low address — a scratch buffer the
host writes into before each call, or a header shared by convention between modules:
clang ... -Wl,--global-base=65536 # reserve the first 64 KiB for the host
Document such conventions next to the build flags; a reserved region that one module honours and another does not is a memory corruption bug that only appears when they share memory.
Expected output
After raising the stack and adding --stack-first, a recursion that used to corrupt data now fails loudly:
RuntimeError: memory access out of bounds
at parse_expr (wasm://wasm/0002a9be:wasm-function[37]:0x1c4f)
at parse_expr (wasm://wasm/0002a9be:wasm-function[37]:0x1d02)
...
A trap at a consistent place, with a deep stack of the same function, identifies the overflow immediately.
Gotchas
- Raising
--initial-memorybelow the data plus stack. The linker reportsinitial memory too small; the initial size must cover everything placed statically. --max-memorysmaller than what the program needs.mallocreturns null once growth fails. That is the point of a maximum, but make sure the program handles it — see handling out-of-memory in Wasm.- Stack size set at compile time, not link time.
-z stack-sizeis a linker option; passed to the compiler without-Wl,it is ignored with a warning or not at all. - Imported memory smaller than the layout needs. With
--import-memory, the host’s memory must be at least the module’s initial size, or instantiation fails.
Performance note
Layout choices cost nothing at runtime — the addresses are constants baked in at link time. The only performance effect is on
startup: a module whose initial memory was 1 MB grew 14 times to reach its working set of 15 MB, adding about 3 ms on a phone; setting
--initial-memory to 16 MB removed those grows, at the cost of committing the memory immediately.
Frequently Asked Questions
Does the stack grow at runtime?
No. The shadow stack has the fixed size chosen at link time. Only the heap grows, through memory.grow.
Do threads each get their own stack?
Yes — each thread’s stack is allocated from the heap when the thread starts, sized by the threading runtime (for example
-sDEFAULT_PTHREAD_STACK_SIZE in Emscripten). The link-time stack size applies to the main thread.
What is the WebAssembly engine’s own stack? Separate. Engines limit the depth of native calls too; very deep recursion can exhaust that even with a large shadow stack. See fixing stack overflow in Wasm.
Can I see the layout from JavaScript?
Export __heap_base and __data_end with --export=__heap_base --export=__data_end; they appear as WebAssembly.Global exports
holding the addresses.
Related
- Importing memory from the host with --import-memory — layout constraints on imported memory.
- Understanding Wasm linear memory limits — engine-level ceilings.
- Catching memory bugs with Emscripten sanitizers — finding the corruption a small stack causes.
- Tracking down memory corruption in Wasm — when the symptom is all you have.
← Back to Linking Wasm Objects with wasm-ld