How the Wasm Operand Stack Works
This page answers one question: WebAssembly is described as a “stack machine” — what exactly is on that stack, how do instructions use it, and what happens to it when the engine compiles the code to machine instructions?
Prerequisites
- [ ] WABT installed (
wat2wasm,wasm-interp) to run the examples. - [ ] Basic familiarity with WAT syntax, as in writing your first WAT module by hand.
Three stacks that are easy to confuse
People talk about “the stack” in WebAssembly and mean three different things. The operand stack is an abstract concept in the
specification: a list of typed values that instructions push and pop. i32.const 5 pushes a value; i32.add pops two and pushes their sum.
The call stack is the engine’s real machine stack, holding return addresses and spilled registers for each active function call; your code
cannot see or address it. The shadow stack is a region of linear memory that compiled C, C++ and Rust use for data whose address is taken,
covered in understanding the shadow stack in linear memory.
This page is about the first one. The operand stack exists in the semantics of WebAssembly — it defines what each instruction means — but in a compiled program it mostly disappears into registers. Understanding it explains why WebAssembly validates so quickly, why its bytecode is compact, and why “stack machine” says nothing about speed.
Step 1 — trace a small function
Take a function that computes a * b + c:
(func $muladd (param $a i32) (param $b i32) (param $c i32) (result i32)
local.get $a ;; stack: [a]
local.get $b ;; stack: [a, b]
i32.mul ;; stack: [a*b]
local.get $c ;; stack: [a*b, c]
i32.add) ;; stack: [a*b+c] ← one i32 left, matching (result i32)
Each instruction has a fixed signature in terms of the stack. local.get pushes one value of the local’s type; i32.mul pops two i32
values and pushes one; at the end of the function exactly the declared results must remain. There is no instruction that reads “the third
item down” or swaps arbitrary positions — the stack is accessed strictly at the top, which is what makes it simple to check.
Step 2 — see validation as type checking on the stack
Before any code runs, the engine validates every function by simulating the operand stack with types instead of values. It walks the instructions once, tracking the type of every stack slot, and rejects the module if any instruction would pop the wrong type or the stack would not match at the end:
(func (result i32)
i32.const 1
f64.const 2.0
i32.add) ;; error: expected [i32, i32], got [i32, f64]
wat2wasm bad.wat -o bad.wasm
bad.wat:4:3: error: type mismatch in i32.add, expected [i32, i32] but got [i32, f64]
Because the stack is only accessed at the top and every block declares its inputs and outputs, validation is a single linear pass with no backtracking. That is one reason browsers can validate and compile WebAssembly as it streams in, as described in streaming instantiation vs ArrayBuffer instantiation.
Step 3 — follow the stack through blocks and calls
Blocks partition the operand stack. A block declares a type — what it consumes and what it leaves — and inside the block, instructions can only touch values pushed inside it. When a branch leaves the block, the engine discards anything extra and keeps exactly the declared results:
(func (param $x i32) (result i32)
(block $done (result i32)
i32.const -1 ;; default result
local.get $x
i32.eqz
br_if $done ;; if x == 0: leave the block with [-1]
drop ;; otherwise discard the default
local.get $x
i32.const 2
i32.mul)) ;; leave the block with [x*2]
A call pops the callee’s parameters from the caller’s stack and, when the callee returns, pushes its results. From the caller’s point of
view, a call is just an instruction with a signature. Each function starts with an empty operand stack of its own — callers cannot see into
callees’ stacks or vice versa — which is part of WebAssembly’s isolation between functions.
Step 4 — see where the stack goes when compiled
The operand stack is a specification device; engines do not keep a literal stack in memory at run time. The baseline compiler (Liftoff in V8) tracks, for each stack slot, where its value currently lives — a register, a constant, or a spill slot on the native stack — and emits machine code that operates on registers directly. The optimizing compiler (TurboFan, Ion, Cranelift) goes further: it converts the whole function into a graph of values with no stack at all and performs register allocation as a native compiler would. The function above becomes something like:
; x86-64, roughly
imul eax, esi ; a * b
add eax, edx ; + c
ret
No pushes, no pops. So the “stack machine” design affects how compact the bytecode is and how quickly it validates and compiles, not how fast the result runs — the details are in how Wasm locals map to machine registers.
Step 5 — run it in an interpreter to watch the stack
WABT’s interpreter can trace execution, which makes the operand stack visible:
wat2wasm muladd.wat -o muladd.wasm
wasm-interp muladd.wasm --run-all-exports --trace 2>&1 | head -12
#0. 4: V:3 | local.get $3
#0. 6: V:4 | local.get $3
#0. 8: V:5 | i32.mul 3, 4
#0. 9: V:4 | local.get $2
#0. 11: V:5 | i32.add 12, 5
#0. 12: V:4 | return
V: is the value-stack height (including locals in this interpreter’s accounting), and each arithmetic instruction shows the operands it
popped. Watching a short function this way is the quickest route to an intuition for how WAT you read in a disassembly actually behaves.
Why WebAssembly chose a stack
A register-based bytecode — instructions naming source and destination registers — would have been a reasonable alternative, and some
virtual machines use one. WebAssembly chose a structured stack machine for size and simplicity: an instruction like i32.add needs no operand
fields at all, because its inputs are implicitly the top of the stack, so typical code encodes in fewer bytes. Validation is a linear pass.
And because every engine performs its own register allocation anyway, nothing is lost at run time. The design trades a little readability of
hand-written code for a compact, quickly checked, easily compiled format — a good trade for code that is downloaded over networks and compiled
on phones.
Expected output
The muladd example called with (3, 4, 5) returns 17; the trace shows the stack growing to two values before each arithmetic instruction
and shrinking to one after.
Gotchas
- Thinking the stack is in memory. It is not addressable and not in linear memory; data that needs an address goes to the shadow stack.
- Leaving extra values at the end. Validation fails if a function or block ends with more or fewer values than declared. Add
dropfor unwanted values. - Branches with the wrong stack shape. A branch to a block must provide that block’s result types. Check what is on the stack at each
br. - Reading performance into the stack. Push and pop counts in WAT say nothing about the machine code; profile instead.
Performance note
Because validation is a single pass over the operand stack’s types, it is fast: a 4 MB module validated in about 15 ms in V8 on a laptop, and with streaming compilation, validation overlaps the download completely. The compact stack encoding is also part of why WebAssembly modules are typically smaller than equivalent JavaScript for the same logic.
Frequently Asked Questions
Is there a maximum operand stack depth? Engines impose limits on function size and nesting that bound it indirectly; real code never approaches them.
Can I access values below the top of the stack?
No. Store them in locals with local.set or local.tee and read them back with local.get.
Does the stack exist in the binary? Only implicitly, through the instructions’ signatures. The binary has no stack declarations — just instructions whose effects on the stack are known.
Do other virtual machines use the same design? The JVM and .NET’s intermediate language are also stack-based bytecodes with typed validation, which is part of why WebAssembly’s design felt familiar to compiler writers. WebAssembly’s structured control flow is stricter than either.
Why does local.tee exist?
It stores the top value into a local and leaves it on the stack, saving a local.get when a value is both stored and used immediately.
Related
- Stack vs heap execution model — the section this page belongs to.
- Understanding structured control flow in Wasm — blocks and branches in depth.
- Decoding Wasm opcodes for debugging — how stack instructions are encoded.
- How V8 compiles Wasm with Liftoff and TurboFan — the compilers that remove the stack.
← Back to Stack vs Heap Execution Model