Comparing Baseline and Optimized Wasm Code
This page answers one task: a hot WebAssembly function runs slower than you expect, and benchmarks and profiles have not explained why. You want to see the actual machine code the engine produces for that function — in the baseline tier and the optimising tier — and use the difference to understand what the optimiser does well, what it cannot do, and how to change your source so it can.
Prerequisites
- [ ] A small reproduction: the hot function in a module you can run in Node or a test browser.
- [ ] A V8 build or Node version that supports code-printing flags (some require builds with disassembler support enabled).
- [ ] Basic familiarity with x86-64 or ARM64 assembly.
Why look at machine code
WebAssembly is already low-level, but it is not what the CPU runs. The baseline compiler (Liftoff in V8) translates each Wasm instruction in a single pass with simple register allocation, often spilling values to the stack and reloading them. The optimising compiler (TurboFan) builds a graph, eliminates redundant work, keeps values in registers across loops, removes some bounds checks, and selects better instructions. Comparing the two for one function shows how much of the gap is register pressure, redundant loads, missed vectorisation or checks that could not be removed — information that profiles alone do not give.
Step 1 — isolate the function
Make a tiny module containing just the hot function and a driver, so the printed output is manageable. Keep the function’s signature and body as in the real code. If the function is generated from Rust or C, compile that small crate or file with the same flags as your real build. Keep the name section so functions are identifiable in output.
Step 2 — print the generated code
V8 offers flags to print generated code for Wasm functions; availability depends on how V8 was built (disassembler support), and names change. Typical usage:
node --print-wasm-code --no-wasm-tier-up --liftoff-only run.mjs > liftoff.txt # baseline code
node --print-wasm-code --no-liftoff run.mjs > turbofan.txt # optimised code
Check node --v8-options | grep -i "print.*wasm" for what your build supports; if printing is unavailable, a debug or release build of V8 with
v8_enable_disassembler (or d8, V8’s shell) provides it. The output lists each function with its machine instructions.
Step 3 — read the differences
Look for a few things in the hot loop:
- Spills and reloads. Liftoff often stores values to the stack between instructions; TurboFan should keep loop variables in registers. Remaining spills in optimised code indicate register pressure — too many live values.
- Bounds checks. On 64-bit platforms, engines usually rely on guard regions around linear memory instead of explicit bounds checks for 32-bit memories, so checks may be absent entirely; where explicit checks remain (memory64, some platforms or configurations), look for compare-and-branch sequences before loads.
- Instruction selection. Multiplications by constants turned into shifts, address arithmetic folded into addressing modes, SIMD instructions used directly.
- Calls. Calls to other functions in the loop that were not inlined — Wasm engines inline less aggressively than JavaScript JITs, so small helper calls in inner loops can cost more than expected.
Step 4 — turn findings into source changes
Typical source-level fixes: reduce live values in the inner loop (split it, hoist invariant computations), replace small non-inlined helper calls in hot loops
with inline code (in Rust, #[inline(always)] before the Wasm is generated — Wasm engines do not reliably inline across Wasm function calls), make loop bounds
and strides simple so the toolchain can vectorise, and use explicit SIMD where auto-vectorisation fails. Re-print the optimised code after each change to confirm
the effect, then benchmark.
Step 5 — keep perspective
Machine code shows how the engine compiled your function, not what the toolchain did. Often the more productive comparison is one level up: read the Wasm the
toolchain produced (wasm2wat or wasm-tools print) to see whether LLVM and wasm-opt already removed the redundancy, inlined helpers and vectorised.
Engines optimise less aggressively than LLVM, so getting the Wasm right matters more than coaxing the engine.
Other engines
SpiderMonkey and JavaScriptCore also have code-printing options, mainly in debug builds or behind environment variables intended for engine developers. The same reading strategy applies; differences between engines’ output explain browser-specific performance gaps for the same module.
A worked reading of one loop
Consider a loop that sums bytes from linear memory. In the Wasm, each iteration does local.get of the pointer, i32.load8_u, an i32.add into the
accumulator, a pointer increment and a compare-and-branch. Liftoff’s output typically shows each of those as separate machine instructions with the
accumulator and pointer loaded from and stored to stack slots — five or six memory operations per iteration that have nothing to do with the bytes being
summed. TurboFan’s output keeps pointer and accumulator in registers, folds the memory base and pointer into one addressing mode for the load, and may unroll
the loop. If TurboFan’s loop still contains a stack load, look for the reason: a call inside the loop (registers are saved around calls), more live values than
available registers, or a value whose type changes in a way the optimiser cannot track. Reading one short loop this way builds intuition that carries over to
larger functions.
Comparing toolchain choices through machine code
Machine code is also a good judge of toolchain options. Build the same function with opt-level = "s" and opt-level = 3, with and without wasm-opt, and with
SIMD enabled or not, and print the optimised machine code for each. Differences that look small in the Wasm — an extra local, a different loop shape — sometimes
produce substantially different machine code, and occasionally a “smaller” Wasm build yields a slower inner loop because a helper was not inlined. Seeing the
actual instructions settles debates about which flags to ship faster than benchmarks alone, because it shows why one build is faster.
Keeping notes
Save the printed code for the key functions alongside the benchmark results and versions. When a later engine or toolchain update changes performance, a diff of old and new machine code for the same function often explains the change immediately.
Expected output
For a hot pixel loop, Liftoff code shows 41 instructions per iteration with 9 stack spills, TurboFan code 17 instructions with none; a remaining non-inlined helper call in the TurboFan loop is removed by inlining it in Rust, cutting the loop to 12 instructions and the benchmark time by 22%.
Gotchas
- Printing code for a whole application. Output is unmanageable. Isolate the function.
- Flags unavailable in release builds. Use a build with disassembler support or d8.
- Assuming engines inline like LLVM. They inline less. Inline before Wasm is generated.
- Optimising machine code before checking the Wasm. Fix the toolchain output first.
- Comparing across versions. Code generation changes. Record versions.
- Judging builds by Wasm size alone. A smaller build can have a slower inner loop. Check the machine code.
Performance note
Inlining a small helper called once per pixel removed a call and its argument shuffling from the optimised loop, reducing a filter from 3.6 ms to 2.8 ms per megapixel.
Frequently Asked Questions
Can DevTools show machine code? Not directly; DevTools shows Wasm disassembly. Machine code needs engine flags.
Are bounds checks always removed? On 64-bit hosts with 32-bit memories engines typically use guard regions; other configurations may keep explicit checks.
Is this worth doing often? Rarely — for the few hottest functions where profiles show unexplained cost.
Does SIMD show up in machine code? Yes — v128 operations map to SSE/AVX on x86-64 and NEON on ARM64.
Why does optimised code still load from the stack in my loop? Usually a call inside the loop, too many live values for the registers, or a value the optimiser cannot keep in a register.
Can machine code help choose build flags? Yes — print the optimised code for each flag combination; it shows why one build’s loop is faster.
Should I keep printed code with benchmark results? Yes — diffs of old and new machine code explain performance changes after engine or toolchain updates.
Does unrolling always help? Not always; it reduces branch overhead but increases code size and register pressure, so check the result.
What is the quickest first check before reading machine code? Read the toolchain’s Wasm for the function with wasm2wat; many problems are already visible there.
Related
- How V8 compiles Wasm with Liftoff and TurboFan — the tiers.
- Using engine flags to experiment with tiering — flags.
- Converting Wasm back to WAT with wasm2wat — reading the Wasm.
- How Wasm locals map to machine registers — register allocation.
← Back to Engine Tiering & JIT Compilation