Comparing Baseline and Optimized Wasm Code

This page answers one task: a hot WebAssembly function runs slower than you expect, and benchmarks and profiles have not explained why. You want to see the actual machine code the engine produces for that function — in the baseline tier and the optimising tier — and use the difference to understand what the optimiser does well, what it cannot do, and how to change your source so it can.

Prerequisites

  • [ ] A small reproduction: the hot function in a module you can run in Node or a test browser.
  • [ ] A V8 build or Node version that supports code-printing flags (some require builds with disassembler support enabled).
  • [ ] Basic familiarity with x86-64 or ARM64 assembly.

Why look at machine code

WebAssembly is already low-level, but it is not what the CPU runs. The baseline compiler (Liftoff in V8) translates each Wasm instruction in a single pass with simple register allocation, often spilling values to the stack and reloading them. The optimising compiler (TurboFan) builds a graph, eliminates redundant work, keeps values in registers across loops, removes some bounds checks, and selects better instructions. Comparing the two for one function shows how much of the gap is register pressure, redundant loads, missed vectorisation or checks that could not be removed — information that profiles alone do not give.

Step 1 — isolate the function

Make a tiny module containing just the hot function and a driver, so the printed output is manageable. Keep the function’s signature and body as in the real code. If the function is generated from Rust or C, compile that small crate or file with the same flags as your real build. Keep the name section so functions are identifiable in output.

Step 2 — print the generated code

V8 offers flags to print generated code for Wasm functions; availability depends on how V8 was built (disassembler support), and names change. Typical usage:

node --print-wasm-code --no-wasm-tier-up --liftoff-only run.mjs > liftoff.txt     # baseline code
node --print-wasm-code --no-liftoff run.mjs > turbofan.txt                        # optimised code

Check node --v8-options | grep -i "print.*wasm" for what your build supports; if printing is unavailable, a debug or release build of V8 with v8_enable_disassembler (or d8, V8’s shell) provides it. The output lists each function with its machine instructions.

Comparing tiers for one hot function Isolate the hot function in a small module. Run it with baseline-only and optimised-only settings while printing generated code. Compare instruction counts, spills, bounds checks and loop structure between tiers. Form a hypothesis about what limits the optimiser, change the source, and check the new optimised code and benchmark. isolate hot function small module print Liftoff code baseline print TurboFan code optimised compare: spills, checks, loops find the limiter change source, re-check benchmark again

Step 3 — read the differences

Look for a few things in the hot loop:

  • Spills and reloads. Liftoff often stores values to the stack between instructions; TurboFan should keep loop variables in registers. Remaining spills in optimised code indicate register pressure — too many live values.
  • Bounds checks. On 64-bit platforms, engines usually rely on guard regions around linear memory instead of explicit bounds checks for 32-bit memories, so checks may be absent entirely; where explicit checks remain (memory64, some platforms or configurations), look for compare-and-branch sequences before loads.
  • Instruction selection. Multiplications by constants turned into shifts, address arithmetic folded into addressing modes, SIMD instructions used directly.
  • Calls. Calls to other functions in the loop that were not inlined — Wasm engines inline less aggressively than JavaScript JITs, so small helper calls in inner loops can cost more than expected.
What typically differs between baseline and optimised code Baseline code translates each instruction separately, spilling values to the stack, recomputing addresses and keeping every operation. Optimised code keeps loop values in registers, folds address arithmetic, removes redundant work and selects better instructions, though it may still contain calls and spills where the source creates register pressure. baseline (Liftoff) per-instruction translation frequent stack spills redundant address math fast to compile optimised (TurboFan) values in registers folded addressing modes redundant work removed fast to run

Step 4 — turn findings into source changes

Typical source-level fixes: reduce live values in the inner loop (split it, hoist invariant computations), replace small non-inlined helper calls in hot loops with inline code (in Rust, #[inline(always)] before the Wasm is generated — Wasm engines do not reliably inline across Wasm function calls), make loop bounds and strides simple so the toolchain can vectorise, and use explicit SIMD where auto-vectorisation fails. Re-print the optimised code after each change to confirm the effect, then benchmark.

Step 5 — keep perspective

Machine code shows how the engine compiled your function, not what the toolchain did. Often the more productive comparison is one level up: read the Wasm the toolchain produced (wasm2wat or wasm-tools print) to see whether LLVM and wasm-opt already removed the redundancy, inlined helpers and vectorised. Engines optimise less aggressively than LLVM, so getting the Wasm right matters more than coaxing the engine.

Other engines

SpiderMonkey and JavaScriptCore also have code-printing options, mainly in debug builds or behind environment variables intended for engine developers. The same reading strategy applies; differences between engines’ output explain browser-specific performance gaps for the same module.

A worked reading of one loop

Consider a loop that sums bytes from linear memory. In the Wasm, each iteration does local.get of the pointer, i32.load8_u, an i32.add into the accumulator, a pointer increment and a compare-and-branch. Liftoff’s output typically shows each of those as separate machine instructions with the accumulator and pointer loaded from and stored to stack slots — five or six memory operations per iteration that have nothing to do with the bytes being summed. TurboFan’s output keeps pointer and accumulator in registers, folds the memory base and pointer into one addressing mode for the load, and may unroll the loop. If TurboFan’s loop still contains a stack load, look for the reason: a call inside the loop (registers are saved around calls), more live values than available registers, or a value whose type changes in a way the optimiser cannot track. Reading one short loop this way builds intuition that carries over to larger functions.

Comparing toolchain choices through machine code

Machine code is also a good judge of toolchain options. Build the same function with opt-level = "s" and opt-level = 3, with and without wasm-opt, and with SIMD enabled or not, and print the optimised machine code for each. Differences that look small in the Wasm — an extra local, a different loop shape — sometimes produce substantially different machine code, and occasionally a “smaller” Wasm build yields a slower inner loop because a helper was not inlined. Seeing the actual instructions settles debates about which flags to ship faster than benchmarks alone, because it shows why one build is faster.

Keeping notes

Save the printed code for the key functions alongside the benchmark results and versions. When a later engine or toolchain update changes performance, a diff of old and new machine code for the same function often explains the change immediately.

Expected output

For a hot pixel loop, Liftoff code shows 41 instructions per iteration with 9 stack spills, TurboFan code 17 instructions with none; a remaining non-inlined helper call in the TurboFan loop is removed by inlining it in Rust, cutting the loop to 12 instructions and the benchmark time by 22%.

Gotchas

  • Printing code for a whole application. Output is unmanageable. Isolate the function.
  • Flags unavailable in release builds. Use a build with disassembler support or d8.
  • Assuming engines inline like LLVM. They inline less. Inline before Wasm is generated.
  • Optimising machine code before checking the Wasm. Fix the toolchain output first.
  • Comparing across versions. Code generation changes. Record versions.
  • Judging builds by Wasm size alone. A smaller build can have a slower inner loop. Check the machine code.

Performance note

Inlining a small helper called once per pixel removed a call and its argument shuffling from the optimised loop, reducing a filter from 3.6 ms to 2.8 ms per megapixel.

Instructions per loop iteration by tier Machine instructions in the inner loop of a pixel kernel for baseline Liftoff code, optimised TurboFan code, and optimised code after inlining a helper function in the Rust source. instructions per iteration Liftoff 41 instr TurboFan 17 instr TurboFan after inlining 12 instr

Frequently Asked Questions

Can DevTools show machine code? Not directly; DevTools shows Wasm disassembly. Machine code needs engine flags.

Are bounds checks always removed? On 64-bit hosts with 32-bit memories engines typically use guard regions; other configurations may keep explicit checks.

Is this worth doing often? Rarely — for the few hottest functions where profiles show unexplained cost.

Does SIMD show up in machine code? Yes — v128 operations map to SSE/AVX on x86-64 and NEON on ARM64.

Why does optimised code still load from the stack in my loop? Usually a call inside the loop, too many live values for the registers, or a value the optimiser cannot keep in a register.

Can machine code help choose build flags? Yes — print the optimised code for each flag combination; it shows why one build’s loop is faster.

Should I keep printed code with benchmark results? Yes — diffs of old and new machine code explain performance changes after engine or toolchain updates.

Does unrolling always help? Not always; it reduces branch overhead but increases code size and register pressure, so check the result.

What is the quickest first check before reading machine code? Read the toolchain’s Wasm for the function with wasm2wat; many problems are already visible there.

← Back to Engine Tiering & JIT Compilation