How V8 Compiles Wasm with Liftoff and TurboFan
This page answers one question: what happens inside V8 — the engine in Chrome, Edge, Node and Deno — between receiving a WebAssembly module and running it fast, and how do its compiler tiers affect your module’s startup and performance?
Prerequisites
- [ ] Chrome or Node for experiments with V8 flags.
- [ ] A module of a few hundred kilobytes or more, so compile times are visible.
Two compilers with opposite goals
V8 compiles WebAssembly with two compilers. Liftoff is the baseline compiler: it makes a single pass over each function’s bytecode and emits machine code directly, with simple register handling and no optimization. It is extremely fast — it can compile tens of megabytes of WebAssembly per second per core — which lets a page start running a module almost as soon as it has downloaded. The code it produces is correct but slower than optimized code, typically by a factor of two to four for compute-heavy loops.
TurboFan is the optimizing compiler, shared with V8’s JavaScript pipeline. It builds a graph of each function, runs inlining, constant folding, loop optimizations and global register allocation, and produces code close to what a native optimizing compiler would. It is much slower to compile — an order of magnitude or more — which is why running it on everything up front would delay startup.
V8 uses both. Every function is compiled with Liftoff first, so the module is usable quickly. Functions that turn out to be hot are recompiled with TurboFan in the background and swapped in. That combination — fast start, fast steady state — is what “tiering” means.
Step 1 — see streaming compilation at work
With instantiateStreaming or compileStreaming, V8 starts Liftoff on each function as soon as its bytes arrive, on background threads. By the
time the last byte is received, most of the module is already compiled. Compiling from an ArrayBuffer instead must wait for the whole download,
which is why streaming is the default advice — see
streaming instantiation vs ArrayBuffer instantiation.
In a Chrome performance trace, look for v8.wasm.streamFromResponse, compile tasks on worker threads, and the gap between the last network byte
and the instantiate call — on a well-served module it is small.
Step 2 — understand lazy compilation
Recent V8 versions go further and compile some functions lazily: a function is compiled with Liftoff only when first called, not when the module loads. Validation still happens up front, but code generation for functions that never run — error paths, rarely used features — is skipped. That cuts startup work for large modules with many cold functions, at the cost of a small compile pause the first time each function runs.
Lazy compilation is one reason the first call to a function can be slower than later calls, a pattern explained in why the first call into Wasm is slow.
Step 3 — watch dynamic tier-up
V8 decides which functions to optimize by counting: Liftoff code includes cheap counters on function entry and loop back-edges, and when a function’s budget runs out it is queued for TurboFan. Compilation happens on background threads; when finished, the next call to the function — and, through on-stack replacement for long loops, sometimes the current invocation — uses the optimized code.
You can observe the decisions with V8 flags in Node:
node --trace-wasm-tier-up bench.mjs 2>&1 | head
[wasm] tier-up: function #412 (parse_header) budget exhausted, queueing for TurboFan
[wasm] tier-up: function #87 (blur_row) budget exhausted, queueing for TurboFan
Flag names change between V8 versions; node --v8-options | grep -i 'wasm.*tier' lists what your version supports. The browser-side workflow is
in watching Wasm tier-up in Chrome.
Step 4 — measure the effect of each tier
Forcing a single tier makes the trade-off concrete. In Node:
node --liftoff --no-wasm-tier-up bench.mjs # Liftoff only: fastest start, slowest steady state
node --no-liftoff bench.mjs # TurboFan only: slowest start, fastest from the first call
node bench.mjs # default tiering
compile first call steady state per call
Liftoff only 38 ms 1.9 ms 1.9 ms
TurboFan only 610 ms 0.6 ms 0.6 ms
default (tiered) 38 ms 1.9 ms 0.6 ms (after ~40 calls)
The default gets the best of both: Liftoff’s compile time and, once hot functions tier up, TurboFan’s steady-state speed. The cost is a warm-up period during which hot code runs at baseline speed — which matters for benchmarks, as discussed in avoiding JIT warm-up errors in Wasm benchmarks.
Step 5 — benefit from code caching
Compiled code can be cached. When a module is loaded with streaming compilation from a URL, Chrome can store the compiled machine code alongside the cached response and reuse it on later visits, skipping compilation entirely. V8 caches code after the module has been used for a while, so the cache tends to contain TurboFan code for hot functions — a returning visitor gets optimized code immediately. Keeping module URLs stable and cacheable, as in setting Cache-Control headers for Wasm, is what makes this work.
Memory and CPU costs of compilation
Compilation is not free in memory either, and on phones that matters as much as time. Liftoff code is larger than TurboFan code for the same function, because it is less optimized; a large module compiled entirely with Liftoff can occupy more memory for code than the module’s own size. TurboFan uses more memory while compiling — its graphs and register allocation are expensive — but produces compact code. V8 manages these trade-offs with limits on background compile threads and by freeing Liftoff code once a function’s optimized version replaces it.
For a page, the visible effects are background CPU usage in the first seconds after a module loads, and a temporary rise in memory. On low-end
devices with few cores, background TurboFan work competes with the page’s own workers and rendering. If startup feels sluggish only on such
devices, a trace showing many TurboFan compile tasks overlapping the page’s own work is a sign to shrink the module or defer the code that
triggers tier-up, rather than a sign that anything is wrong with the module’s logic.
What this means for module authors
Several practical consequences follow. Smaller modules start faster because Liftoff time and download time both scale with size, so size work pays off at startup even though Liftoff is fast. Hot code should be concentrated in a few functions, because tier-up happens per function — a hot loop spread across many tiny functions that are not inlined spends longer in baseline code. Very large functions can take TurboFan a long time to optimize, delaying the speed-up for exactly the code that needs it; splitting a giant generated function can help. And short-lived pages that run a computation once may never benefit from TurboFan at all — for them, Liftoff speed and module size are what matter, which is worth remembering before optimizing a kernel that only ever runs on baseline code.
Expected output
With --trace-wasm-tier-up, a workload prints a handful of tier-up lines for the functions that do the real work, and none for functions called a
few times. The steady-state timing after those lines matches the --no-liftoff run.
Gotchas
- Benchmarking only the first calls. They run Liftoff code. Warm up until tier-up finishes.
- Assuming all code is optimized. Cold functions stay on Liftoff indefinitely. That is fine; they are cold.
- Relying on flag names. V8 flags are not a stable interface. Check the current names before scripting them.
- Disabling streaming to “simplify” loading. It delays compilation until the download ends and can bypass the code cache.
Performance note
For a 4 MB module on a laptop, Liftoff compiled the whole module in 58 ms; compiling everything with TurboFan up front would have taken about 1.1 s. Tiering optimized the 37 functions that mattered within the first second of use, in the background, without delaying startup.
Frequently Asked Questions
Does V8 ever deoptimize Wasm like it does JavaScript? No. WebAssembly types are fixed, so optimized code never needs to bail out because of unexpected types. Tier-up is one-way.
Is there a tier between Liftoff and TurboFan? For WebAssembly, the two tiers are the main ones; V8’s JavaScript pipeline has more (Ignition, Sparkplug, Maglev). Engines evolve, so check current V8 documentation for changes.
Can I make V8 optimize everything up front?
--no-liftoff does that in Node, and Chrome exposes similar flags for testing. It is a benchmarking tool, not a deployment setting.
Does this apply in Node and Deno?
Yes — both use V8 with the same Wasm pipeline. Code caching in Node is available through v8.serialize-based or loader-level caches rather than
the browser’s HTTP cache.
Related
- How SpiderMonkey compiles Wasm — Firefox’s equivalent design.
- Pinning a compiler tier for benchmarks — using the flags above deliberately.
- How Wasm locals map to machine registers — what each tier does with locals.
- Reducing Wasm cold-start latency — making the Liftoff phase short.
← Back to Engine Tiering & JIT Compilation