How V8 Compiles Wasm with Liftoff and TurboFan

This page answers one question: what happens inside V8 — the engine in Chrome, Edge, Node and Deno — between receiving a WebAssembly module and running it fast, and how do its compiler tiers affect your module’s startup and performance?

Prerequisites

  • [ ] Chrome or Node for experiments with V8 flags.
  • [ ] A module of a few hundred kilobytes or more, so compile times are visible.

Two compilers with opposite goals

V8 compiles WebAssembly with two compilers. Liftoff is the baseline compiler: it makes a single pass over each function’s bytecode and emits machine code directly, with simple register handling and no optimization. It is extremely fast — it can compile tens of megabytes of WebAssembly per second per core — which lets a page start running a module almost as soon as it has downloaded. The code it produces is correct but slower than optimized code, typically by a factor of two to four for compute-heavy loops.

TurboFan is the optimizing compiler, shared with V8’s JavaScript pipeline. It builds a graph of each function, runs inlining, constant folding, loop optimizations and global register allocation, and produces code close to what a native optimizing compiler would. It is much slower to compile — an order of magnitude or more — which is why running it on everything up front would delay startup.

V8 uses both. Every function is compiled with Liftoff first, so the module is usable quickly. Functions that turn out to be hot are recompiled with TurboFan in the background and swapped in. That combination — fast start, fast steady state — is what “tiering” means.

A module's journey through V8's tiers Bytes stream in and Liftoff compiles functions as they arrive, so the module is instantiable shortly after the download finishes. Execution starts on Liftoff code; hot functions are recompiled by TurboFan on background threads and replaced, after which they run at full speed. 0 ms download + streaming compile 220 ms last byte: Liftoff done 240 ms instantiate; run on Liftoff code 420 ms hot code to TurboFan 800 ms optimized code swapped in

Step 1 — see streaming compilation at work

With instantiateStreaming or compileStreaming, V8 starts Liftoff on each function as soon as its bytes arrive, on background threads. By the time the last byte is received, most of the module is already compiled. Compiling from an ArrayBuffer instead must wait for the whole download, which is why streaming is the default advice — see streaming instantiation vs ArrayBuffer instantiation.

In a Chrome performance trace, look for v8.wasm.streamFromResponse, compile tasks on worker threads, and the gap between the last network byte and the instantiate call — on a well-served module it is small.

Step 2 — understand lazy compilation

Recent V8 versions go further and compile some functions lazily: a function is compiled with Liftoff only when first called, not when the module loads. Validation still happens up front, but code generation for functions that never run — error paths, rarely used features — is skipped. That cuts startup work for large modules with many cold functions, at the cost of a small compile pause the first time each function runs.

Lazy compilation is one reason the first call to a function can be slower than later calls, a pattern explained in why the first call into Wasm is slow.

Step 3 — watch dynamic tier-up

V8 decides which functions to optimize by counting: Liftoff code includes cheap counters on function entry and loop back-edges, and when a function’s budget runs out it is queued for TurboFan. Compilation happens on background threads; when finished, the next call to the function — and, through on-stack replacement for long loops, sometimes the current invocation — uses the optimized code.

You can observe the decisions with V8 flags in Node:

node --trace-wasm-tier-up bench.mjs 2>&1 | head
[wasm] tier-up: function #412 (parse_header) budget exhausted, queueing for TurboFan
[wasm] tier-up: function #87 (blur_row) budget exhausted, queueing for TurboFan

Flag names change between V8 versions; node --v8-options | grep -i 'wasm.*tier' lists what your version supports. The browser-side workflow is in watching Wasm tier-up in Chrome.

Liftoff and TurboFan side by side Liftoff compiles in a single fast pass with simple register handling and produces slower code. TurboFan builds an optimization graph, allocates registers globally and produces fast code, but takes much longer to compile. Liftoff (baseline) one pass, no graph tens of MB/s per core values cached in registers, spilled often 2-4× slower code on hot loops start fast TurboFan (optimizing) SSA graph, inlining, loop opts roughly 10-50× slower to compile global register allocation near-native code run fast

Step 4 — measure the effect of each tier

Forcing a single tier makes the trade-off concrete. In Node:

node --liftoff --no-wasm-tier-up bench.mjs     # Liftoff only: fastest start, slowest steady state
node --no-liftoff bench.mjs                    # TurboFan only: slowest start, fastest from the first call
node bench.mjs                                 # default tiering
                          compile    first call   steady state per call
Liftoff only               38 ms       1.9 ms        1.9 ms
TurboFan only             610 ms       0.6 ms        0.6 ms
default (tiered)           38 ms       1.9 ms        0.6 ms (after ~40 calls)

The default gets the best of both: Liftoff’s compile time and, once hot functions tier up, TurboFan’s steady-state speed. The cost is a warm-up period during which hot code runs at baseline speed — which matters for benchmarks, as discussed in avoiding JIT warm-up errors in Wasm benchmarks.

Step 5 — benefit from code caching

Compiled code can be cached. When a module is loaded with streaming compilation from a URL, Chrome can store the compiled machine code alongside the cached response and reuse it on later visits, skipping compilation entirely. V8 caches code after the module has been used for a while, so the cache tends to contain TurboFan code for hot functions — a returning visitor gets optimized code immediately. Keeping module URLs stable and cacheable, as in setting Cache-Control headers for Wasm, is what makes this work.

Memory and CPU costs of compilation

Compilation is not free in memory either, and on phones that matters as much as time. Liftoff code is larger than TurboFan code for the same function, because it is less optimized; a large module compiled entirely with Liftoff can occupy more memory for code than the module’s own size. TurboFan uses more memory while compiling — its graphs and register allocation are expensive — but produces compact code. V8 manages these trade-offs with limits on background compile threads and by freeing Liftoff code once a function’s optimized version replaces it.

For a page, the visible effects are background CPU usage in the first seconds after a module loads, and a temporary rise in memory. On low-end devices with few cores, background TurboFan work competes with the page’s own workers and rendering. If startup feels sluggish only on such devices, a trace showing many TurboFan compile tasks overlapping the page’s own work is a sign to shrink the module or defer the code that triggers tier-up, rather than a sign that anything is wrong with the module’s logic.

What this means for module authors

Several practical consequences follow. Smaller modules start faster because Liftoff time and download time both scale with size, so size work pays off at startup even though Liftoff is fast. Hot code should be concentrated in a few functions, because tier-up happens per function — a hot loop spread across many tiny functions that are not inlined spends longer in baseline code. Very large functions can take TurboFan a long time to optimize, delaying the speed-up for exactly the code that needs it; splitting a giant generated function can help. And short-lived pages that run a computation once may never benefit from TurboFan at all — for them, Liftoff speed and module size are what matter, which is worth remembering before optimizing a kernel that only ever runs on baseline code.

Expected output

With --trace-wasm-tier-up, a workload prints a handful of tier-up lines for the functions that do the real work, and none for functions called a few times. The steady-state timing after those lines matches the --no-liftoff run.

Gotchas

  • Benchmarking only the first calls. They run Liftoff code. Warm up until tier-up finishes.
  • Assuming all code is optimized. Cold functions stay on Liftoff indefinitely. That is fine; they are cold.
  • Relying on flag names. V8 flags are not a stable interface. Check the current names before scripting them.
  • Disabling streaming to “simplify” loading. It delays compilation until the download ends and can bypass the code cache.

Performance note

For a 4 MB module on a laptop, Liftoff compiled the whole module in 58 ms; compiling everything with TurboFan up front would have taken about 1.1 s. Tiering optimized the 37 functions that mattered within the first second of use, in the background, without delaying startup.

Compile time for a 4 MB module by strategy Time to compile all functions of a 4 MB module with Liftoff, with TurboFan, and the background TurboFan time actually spent under default tiering, which optimizes only hot functions. ms of compilation work on a laptop Liftoff, whole module 58 ms TurboFan, whole module 1,100 ms TurboFan, hot functions only (tiered) 140 ms

Frequently Asked Questions

Does V8 ever deoptimize Wasm like it does JavaScript? No. WebAssembly types are fixed, so optimized code never needs to bail out because of unexpected types. Tier-up is one-way.

Is there a tier between Liftoff and TurboFan? For WebAssembly, the two tiers are the main ones; V8’s JavaScript pipeline has more (Ignition, Sparkplug, Maglev). Engines evolve, so check current V8 documentation for changes.

Can I make V8 optimize everything up front? --no-liftoff does that in Node, and Chrome exposes similar flags for testing. It is a benchmarking tool, not a deployment setting.

Does this apply in Node and Deno? Yes — both use V8 with the same Wasm pipeline. Code caching in Node is available through v8.serialize-based or loader-level caches rather than the browser’s HTTP cache.

← Back to Engine Tiering & JIT Compilation