Understanding Lazy Compilation in V8
This page answers one question: why does a large WebAssembly module instantiate quickly in Chrome or Node, yet the first call to some functions takes noticeably longer than later calls? The answer is lazy compilation. You want to understand how V8 decides when to compile each function, what that means for your module’s startup and responsiveness, and how to work with it.
Prerequisites
- [ ] A Wasm module large enough for compile time to matter (hundreds of kilobytes or more).
- [ ] Chrome or Node.js (both use V8).
- [ ] Optionally, Chrome’s Performance panel for viewing compile events.
From eager to lazy compilation
Early WebAssembly engines compiled every function in a module before instantiation finished. For a 10 MB module, that meant hundreds of milliseconds or more of compilation before any code could run — even for functions never called. V8 introduced a fast baseline compiler (Liftoff) to make that cheaper, then moved to lazy compilation: when a module is compiled, V8 validates it (checking that it is well-formed and type-correct) but compiles function bodies only when they are first called. Each function gets a small stub; the first call into the stub compiles the function with Liftoff and patches the call target, so later calls go straight to compiled code.
On top of that, dynamic tiering decides which functions deserve the optimising compiler (TurboFan): functions that run often are recompiled with TurboFan in the background, and execution switches to the optimised code when ready. The combination means startup cost is proportional to the code actually used, and optimisation effort goes to hot code.
What this means for your module
Fast instantiation. WebAssembly.compile and instantiate complete sooner because only validation happens up front; startup cost no longer scales with
total code size so directly.
Slower first calls. The first call to each function pays its Liftoff compilation — typically microseconds for small functions, more for large ones. A user action that touches hundreds of previously unused functions (opening a complex dialog, the first export of a document) can therefore pause briefly while they compile.
Gradual speed-up. Hot functions run Liftoff code first and switch to TurboFan code later; benchmarks that do not warm up measure a mix of tiers.
Step 1 — see it in a profile
Record a Performance trace in Chrome around the first use of a feature. Compile events (with names such as wasm.CompileLazy or similar in the trace) appear
interleaved with execution, attributed to the functions being compiled. Compare with a second use of the same feature: the compile events disappear.
Step 2 — warm up predictable paths
If a feature’s first use must feel instant, call its main functions once with a small input during idle time after instantiation — for example in
requestIdleCallback or in a worker right after loading. That triggers Liftoff compilation of the code path before the user needs it. Keep warm-up inputs
small and side-effect free.
await init();
requestIdleCallback(() => {
wasm.export_pdf(tinyDocument); // compiles the export path's functions ahead of the user's click
});
Step 3 — understand code caching’s role
When a module is loaded with streaming APIs from a cacheable URL, Chrome can cache compiled code for later visits. With lazy compilation, what gets cached is the code that was compiled — functions that ran — once the cache entry is written (Chrome writes it after the module has been used for a while). Returning users therefore start with the functions they actually used already compiled. Details are in how Wasm code caching works in browsers.
Step 4 — turn it off only for experiments
V8 flags can disable lazy compilation (--no-wasm-lazy-compilation) or tiering, which is useful for experiments — measuring total compile time, comparing
tiers — but flags are for testing, not production, and their names and defaults change between versions. In Node, pass them on the command line; in Chrome, use
--js-flags when launching a test browser. See
using engine flags to experiment with tiering.
Step 5 — shape your module for lazy compilation
Lazy compilation rewards modules where rarely used functionality is separate from common paths. Very large functions (generated parsers, huge match
statements) compile slowly even with Liftoff; splitting them can reduce first-call pauses. Functions that are always called during startup gain nothing from
laziness — they compile immediately anyway — so startup-critical code should be lean.
Other engines
Firefox’s SpiderMonkey and Safari’s JavaScriptCore have their own strategies: SpiderMonkey uses a fast baseline compiler and an optimising compiler with background compilation, and has used eager baseline compilation of the whole module with parallel threads; JavaScriptCore uses an interpreter and multiple compiler tiers. Behaviour on first calls therefore differs between browsers; measure in each when first-interaction latency matters.
Lazy compilation and validation errors
Because validation still happens for the whole module up front, a malformed module fails at WebAssembly.compile exactly as before; laziness never turns a
validation error into a runtime error. What changes is where compile-time resource problems appear. A function too large for the engine’s internal limits,
or code that triggers a compiler bug, now fails or crashes when the function is first called rather than at instantiation, which can make such problems look
like runtime bugs in a specific feature. If a crash happens consistently on the first use of a particular function and never afterwards in the same session,
suspect compilation; reproducing with lazy compilation disabled in a test browser moves the failure to startup and confirms it.
Interaction with workers and multiple instances
Compiled code belongs to the WebAssembly.Module, not the instance. Instantiating the same module several times — or posting the module to workers in the same
process — lets instances benefit from functions another instance already compiled, depending on engine version and how the module is shared. In practice, a
pattern of “compile once, instantiate many” in workers means the first worker to call a function pays its compilation, and later workers find it ready. If
each worker compiles the module from bytes independently, each pays separately; that is one more reason to compile once on the main thread and post the
Module to workers.
Large generated functions
Code generators — parser generators, interpreters’ dispatch loops, generated bindings — often produce single functions of tens of thousands of instructions. These compile slowly with every tier and can delay the first call noticeably. Splitting generated code into smaller functions, or restructuring dispatch loops into tables of handlers, reduces both first-call pauses and the time TurboFan spends optimising them.
Expected output
A 9 MB module instantiates in about 40 ms in Chrome; the first PDF export pauses 80 ms for lazy compilation of about 1,200 functions, visible as compile events in the trace; an idle-time warm-up call removes that pause from the user’s first click; and a later visit with code caching shows almost no compile events.
Gotchas
- Measuring first calls as steady state. They include compilation. Warm up before benchmarking.
- Assuming instantiation compiled everything. Only validation happened. Expect first-call costs.
- Using V8 flags in production. They are for experiments and change between versions.
- Huge single functions. Slow to compile on first call. Split them.
- Assuming all browsers behave like V8. Measure first-call latency in each engine.
- Compiling the module separately in every worker. Each pays compilation. Compile once and post the Module.
Performance note
For the 9 MB module, eager compilation with Liftoff took about 310 ms before the first call; lazy compilation made instantiation about 40 ms, with compile cost spread across first calls — 80 ms for the export path.
Frequently Asked Questions
Is validation still done up front? Yes — the whole module is validated before instantiation, so invalid modules fail early.
Does lazy compilation affect correctness? No — only timing.
Can I force compilation of specific functions? Not directly; calling them once (warm-up) compiles them.
Does WebAssembly.Module sent to a worker share compiled code?
Compiled code can be shared within the process; behaviour details vary by version.
Why does a crash happen only on the first use of a feature? With lazy compilation, compile-time failures appear on a function’s first call; reproduce with lazy compilation disabled to confirm.
Do workers share functions another worker already compiled? When they share one compiled Module in the same process, they can; compiling from bytes in each worker does not share.
Should generated code be split into smaller functions? Yes, where practical — huge functions compile slowly in every tier and delay first calls.
Related
- How V8 compiles Wasm with Liftoff and TurboFan — the tiers.
- Why the first call into Wasm is slow — first-call costs.
- Watching Wasm tier-up in Chrome — seeing tiering.
- Measuring Wasm startup time end to end — startup phases.
← Back to Engine Tiering & JIT Compilation