Engine Tiering & JIT Compilation
WebAssembly is often described as “compiled ahead of time”, and from the developer’s side it is: Rust, C or C++ is compiled to a .wasm file
before it ever reaches a user. But that file is bytecode, not machine code, and every browser engine still has to compile it again — on the user’s
device, at load time, under a time budget. How engines do that second compilation decides how quickly a module becomes usable, how long it takes
to reach full speed, and why benchmarks so often disagree with each other. This topic explains the tiered compilation strategies of the three
major engines, how to observe them, and how to design and measure modules with them in mind.
Prerequisites
- [ ] A module of realistic size — a few hundred kilobytes or more — so compilation is measurable.
- [ ] Chrome and Node for V8 experiments, Firefox for SpiderMonkey, and Safari (ideally on an iPhone) for JavaScriptCore.
- [ ] A benchmark harness with warm-up and medians, as in building a reproducible Wasm benchmark harness.
- [ ] A build that keeps the
namesection, so compiled functions are identifiable in traces.
Why engines compile in tiers
An engine faces a trade-off it cannot escape with a single compiler. A compiler that produces excellent machine code — inlining, register allocation across a whole function, loop optimizations — is slow, and running it on every function of a multi-megabyte module before execution starts would delay pages by seconds on phones. A compiler that is fast enough to keep pace with the network produces code that runs two to four times slower. Users want both: a module that starts immediately and runs fast once it is doing real work.
Tiering resolves the trade-off by using more than one compiler. Every function is first compiled — or interpreted — by a fast tier, so execution starts as soon as possible. Functions that turn out to matter are then recompiled by an optimizing tier, usually on background threads, and the faster code replaces the slower code while the program keeps running. Because WebAssembly’s types are fixed and fully known, optimized code never needs to fall back to slower code the way JavaScript JITs sometimes must — tier-up is a one-way street.
How engines got here
Tiering for WebAssembly was not the first design engines tried, and the history explains the current shape. The earliest implementations, around 2017, compiled modules entirely with their optimizing compilers before running anything. That produced fast code and slow startup: large modules — game engines exported from Unity and Unreal were the early stress tests — could take many seconds to become usable on laptops and far longer on phones. Engines responded by adding fast baseline compilers whose job was purely to get a module running, then layering optimization on top in the background.
The next problem was memory and CPU on constrained devices. Compiling an entire module twice — once with the baseline tier, once with the optimizing tier — doubles compile work and temporarily doubles code memory. That pushed V8 towards optimizing only hot functions and compiling lazily, so code that never runs is never compiled. JavaScriptCore went further with an interpreter tier that needs no compilation at all before first execution. The strategies continue to evolve: each engine retunes its thresholds, threading and caching as real-world modules and devices change.
Two constants have held throughout. Validation is always done up front and in one pass, which the format was designed to make cheap. And WebAssembly never deoptimizes: because types are fixed, optimized code stays valid, so tier-up is a one-way, low-risk transition that engines can apply aggressively.
How the three engines differ
All three major engines tier, but they make different choices about which tiers exist and when to move up.
V8 — Chrome, Edge, Node and Deno — uses Liftoff as its baseline compiler and TurboFan as its optimizing compiler. Liftoff compiles quickly as bytes stream in (and, in recent versions, lazily on first call), and V8 optimizes functions on demand: counters in Liftoff code track how much work each function does, and those that exhaust their budget are queued for TurboFan. The full story is in how V8 compiles Wasm with Liftoff and TurboFan.
SpiderMonkey — Firefox — uses its own baseline compiler and Ion. Its strategy has favoured starting optimization of the whole module in the background right after the baseline compile, so optimized code arrives soon for every function rather than only for hot ones. See how SpiderMonkey compiles Wasm.
JavaScriptCore — Safari and every browser on iOS — adds an interpreter below its two compilers: functions start in the in-place interpreter, move to BBQ when warm and to OMG when hot. That gives the fastest possible start at the cost of slow early execution, as described in how Safari runs Wasm.
Server-side runtimes take a different path again. wasmtime compiles every function ahead of time with Cranelift — optionally caching or precompiling the result — so there is no tier-up within a run. That makes server-side benchmarks more stable, and it is why the techniques on this page matter mostly for browsers.
Step-by-step: observing tiering in your own module
Seeing tiering on a real module takes a few minutes and makes every later performance question easier to reason about.
- Time calls individually. Call the hot export a few hundred times and record each call’s duration. A step down after some dozens of calls is the optimized code arriving:
const t = [];
for (let i = 0; i < 300; i++) { const s = performance.now(); exports.kernel(ptr, n); t.push(performance.now() - s); }
console.log(t.filter((_, i) => i % 25 === 0).map((x) => x.toFixed(2)).join(" "));
-
Find the compile work in a trace. Record a Performance panel trace and look for compilation tasks on V8’s background threads just before the step — the method is in watching Wasm tier-up in Chrome.
-
Pin tiers to bracket the behaviour. Run the same benchmark in Node with baseline only and optimized only to see the two speeds the default moves between:
node --liftoff --no-wasm-tier-up bench.mjs
node --no-liftoff bench.mjs
- Repeat on a phone. The ratios and timings change substantially on mobile hardware, and that is usually where tiering is most visible to users.
The JavaScript side: what callers experience
From JavaScript, tiering is invisible in the API — exports are the same function objects before and after tier-up — but it shapes three things callers notice. The first call to an export can be several times slower than later ones, because it may include lazy compilation and runs on baseline code; why the first call into Wasm is slow breaks the costs down. Short tasks may never reach optimized code at all. And timing measurements taken early in a session understate steady-state performance. Wrappers that expose Wasm functionality can help by warming up hot paths during idle time and by keeping one instance alive for repeated work rather than creating fresh ones per call:
// warm the hot path once the page is idle, so the user's first request runs on warm code
requestIdleCallback(() => engine.process(smallRepresentativeInput));
Optimization flags and trade-offs
Compile-time choices interact with engine tiering in ways worth knowing.
Module size. Baseline compilation time scales roughly with module size, so smaller modules start faster in every engine. Size optimizations —
-Oz, LTO, removing dead dependencies — therefore help startup even when they slow optimized code slightly.
Function size. Optimizing compilers work per function, and their cost grows faster than linearly with function size. A huge generated function — an interpreter’s dispatch loop, a fully unrolled table — can take hundreds of milliseconds to optimize, during which it runs on baseline code. Splitting such functions lets the hot parts reach optimized code sooner.
Hot-code concentration. Engines tier up per function. Work concentrated in a few functions reaches optimized code quickly; work spread thinly across many small functions that the source compiler did not inline may take longer to warm up.
Caching. All three browsers can cache compiled code for modules loaded via streaming APIs from stable URLs. A returning visitor then skips most compilation and often starts directly on optimized code — which makes content-hashed, long-cached module URLs one of the most effective tiering optimizations available, as described in setting Cache-Control headers for Wasm.
Designing modules for tiered engines
Knowing how engines tier suggests a handful of design habits that make modules start faster and reach full speed sooner, without targeting any one engine.
Keep the startup path small. The code needed for the first screen or the first interaction should be a small fraction of the module, so it compiles quickly in every baseline tier and can be warmed up cheaply. Code needed later can live in the same module — lazy compilation means it costs little until used — or in a separately loaded module, as in splitting a Wasm module for lazy loading.
Give hot code a clear home. A kernel that does the heavy work should be a few functions of moderate size, called repeatedly, so engines identify it as hot early and optimize it quickly. Code generators that emit one enormous function — some parser generators, unrolled lookup tables, interpreters compiled into a single dispatch function — deserve a look: splitting their output into smaller functions often shortens time-to-optimized-code dramatically.
Separate one-time setup from repeated work. Initialisation that runs once will run on baseline code in every engine; that is fine if it is cheap, and a reason to precompute expensive setup at build time if it is not. Repeated work, by contrast, benefits from tier-up and from warm-up during idle time.
Finally, measure in the engines and on the devices that matter. The same module can have a fast first second in one engine and a slower one in another, for reasons entirely outside its code. Real-user timing segmented by browser, as in measuring Wasm performance with real-user monitoring, shows where the differences actually affect people.
Tiering beyond the browser
Server and edge runtimes face the same trade-off with different constraints, and they resolve it differently. A long-running server compiles a module once and runs it for hours, so compile time is amortised and full optimization up front is the obvious choice — which is what wasmtime with Cranelift does. Serverless and edge platforms sit in between: they may instantiate a module for every request, so instantiation must take microseconds, but compilation can happen once per deployment. They precompile modules when they are deployed and cache the machine code, so a request never pays for compilation at all, as described in cold start characteristics of server-side Wasm. Some runtimes also offer interpreters or fast single-pass compilers for workloads where modules are many, small and short-lived — plugin systems that load user code on demand, for example — where optimization would cost more than it saves. The browser’s tiers are one point on a spectrum; the right point depends on how long code lives after it is loaded.
Gotchas and failure modes
- Benchmarks that disagree between runs. Tier-up happened at different moments. Warm up longer or pin a tier.
- Cross-engine comparisons of short benchmarks. They compare tiering strategies, not code quality. State which phase was measured.
- A slow first interaction. Lazy compilation and baseline code land on the user’s first call. Warm up during idle time.
- A hot function that never speeds up. It is huge and still being optimized, or it is not actually hot in the real page. Check a trace.
- Assuming server numbers apply to browsers. wasmtime compiles everything optimized up front; browsers do not.
- Relying on V8 flag names. Diagnostic flags change between releases. Check them before scripting.
Reading engine changes over time
Engines update every few weeks, and their WebAssembly compilers change with them: new tiers, different thresholds, better code for particular instruction patterns. Most changes are invisible improvements, but occasionally a browser release shifts a module’s startup or steady-state numbers noticeably. Keeping a small, pinned benchmark and a startup measurement in CI — run against the current browser versions — turns those shifts into data rather than user reports, and makes it clear whether a change came from your code or from the engine.
Verification
To confirm your understanding of how a specific module behaves, produce four numbers on a representative device: baseline compile time, first-call time, per-call time before tier-up, and per-call time after tier-up. The first two say how quickly the module becomes useful; the last two say how much tier-up buys and therefore how much warm-up matters. Record them in each engine you support and compare against the user-facing interaction budget. When they change unexpectedly between releases of your module, compare module size and function sizes first — those are the inputs engines’ tiering responds to. For benchmarks of code changes, use pinned tiers as in pinning a compiler tier for benchmarks, so the comparison reflects the code rather than the timing of tier-up.
Guides in this topic
- How V8 compiles Wasm with Liftoff and TurboFan — the baseline and optimizing tiers in Chrome and Node, lazy compilation and dynamic tier-up.
- How SpiderMonkey compiles Wasm — Firefox’s baseline compiler, Ion, and its eager background optimization.
- How Safari runs Wasm — JavaScriptCore’s interpreter, BBQ and OMG tiers, and what they mean on iPhones.
- Watching Wasm tier-up in Chrome — per-call timing, traces and V8 flags that show which functions get optimized.
- Why the first call into Wasm is slow — the one-time costs behind slow first calls and how to move them off the user’s path.
- Pinning a compiler tier for benchmarks — flags that force one tier so builds can be compared fairly.
Frequently Asked Questions
Is WebAssembly interpreted? Mostly not. Engines compile it to machine code, quickly at first and optimized later. JavaScriptCore does start functions in an interpreter, but hot code moves to compiled tiers quickly.
Why not ship machine code instead of bytecode? Machine code is specific to a CPU and operating system, and running untrusted machine code would bypass the sandbox. Bytecode that the engine validates and compiles itself keeps WebAssembly portable and safe.
Does tiering affect correctness? No. Every tier implements exactly the same semantics; only speed differs. WebAssembly never deoptimizes, so there is no behavioural change when code tiers up.
Can I precompile modules for browsers? Not into the browser’s machine code. Browsers cache their own compiled code for modules from stable URLs, which gives much of the same benefit. Server runtimes such as wasmtime can precompile to a file.
Do Web Workers share compiled code?
Within one browser process, compiled code for a module can be shared between instances, including those in workers, especially when the compiled
WebAssembly.Module is posted rather than recompiled from bytes; see
instantiating one module many times.
Does SIMD change tiering? No. Both baseline and optimizing tiers support SIMD instructions; optimized SIMD code benefits most from register allocation, so warm-up matters for SIMD kernels as for any other hot code.
How do I make my module tier up faster? Keep hot code in a moderate number of functions of reasonable size, warm up hot paths during idle time, and keep the module small so baseline compilation finishes quickly.
Related
- WebAssembly Core Concepts & Browser Runtime — the section overview this topic belongs to.
- Module Caching & Startup Performance — reducing what tiering has to do.
- Wasm Performance Benchmarking — measuring with tiering in mind.
- Stack vs Heap Execution Model — what the compilers do with locals and the stack.