Watching Wasm Tier-up in Chrome
This page answers one task: observe V8’s tier-up for your module — which functions get optimized, how long that takes, and when the faster code starts running — so you can tell whether a slow phase is baseline code, compilation, or your own logic.
Prerequisites
- [ ] Chrome (or another Chromium browser) with DevTools, and Node for flag-based traces.
- [ ] A module with its
namesection kept, so functions are identifiable. - [ ] The background in how V8 compiles Wasm with Liftoff and TurboFan.
Why tier-up is worth seeing
Tier-up is invisible in normal use: a function simply gets faster after a while. It becomes visible — and worth understanding — in three situations. A benchmark that gives different answers on different runs is often catching functions mid-tier-up. A page that feels sluggish for its first second and then fine is often running hot code on Liftoff while TurboFan works in the background. And a function that never seems to reach the speed you measured in isolation may never be tiering up at all, because it is not called often enough in the real page, or because it is so large that optimization takes a long time.
In each case the question is the same — what tier is this code running in, right now? — and Chrome offers several ways to answer it, from a quick look at compile tasks in a trace to a function-by-function log of tier-up decisions.
Step 1 — see the effect in per-call timing
The simplest observation needs no special tools: time each call of the hot function and plot or print the series.
const times = [];
for (let i = 0; i < 300; i++) {
const t0 = performance.now();
exports.blur_rows(ptr, width, height);
times.push(performance.now() - t0);
}
console.log(times.filter((_, i) => i % 20 === 0).map((t) => t.toFixed(2)).join(" "));
4.10 4.05 4.08 1.31 1.29 1.30 1.28 1.30 1.29 1.31 1.28 1.30 1.29 1.30 1.29
The drop between the third and fourth samples — around call 50 — is the optimized code taking over. Before it, the function ran on Liftoff code; after it, on TurboFan code. The per-call ratio, about 3× here, is the benefit of tier-up for this function.
Step 2 — find compile tasks in the Performance panel
Record a trace while the page loads and runs the hot path. In the Performance panel, expand the thread tracks below Main: V8’s background compile work appears on worker threads, with slices for WebAssembly compilation. Zoom into the period where the per-call time dropped; you should find a TurboFan compile task for the function — or a batch of functions — finishing just before.
Clicking a compile slice shows its duration. Long optimization tasks for very large functions are worth noting: a function that takes 300 ms to optimize runs on baseline code for those 300 ms, whatever its call count. Splitting oversized generated functions — giant interpreters, unrolled tables — can bring their optimized code online much sooner. Reading these traces in general is covered in profiling Wasm with the Chrome Performance panel.
Step 3 — use chrome://tracing for detailed events
For a detailed view, record with Chrome’s tracing tool. Open chrome://tracing (or the newer Perfetto UI at ui.perfetto.dev with a Chrome
trace), start a recording with the v8.wasm and v8.wasm.detailed categories enabled, exercise the page, and stop. Search the result for compile
events; each records the function index or name and the compilation tier.
v8.wasm.compileLiftoff func #412 0.18 ms
v8.wasm.compileTopTier func #412 6.20 ms (background)
v8.wasm.compileTopTier func #87 3.95 ms (background)
Category and event names have changed across Chrome versions, so search for “wasm” if the exact names differ. The detailed categories add overhead; use them for investigation, not for timing.
Step 4 — log decisions by function in Node
Node exposes V8’s tracing flags directly, which makes it the easiest place to get a per-function log of tier-up decisions:
node --trace-wasm-compilation-times --trace-wasm-tier-up bench.mjs 2>&1 | grep -E 'tier|compil' | head -20
The output names each function (by index, or by name with a name section) and when it was compiled at each tier. Because Node and Chrome share
V8, the decisions are representative of the browser for the same workload, minus the page’s other activity. node --v8-options | grep wasm lists
the flags available in your Node version; they are diagnostic flags and do change between releases.
Step 5 — act on what you see
Most tier-up observations lead to one of four actions. If hot code is optimized quickly and the slow phase is short, do nothing — the page simply warms up. If a benchmark’s results vary, add warm-up until timings are stable, as described in avoiding JIT warm-up errors in Wasm benchmarks. If an important function optimizes slowly because it is huge, split it. And if the slow first second matters to users, consider precomputing or deferring the heaviest initial work, or warming the code path with a small input before the real one — a technique discussed in why the first call into Wasm is slow.
How tier-up interacts with code caching
Tier-up observations look different on a returning visit. When Chrome has cached compiled code for a module — which it does for modules loaded with streaming APIs from cacheable URLs, after they have run for a while — the cache contains optimized code for functions that were hot in previous sessions. Those functions start in their optimized form on the next visit, and the per-call drop from Step 1 disappears: the first calls are already fast. That is why a performance investigation should state whether it measured a cold or a returning visit, and why clearing the cache, or using a fresh profile, is part of reproducing first-visit behaviour reliably.
Keep in mind that every observation tool here perturbs what it observes a little: tracing categories add overhead, DevTools can change tiering
decisions, and per-call timing adds performance.now() calls. Use the lightest tool that answers the question, and confirm important findings
with a second method.
Expected output
Per-call timing shows a clear step down after a few dozen calls; the trace shows a TurboFan compile task for the function just before the step; and the Node log names the function as tiered up at about the same call count.
Gotchas
- No step in the timing. The function was already optimized from the code cache, or it was optimized before measurement started. Use a fresh profile or a cold cache.
- Step never appears. The function is not called enough to tier up in the measured window, or it is tiny and inlined into its caller. Check the caller.
- Different results with DevTools open. DevTools can change tiering behaviour. Confirm findings with DevTools closed, using timing logs.
- Function indices instead of names. The module was stripped, so traces show
func #412. Map indices with an unstripped build of the same commit. - Flags that no longer exist. V8 renames diagnostic flags. List current flags before scripting them.
Performance note
In the image filter above, the hottest function tiered up after about 50 calls — roughly 200 ms into the workload — and its optimized compile took 6 ms on a background thread. Over a ten-second session the warm-up period accounted for 2% of the total time; for a one-second task it accounted for 20%, which is why short tasks benefit most from warm-up strategies and caching.
Frequently Asked Questions
Why do some functions tier up much later than others? Tier-up is driven by how much work a function does, counted on calls and loop iterations. A function called rarely but with long loops can tier up quickly; one called often but doing little may take longer.
Can I force a function to tier up early? Not directly from a page. Calling it with a small input during idle time warms it up so the optimized code is ready when real work starts.
Do workers tier up independently? Compiled code belongs to the module, so optimized code produced from one worker’s activity can benefit other instances in the same process.
Is tier-up the same in Node and Chrome? The mechanism is the same V8 pipeline. Page activity, DevTools and caching differences mean exact timings differ.
Can I see tier-up in a production session? Not directly — the tools here are for development. In production, per-call timings collected through real-user monitoring show the same step down in aggregate.
Does tier-up ever make code slower? No — optimized code replaces baseline code only when it is ready, and WebAssembly never deoptimizes.
Related
- Pinning a compiler tier for benchmarks — removing tier-up from measurements.
- Measuring Wasm compile time in DevTools — the baseline compile side.
- How SpiderMonkey compiles Wasm — a different strategy to compare.
- Generating flame graphs for Wasm — seeing where optimized time goes.
← Back to Engine Tiering & JIT Compilation