Avoiding JIT Warm-up Errors in Wasm Benchmarks
This page answers one question: why do the first measurements of a WebAssembly function differ so much from later ones, and how do you structure a benchmark so the number you report is the one users actually experience?
Prerequisites
- [ ] A function to benchmark, exported from a module and callable in a loop.
- [ ] Chrome, Firefox or Node — all use tiered compilation for WebAssembly.
- [ ] A rough sense of how long one call takes: microseconds, milliseconds or seconds.
Two compilers, two speeds
Browsers do not compile a WebAssembly module once. To start quickly, they first compile every function with a fast baseline compiler — Liftoff in V8, the baseline compiler in SpiderMonkey, BBQ in JavaScriptCore — that produces correct but unoptimized machine code in a single pass. Then, in the background or when a function proves hot, they compile it again with an optimizing compiler — TurboFan, Ion, OMG — and swap the faster code in. The engine-level detail is in how V8 compiles Wasm with Liftoff and TurboFan.
For a benchmark, this means the same function has two very different speeds depending on when you call it. Baseline code is commonly two to five times slower than optimized code for tight numeric loops. A benchmark that calls the function once, or averages the first few calls with later ones, reports a number that corresponds to neither — and a comparison between two implementations can come out backwards if one of them happened to tier up sooner.
Step 1 — warm up before you time
The simplest fix is to call the function enough times before measuring that the engine has optimized it. How many is “enough” depends on the engine and on how long each call takes, so make the warm-up time-based rather than count-based:
function warmUp(fn, minMs = 500) {
const end = performance.now() + minMs;
let n = 0;
while (performance.now() < end) { fn(); n++; }
return n;
}
function measure(fn, iterations = 200) {
const samples = new Float64Array(iterations);
for (let i = 0; i < iterations; i++) {
const t0 = performance.now();
fn();
samples[i] = performance.now() - t0;
}
samples.sort();
return { median: samples[iterations >> 1], p90: samples[Math.floor(iterations * 0.9)] };
}
warmUp(() => exports.blur(ptr, w, h));
console.log(measure(() => exports.blur(ptr, w, h)));
Half a second of warm-up is a reasonable default for functions that take microseconds to milliseconds. For functions that take hundreds of milliseconds per call, warm up by count — five or ten calls — since half a second may be a single call.
Step 2 — report the median and a high percentile
Even after warm-up, individual samples vary: garbage collection in the JavaScript around the call, other tabs, thermal throttling, timer resolution. Report the median, which ignores outliers, and a high percentile such as p90, which shows whether the outliers are frequent. Never report the minimum alone; it rewards lucky samples and hides variance. Never report the mean of all samples including warm-up; it is dominated by the slow early calls.
Timer resolution deserves a mention. Browsers coarsen performance.now() for security — to 100 µs or more in some contexts
unless the page is cross-origin isolated, where it is 5 µs. A function that takes 20 µs cannot be timed meaningfully one call
at a time at that resolution. Time a batch of calls and divide:
function measureBatched(fn, batch = 1000, rounds = 50) {
const per = [];
for (let r = 0; r < rounds; r++) {
const t0 = performance.now();
for (let i = 0; i < batch; i++) fn();
per.push((performance.now() - t0) / batch);
}
per.sort((a, b) => a - b);
return per[rounds >> 1];
}
Step 3 — confirm the optimized tier is running
Warm-up is an assumption until verified. In V8, flags can force a tier so you can see the difference directly — useful for confirming that your steady-state number matches optimized code:
# Node: baseline only (Liftoff), then optimized only (TurboFan)
node --liftoff --no-wasm-tier-up bench.mjs
node --no-liftoff bench.mjs
If the steady-state number from a normal run matches the --no-liftoff run, the warm-up worked. If it sits between the two,
the function had not finished tiering up — warm up longer. In Chrome, the Performance panel shows background compile tasks
named v8.wasm.compileTopTier or similar; a measurement that overlaps them is measuring a moving target. Pinning tiers for a
whole benchmark is described in
pinning a compiler tier for benchmarks.
Step 4 — compare like with like
When comparing two implementations — a Wasm kernel against a JavaScript one, or two Wasm builds — give each its own warm-up and measure them in an interleaved order rather than one after the other. Running A for a minute and then B for a minute lets thermal throttling or background activity favour whichever ran first. Interleaving spreads that noise evenly:
const results = { a: [], b: [] };
warmUp(implA); warmUp(implB);
for (let round = 0; round < 30; round++) {
results.a.push(measureBatched(implA, 100, 1));
results.b.push(measureBatched(implB, 100, 1));
}
JavaScript has the same tiering problem — V8’s Ignition, Sparkplug, Maglev and TurboFan tiers — so a JavaScript baseline needs exactly the same warm-up treatment. Comparisons that warm up the Wasm version and not the JavaScript one exaggerate Wasm’s advantage; measuring Wasm vs JavaScript throughput covers that comparison in full.
Step 5 — decide whether cold performance is the real question
Warm-up is the right correction when users call the function many times: an editor applying filters, a game simulating every frame. It is the wrong correction when users call it once: a page that formats one document on load, a serverless function that handles one request per instance. For those, the cold number — first call, including compilation — is the user’s experience, and you should measure it deliberately, in a fresh page or process each time, rather than warming it away.
Report both when it matters: “first call 9.8 ms, steady state 1.3 ms” tells a reader far more than either number alone, and it points to different fixes. A slow first call is improved by smaller modules, caching compiled code, and compiling early; a slow steady state is improved by better code.
Expected output
A harness that follows these steps prints both views:
blur 1920x1080 cold: 9.8 ms steady: median 1.31 ms p90 1.42 ms (warm-up 500 ms, 200 samples)
blur_js cold: 41.2 ms steady: median 4.87 ms p90 5.30 ms
Gotchas
- Benchmark results flip between runs. One implementation tiered up during measurement in one run and before it in another. Lengthen the warm-up and interleave.
- DevTools open changes the numbers. An open DevTools window can disable some optimizations and adds overhead to timers. Close it, or run in a headless browser.
- The optimizer removes the work. A call whose result is unused can, in principle, be optimized away. Accumulate results into a value you print at the end.
- Measuring in a background tab. Browsers throttle timers and scheduling in hidden tabs. Keep the benchmark tab in front, or run in a worker.
Performance note
On the kernel above, V8 needed about 40 calls — roughly 200 ms — before the optimized code took over, while SpiderMonkey tiered up within the first 10 calls because it compiles optimized code eagerly in the background for most modules. Engine differences like that are why warm-up should be time-based and verified, not a fixed count copied from another benchmark.
Frequently Asked Questions
How long should a benchmark run in total? Long enough that the median stops moving when you add rounds — typically a few seconds per variant for millisecond-scale functions.
Does caching compiled code remove the need for warm-up? It removes compilation time on later page loads, but the cached code may still be baseline code for functions that had not tiered up when it was cached. Measure warm and cold separately.
Can I just use a benchmarking library? Libraries such as tinybench and mitata handle warm-up and statistics for you. Check that their defaults suit Wasm’s longer tier-up and that they report medians.
Does running in a worker change warm-up? Workers have their own instance of the module but share the engine’s compiled code within a process, so tier-up in one worker can benefit another. Warm up in the same worker you measure in, and do not assume a fresh worker starts cold.
Is the same true for WASI runtimes? Wasmtime compiles ahead of time with Cranelift, so there is no tier-up within a run, but the first instantiation still pays for compilation unless a precompiled module is used.
Related
- Building a reproducible Wasm benchmark harness — a harness that applies these rules.
- Why the first call into Wasm is slow — the cold side in detail.
- Watching Wasm tier-up in Chrome — observing the swap.
- Measuring JS-to-Wasm call overhead — where batching matters most.
← Back to Wasm Performance Benchmarking