Avoiding JIT Warm-up Errors in Wasm Benchmarks

This page answers one question: why do the first measurements of a WebAssembly function differ so much from later ones, and how do you structure a benchmark so the number you report is the one users actually experience?

Prerequisites

  • [ ] A function to benchmark, exported from a module and callable in a loop.
  • [ ] Chrome, Firefox or Node — all use tiered compilation for WebAssembly.
  • [ ] A rough sense of how long one call takes: microseconds, milliseconds or seconds.

Two compilers, two speeds

Browsers do not compile a WebAssembly module once. To start quickly, they first compile every function with a fast baseline compiler — Liftoff in V8, the baseline compiler in SpiderMonkey, BBQ in JavaScriptCore — that produces correct but unoptimized machine code in a single pass. Then, in the background or when a function proves hot, they compile it again with an optimizing compiler — TurboFan, Ion, OMG — and swap the faster code in. The engine-level detail is in how V8 compiles Wasm with Liftoff and TurboFan.

For a benchmark, this means the same function has two very different speeds depending on when you call it. Baseline code is commonly two to five times slower than optimized code for tight numeric loops. A benchmark that calls the function once, or averages the first few calls with later ones, reports a number that corresponds to neither — and a comparison between two implementations can come out backwards if one of them happened to tier up sooner.

Per-call time of one function over its first calls A function's per-call time across its first two hundred calls in Chrome. The first calls run baseline code, a spike appears around the swap, and from roughly call 40 onwards the optimized code runs at a stable, much lower time. 1 calls first call 9.8 ms 12 calls baseline steady 4.1 ms 38 calls optimized code swapped in 60 calls stable 1.3 ms 200 calls still 1.3 ms

Step 1 — warm up before you time

The simplest fix is to call the function enough times before measuring that the engine has optimized it. How many is “enough” depends on the engine and on how long each call takes, so make the warm-up time-based rather than count-based:

function warmUp(fn, minMs = 500) {
  const end = performance.now() + minMs;
  let n = 0;
  while (performance.now() < end) { fn(); n++; }
  return n;
}

function measure(fn, iterations = 200) {
  const samples = new Float64Array(iterations);
  for (let i = 0; i < iterations; i++) {
    const t0 = performance.now();
    fn();
    samples[i] = performance.now() - t0;
  }
  samples.sort();
  return { median: samples[iterations >> 1], p90: samples[Math.floor(iterations * 0.9)] };
}

warmUp(() => exports.blur(ptr, w, h));
console.log(measure(() => exports.blur(ptr, w, h)));

Half a second of warm-up is a reasonable default for functions that take microseconds to milliseconds. For functions that take hundreds of milliseconds per call, warm up by count — five or ten calls — since half a second may be a single call.

Step 2 — report the median and a high percentile

Even after warm-up, individual samples vary: garbage collection in the JavaScript around the call, other tabs, thermal throttling, timer resolution. Report the median, which ignores outliers, and a high percentile such as p90, which shows whether the outliers are frequent. Never report the minimum alone; it rewards lucky samples and hides variance. Never report the mean of all samples including warm-up; it is dominated by the slow early calls.

Timer resolution deserves a mention. Browsers coarsen performance.now() for security — to 100 µs or more in some contexts unless the page is cross-origin isolated, where it is 5 µs. A function that takes 20 µs cannot be timed meaningfully one call at a time at that resolution. Time a batch of calls and divide:

function measureBatched(fn, batch = 1000, rounds = 50) {
  const per = [];
  for (let r = 0; r < rounds; r++) {
    const t0 = performance.now();
    for (let i = 0; i < batch; i++) fn();
    per.push((performance.now() - t0) / batch);
  }
  per.sort((a, b) => a - b);
  return per[rounds >> 1];
}

Step 3 — confirm the optimized tier is running

Warm-up is an assumption until verified. In V8, flags can force a tier so you can see the difference directly — useful for confirming that your steady-state number matches optimized code:

# Node: baseline only (Liftoff), then optimized only (TurboFan)
node --liftoff --no-wasm-tier-up bench.mjs
node --no-liftoff bench.mjs

If the steady-state number from a normal run matches the --no-liftoff run, the warm-up worked. If it sits between the two, the function had not finished tiering up — warm up longer. In Chrome, the Performance panel shows background compile tasks named v8.wasm.compileTopTier or similar; a measurement that overlaps them is measuring a moving target. Pinning tiers for a whole benchmark is described in pinning a compiler tier for benchmarks.

The same function measured four ways Reported time for one image kernel depending on benchmark method. Timing the first call or averaging including warm-up overstates the cost; the median after warm-up matches the optimized-only run. reported ms per call first call only 9.8 ms mean of first 50 calls 3.6 ms median after 500 ms warm-up 1.3 ms --no-liftoff (optimized only) 1.3 ms

Step 4 — compare like with like

When comparing two implementations — a Wasm kernel against a JavaScript one, or two Wasm builds — give each its own warm-up and measure them in an interleaved order rather than one after the other. Running A for a minute and then B for a minute lets thermal throttling or background activity favour whichever ran first. Interleaving spreads that noise evenly:

const results = { a: [], b: [] };
warmUp(implA); warmUp(implB);
for (let round = 0; round < 30; round++) {
  results.a.push(measureBatched(implA, 100, 1));
  results.b.push(measureBatched(implB, 100, 1));
}

JavaScript has the same tiering problem — V8’s Ignition, Sparkplug, Maglev and TurboFan tiers — so a JavaScript baseline needs exactly the same warm-up treatment. Comparisons that warm up the Wasm version and not the JavaScript one exaggerate Wasm’s advantage; measuring Wasm vs JavaScript throughput covers that comparison in full.

Step 5 — decide whether cold performance is the real question

Warm-up is the right correction when users call the function many times: an editor applying filters, a game simulating every frame. It is the wrong correction when users call it once: a page that formats one document on load, a serverless function that handles one request per instance. For those, the cold number — first call, including compilation — is the user’s experience, and you should measure it deliberately, in a fresh page or process each time, rather than warming it away.

Report both when it matters: “first call 9.8 ms, steady state 1.3 ms” tells a reader far more than either number alone, and it points to different fixes. A slow first call is improved by smaller modules, caching compiled code, and compiling early; a slow steady state is improved by better code.

Warm and cold measurements answer different questions A steady-state measurement after warm-up describes repeated use such as an editor or a game loop. A cold measurement in a fresh page or process describes one-shot use such as a page-load task or a serverless request. steady state (after warm-up) optimized tier, caches warm median of many calls improved by better code and SIMD editors, games, repeated work cold (fresh page or process) includes compile and baseline code one call per fresh context improved by size, caching, early compile page-load tasks, serverless requests

Expected output

A harness that follows these steps prints both views:

blur 1920x1080   cold: 9.8 ms   steady: median 1.31 ms  p90 1.42 ms  (warm-up 500 ms, 200 samples)
blur_js          cold: 41.2 ms  steady: median 4.87 ms  p90 5.30 ms

Gotchas

  • Benchmark results flip between runs. One implementation tiered up during measurement in one run and before it in another. Lengthen the warm-up and interleave.
  • DevTools open changes the numbers. An open DevTools window can disable some optimizations and adds overhead to timers. Close it, or run in a headless browser.
  • The optimizer removes the work. A call whose result is unused can, in principle, be optimized away. Accumulate results into a value you print at the end.
  • Measuring in a background tab. Browsers throttle timers and scheduling in hidden tabs. Keep the benchmark tab in front, or run in a worker.

Performance note

On the kernel above, V8 needed about 40 calls — roughly 200 ms — before the optimized code took over, while SpiderMonkey tiered up within the first 10 calls because it compiles optimized code eagerly in the background for most modules. Engine differences like that are why warm-up should be time-based and verified, not a fixed count copied from another benchmark.

Frequently Asked Questions

How long should a benchmark run in total? Long enough that the median stops moving when you add rounds — typically a few seconds per variant for millisecond-scale functions.

Does caching compiled code remove the need for warm-up? It removes compilation time on later page loads, but the cached code may still be baseline code for functions that had not tiered up when it was cached. Measure warm and cold separately.

Can I just use a benchmarking library? Libraries such as tinybench and mitata handle warm-up and statistics for you. Check that their defaults suit Wasm’s longer tier-up and that they report medians.

Does running in a worker change warm-up? Workers have their own instance of the module but share the engine’s compiled code within a process, so tier-up in one worker can benefit another. Warm up in the same worker you measure in, and do not assume a fresh worker starts cold.

Is the same true for WASI runtimes? Wasmtime compiles ahead of time with Cranelift, so there is no tier-up within a run, but the first instantiation still pays for compilation unless a precompiled module is used.

← Back to Wasm Performance Benchmarking