Comparing Wasm Performance Across Browsers Fairly

This page answers one task: you need to know how a WebAssembly feature performs in each major browser — to set expectations, to choose a build, or to file a bug with an engine team — and you want the comparison to reflect the engines rather than accidents of the measurement.

Prerequisites

  • [ ] A benchmark harness that runs in the browser and reports per-iteration timings.
  • [ ] Current stable Chrome, Firefox and Safari on the same machine (Safari requires macOS), or Playwright’s engines for automation.
  • [ ] A production build of the module and representative inputs.

Why cross-browser comparisons go wrong

The three engines — V8 in Chrome and Edge, SpiderMonkey in Firefox, JavaScriptCore in Safari — compile WebAssembly differently: they use different baseline compilers, different optimising compilers, different tier-up heuristics, and different strategies for bounds checks, SIMD lowering and calls between JavaScript and Wasm. Those are real differences worth measuring. But naive comparisons mostly measure other things: one browser’s timer is coarser, one ran the benchmark before tier-up finished, one did not have threads because isolation headers were missing, one used a different code path because feature detection selected a fallback, or one was in the background tab. Each of these produces differences of 2× or more that have nothing to do with engine quality.

A fair comparison holds everything constant except the engine: same machine and power state, same build, same inputs, same warm-up rules, same features enabled, and the same definition of what is measured.

Variables to hold constant across browsers Use the same machine and power state, the same Wasm build and inputs, identical warm-up and iteration rules, the same enabled features including threads and SIMD, and check timer resolution in each browser before trusting small differences. variable why it matters how to control machine + power throttling, core choice same device, plugged in build + inputs different code paths one artefact for all warm-up rules tier-up timing differs fixed warm-up + check plateau features enabled threads, SIMD fallbacks assert detected features timer resolution coarse timers distort short runs time batches

Step 1 — use one build and assert the code path

Ship the same .wasm to every browser, and make the benchmark page report which features the loader detected and which build variant it loaded. If the loader picks a SIMD build in one browser and a scalar build in another, the comparison is between builds, not engines — which may be the question you want answered, but should be labelled as such:

const features = { simd: await simd(), threads: await threads(), isolated: crossOriginIsolated };
const variant = features.simd ? "simd" : "scalar";
const mod = await load(`/bench/app.${variant}.wasm`);
report({ ua: navigator.userAgent, features, variant });

Serve the page with the production headers, including COOP and COEP if the module uses threads, and confirm crossOriginIsolated in each browser.

Step 2 — apply identical warm-up and iteration rules

Engines tier up at different times: one may optimise a hot function after a few calls, another after many more, and some compile optimised code in the background while the baseline version keeps running. Use a warm-up that runs until timings plateau, then measure:

async function measure(fn, { minWarmup = 20, maxWarmup = 500, runs = 30 } = {}) {
  let last = Infinity;
  for (let i = 0; i < maxWarmup; i++) {
    const t0 = performance.now(); fn(); const t = performance.now() - t0;
    if (i >= minWarmup && Math.abs(t - last) / last < 0.02) break;   // plateau reached
    last = t;
    if (i % 25 === 0) await new Promise((r) => setTimeout(r, 0));     // let background tier-up finish
  }
  const times = [];
  for (let i = 0; i < runs; i++) { const t0 = performance.now(); fn(); times.push(performance.now() - t0); }
  return times.sort((a, b) => a - b)[runs >> 1];
}

Yielding to the event loop during warm-up gives background compilation a chance to install optimised code. Record the number of warm-up iterations each browser needed — it is itself a useful result about tier-up behaviour.

Step 3 — handle timer resolution

performance.now() is coarsened in browsers to mitigate timing attacks, and the coarsening differs: cross-origin isolated pages get finer resolution (around 5 µs in Chrome), non-isolated pages coarser (100 µs or more), and browsers have changed these values over time. A workload that takes 50 µs cannot be timed reliably per call. Time batches instead — run the function enough times that each measurement lasts at least tens of milliseconds — and divide. Check the effective resolution in each browser by measuring the smallest non-zero difference between consecutive performance.now() calls.

Timing a short function per call versus in batches Timing one 50 microsecond call with a 100 microsecond timer returns either zero or one tick, which is meaningless. Timing a batch of a thousand calls takes about 50 milliseconds, far above timer resolution, and dividing gives an accurate per-call figure. per-call timing 50 µs call, 100 µs timer results are 0 or 100 µs differences meaningless misleading batched timing 1,000 calls per sample ~50 ms per sample per-call = sample / 1,000 reliable

Step 4 — separate startup from throughput

Report startup (compile and instantiate) and steady-state throughput separately. Engines that compile quickly with a fast baseline tier may start sooner and run slower; engines with eager optimisation may start later and run faster. Combining them into one number hides the trade-off, and which matters depends on the feature: a one-off conversion is dominated by startup, a long-running editor by throughput. Startup measurement is described in measuring Wasm startup time end to end.

Step 5 — repeat, interleave and report spread

Run each browser several times, alternating between them rather than finishing one before starting the next, so slow drift in machine state — thermal throttling, background updates — affects all browsers equally. Close other applications, keep the benchmark tab in the foreground, and disable extensions (a clean profile helps). Report medians with the spread, and do not claim a difference smaller than the spread.

Reading the results

A difference that survives these controls is real, but may still be specific to the workload. Engines differ in particular operations: call overhead between JavaScript and Wasm, memory.grow cost, SIMD lowering of specific instructions, bounds-check elimination, handling of i64 at the JavaScript boundary, and so on. When one browser is much slower on a workload, profile it in that browser to find which operation dominates — Firefox Profiler and Safari’s Web Inspector both show Wasm frames — and check whether a small code change removes the gap. If a gap looks like an engine performance bug, engine teams accept reduced test cases: a minimal page with the module and the harness that reproduces the difference.

Automating with Playwright

Playwright drives Chromium, Firefox and WebKit, which makes nightly comparisons possible. Its engines track the browsers closely but are not identical to stable Chrome, Firefox and Safari releases, and WebKit on Linux differs from Safari on macOS in platform-specific parts. Use Playwright for trends and regressions, and the real browsers on a Mac for published comparisons.

Comparing against JavaScript in each browser

A cross-browser Wasm comparison is more informative alongside a JavaScript baseline in the same browsers. Engines optimise JavaScript differently too, and the ratio of Wasm to JavaScript performance varies by engine: a workload might be 3× faster in Wasm in one browser and 1.5× in another, either because the second browser’s Wasm is slower or because its JavaScript is faster. Include the JavaScript version of the workload in the same harness, with the same warm-up rules, and report both absolute times and the ratio. Decisions such as “is porting this to Wasm worth it for Safari users?” depend on the ratio in that browser, not on the Chrome result.

Publishing comparisons responsibly

Cross-browser numbers travel. A chart shared internally ends up in a slide deck or a blog post, stripped of its caveats, and becomes “browser X is slow at WebAssembly”. Publish the full method with any comparison: browser versions, operating system, hardware, build flags, features enabled, harness code and inputs. Prefer describing specific operations (“call overhead for many small calls is higher in browser X”) over general claims. Re-run before reusing old numbers, since engines change every few weeks, and note the date prominently.

Expected output

A table with startup and throughput for three workloads in three browsers, each with median, spread, warm-up iterations needed, detected features and build variant; one workload where Safari is 1.6× slower, traced to call overhead on a chatty JS↔Wasm boundary; and a reduced test case for that workload attached to an engine bug report.

Gotchas

  • Different builds per browser. Feature detection may load different variants. Report and control the variant.
  • Fixed short warm-up. Engines tier up at different times. Warm up until timings plateau.
  • Per-call timing of short functions. Timer resolution dominates. Time batches.
  • Background tabs and extensions. Throttling and interference. Foreground, clean profile.
  • One combined number. Startup and throughput trade off differently per engine. Report both.

Performance note

On the image workload, the median per-frame time varied by 41% between the first and fifth complete runs in one browser when warm-up was fixed at 10 iterations, and by 3% when warm-up ran to a plateau — the fixed warm-up had been measuring tier-up, not steady state.

Run-to-run variation with fixed versus plateau warm-up Percentage variation of the median per-frame time across five runs in one browser with a fixed ten-iteration warm-up and with warm-up continued until timings plateau. variation across runs (%) fixed 10-iteration warm-up 41 % warm-up to plateau 3 %

Frequently Asked Questions

Should I include Edge separately? It uses the same engine as Chrome; include it only if you suspect configuration differences.

Can I compare mobile browsers the same way? Yes, with the extra controls for thermal state and power described in the mobile benchmarking guide.

Is it fair to compare with threads in one browser only? Only if labelled as such; otherwise disable threads everywhere for the comparison.

How often should comparisons be refreshed? Engines release every four to six weeks; refresh before decisions that depend on them.

Should I report Wasm alone or against JavaScript? Both — the Wasm-to-JavaScript ratio per browser is what decides whether Wasm is worth it for that browser’s users.

← Back to Wasm Performance Benchmarking