Measuring Wasm Performance with Real-User Monitoring

This page answers one task: a feature relies on WebAssembly, and you need to know how long it takes for real users — on their devices and networks — to download, compile and run it, so you can find regressions and decide where optimisation will pay off.

Prerequisites

  • [ ] A Wasm module loaded through code you control (your own loader or a wrapper around generated glue).
  • [ ] A real-user monitoring (RUM) endpoint or analytics pipeline that accepts custom metrics.
  • [ ] Familiarity with performance.mark, performance.measure and PerformanceObserver.

Lab numbers are not user numbers

Benchmarks on a developer laptop say little about what users experience. Compilation time scales with CPU speed, and the median phone is several times slower than a developer machine. Download time depends on network and caching. Engines tier up differently, so a function that is fast after warm-up may be slow on its first call — which is the call users feel. Only measuring in the field shows the actual distribution, and the tail of that distribution — the 75th and 95th percentile — is usually where the problems are.

Wasm work has four phases worth measuring separately, because each has different fixes: download (size, compression, caching), compile (size, streaming compilation, code caching), instantiate (start functions, memory initialisation), and execution of the operations users wait for. Measuring only the total hides which one to fix.

The phases of a Wasm feature's first use A user opens a feature. The .wasm download starts, streaming compilation overlaps it, instantiation follows, and then the first call runs, initially in a baseline tier. Each phase is measured with marks so the field data shows which one dominates. 0 ms feature opened 60 ms download starts 420 ms download + compile done 470 ms instantiated 610 ms first call returns 820 ms later calls faster

Step 1 — mark each phase

Wrap the loader with marks and measures:

performance.mark("wasm:start");
const response = fetch(new URL("./engine.wasm", import.meta.url));
const module = await WebAssembly.compileStreaming(response);           // download + compile overlap
performance.mark("wasm:compiled");
const instance = await WebAssembly.instantiate(module, imports);
performance.mark("wasm:instantiated");

performance.measure("wasm-load", "wasm:start", "wasm:compiled");
performance.measure("wasm-instantiate", "wasm:compiled", "wasm:instantiated");

With streaming compilation, download and compile overlap, so measure them together, and use Resource Timing (step 2) to separate the network part. Name marks consistently with a prefix so they are easy to filter, and mark the first and later calls of the operations that matter — the first decode(), the first search — since first-call cost includes baseline-tier code and lazy initialisation.

Step 2 — add Resource Timing for the .wasm file

const [entry] = performance.getEntriesByName(new URL("./engine.wasm", import.meta.url).href);
const net = entry && {
  transfer: entry.transferSize,                    // 0 when served from cache
  encoded: entry.encodedBodySize,
  decoded: entry.decodedBodySize,
  download: entry.responseEnd - entry.requestStart,
};

transferSize of zero means the file came from the HTTP cache; comparing cached and uncached loads shows what caching saves. Cross-origin .wasm files need a Timing-Allow-Origin header for these fields to be populated. Caching strategy is covered in setting cache-control headers for Wasm.

Step 3 — time hot operations without overhead

For operations called often, timing every call costs more than it is worth. Sample: time one call in a hundred, or aggregate locally and send summary statistics:

const stats = { n: 0, sum: 0, max: 0 };
export function timedDecode(bytes) {
  const sample = Math.random() < 0.01;
  const t0 = sample ? performance.now() : 0;
  const out = engine.decode(bytes);
  if (sample) { const d = performance.now() - t0; stats.n++; stats.sum += d; stats.max = Math.max(stats.max, d); }
  return out;
}

Record the input size with each sample — a decode of a 50 KB image and of a 20 MB image are different measurements — and bucket by size when aggregating. performance.now() is coarsened to between 5 µs and 100 µs depending on the browser and isolation, which is fine for operations in the millisecond range.

What to measure for a Wasm feature Download and compile are measured together with streaming compilation, with Resource Timing separating network time and cache hits. Instantiation, the first call and steady-state calls are measured separately, each pointing to different optimisations. metric how typical fix download Resource Timing smaller binary, Brotli, caching download + compile marks around compileStreaming streaming, code caching instantiate marks fewer start-up tasks first call mark the first operation lazy init, warm-up steady-state calls sampled timing profile the hot path

Step 4 — send, aggregate and segment

Send metrics with navigator.sendBeacon when the page is hidden, so they survive navigation:

addEventListener("visibilitychange", () => {
  if (document.visibilityState !== "hidden") return;
  navigator.sendBeacon("/rum", JSON.stringify({
    build: WASM_BUILD_ID,
    load: performance.getEntriesByName("wasm-load")[0]?.duration,
    instantiate: performance.getEntriesByName("wasm-instantiate")[0]?.duration,
    net, decode: stats,
    device: { cores: navigator.hardwareConcurrency, memory: navigator.deviceMemory },
  }));
});

Aggregate as percentiles — median, 75th, 95th — never averages, which hide the slow tail behind a comfortable-looking mean. Segment by device class (cores, memory), browser and engine, cache state and build. A regression often appears only in one segment, such as a new build that compiles slowly on older Safari.

Step 5 — connect Wasm timings to user-facing metrics

Wasm timings matter when they affect what users perceive. If the module loads before first render, its download and compile delay Largest Contentful Paint. If heavy calls run on the main thread, they appear as long tasks and hurt Interaction to Next Paint. Collect Web Vitals alongside the Wasm metrics, in the same beacon (the web-vitals library does this) and compare: sessions with slow Wasm loads should show correspondingly slower vitals if the module is on the critical path. If they do not, the module is already off the critical path, and optimising its load is less urgent than its execution.

Measuring work in workers

Many Wasm features run in workers, and their timings need to reach the page’s RUM pipeline. performance.mark and measure work inside workers, but those entries live in the worker’s own performance timeline, invisible to code on the page. Collect the worker’s measurements there and post them to the page with the results, or send beacons directly from the worker, since navigator.sendBeacon and fetch are available in worker scope. Use performance.timeOrigin to align worker and page timestamps when you need to place worker work on the page’s timeline, for example to show that a slow decode in the worker delayed a visible update. Measure queueing as well as execution: a job that waits 300 ms behind other jobs in a busy worker feels just as slow to the user as one that computes for 300 ms, and the remedy — more workers, smaller jobs, prioritisation — is different.

Turning measurements into decisions

Field data is only useful if it changes what you do. Set budgets per phase — for example, “95th percentile load under 1.5 s on mobile, first decode under 200 ms” — and review them on a dashboard per release. When a budget is breached, the phase breakdown points at the fix: a download regression suggests checking binary size and compression in CI, a compile regression suggests a toolchain or optimisation-level change, a first-call regression suggests new lazy initialisation. Run controlled experiments where the decision is not obvious: serve the SIMD build to half of capable users and compare decode percentiles, or test lazy loading against eager loading and compare both Wasm timings and Web Vitals. Keep a few months of history; seasonal changes in the device mix can look like regressions unless the segmentation shows otherwise. And feed the findings back to the lab: if field data shows mid-range Android phones dominate the slow tail, benchmark on one in CI, as described in tracking benchmark results in CI.

Expected output

The RUM dashboard shows, per build and device class, the 50th, 75th and 95th percentiles for Wasm load, instantiation, first call and steady-state calls, plus cache-hit rates for the .wasm file — and a release that doubled compile time on mid-range Android is visible within a day.

Gotchas

  • Averages instead of percentiles. The slow tail disappears. Report percentiles.
  • Timing every hot call. Measurement overhead distorts results. Sample.
  • Missing Timing-Allow-Origin. Cross-origin Resource Timing fields are zero. Add the header.
  • No input-size context. Timings vary with input. Record and bucket by size.
  • Ignoring worker timings. Worker marks are invisible to the page. Post them back or beacon from the worker.
  • Mixing builds in one chart. A regression hides among old releases. Segment every metric by build.
  • Losing data on navigation. Use sendBeacon on visibilitychange.

Performance note

Field data for one application showed a median Wasm load of 310 ms but a 95th percentile of 2.4 s, almost entirely on uncached mid-range Android devices. Enabling Brotli and a service-worker cache moved the 95th percentile to 900 ms; the median barely changed.

Wasm load time percentiles before and after caching fixes Milliseconds to download and compile the module at the 50th, 75th and 95th percentiles across real users, before and after enabling Brotli and a service-worker cache. ms to compiled module p50 before 310 ms p95 before 2,400 ms p50 after 280 ms p95 after 900 ms

Frequently Asked Questions

Can I measure compile time separately from download? Not exactly with streaming compilation, since they overlap. Compare the total with Resource Timing’s download duration, or measure non-streaming compile in a sample of sessions.

Does the browser report Wasm compile times itself? Not through a standard API. DevTools shows them locally; in the field, use your own marks.

How much data should I collect? A few percent of sessions is usually enough for stable percentiles; increase sampling for rare device segments.

What about server-side Wasm? Measure the same phases in the host — compile, instantiate, call — and export them as metrics; see tracing Wasm requests with OpenTelemetry.

Should I measure in development builds? No — development builds are larger and unoptimised. Measure production builds, in the field and in CI.

Do ad blockers affect RUM data? Some block common analytics endpoints. Serve the beacon endpoint from your own domain to reduce gaps.

← Back to Observability & Error Reporting