Measuring Wasm Startup Time End to End

This page answers one task: users wait for a WebAssembly feature to become usable, and you want to know exactly where that time goes — network, compilation, instantiation, start-up code or the first call — so you optimise the phase that matters instead of guessing.

Prerequisites

  • [ ] Access to the code that loads the module (your loader or the generated glue).
  • [ ] Chrome DevTools, and ideally a way to collect timings from real users.
  • [ ] A production build served with production headers and compression.

The phases of Wasm startup

From the user’s point of view, startup is one wait. Underneath, it is a sequence: the request is issued (possibly delayed by other requests or by when the page decides to load the module); bytes download; the engine compiles the module — with streaming compilation, mostly while bytes arrive; the module is instantiated, which allocates memory, initialises data segments and links imports; initialisation code runs (_start, constructors, __wbindgen_start, your own init function); and finally the first real call does its work, often slower than later calls because code is still in the baseline tier and caches are cold.

Each phase has different fixes. Download time responds to compression and size; compile time to size and code caching; instantiation to memory size and data segments; initialisation to what start-up code does; first-call cost to warm-up strategy. Measuring the phases separately is what tells you which fix is worth doing.

Phases of a cold Wasm startup On a cold load over a fast connection, the request starts at zero, the first bytes arrive at 90 milliseconds, the last byte at 340, streaming compilation finishes at 380, instantiation at 395, initialisation code at 470, and the first useful call completes at 540 milliseconds. 0 ms request issued 90 ms first byte 340 ms last byte 380 ms compiled 395 ms instantiated 470 ms init done 540 ms first call done

Step 1 — mark each phase in the loader

Use performance.mark and performance.measure around each step. If you use instantiateStreaming, compile and instantiate are fused; split them with compileStreaming plus instantiate for measurement:

performance.mark("wasm:request");
const response = fetch(new URL("./app_bg.wasm", import.meta.url));
const module = await WebAssembly.compileStreaming(response);
performance.mark("wasm:compiled");
const instance = await WebAssembly.instantiate(module, imports);
performance.mark("wasm:instantiated");
instance.exports.init();
performance.mark("wasm:initialised");
instance.exports.process(firstInput);
performance.mark("wasm:first-call");

performance.measure("wasm compile (incl. download)", "wasm:request", "wasm:compiled");
performance.measure("wasm instantiate", "wasm:compiled", "wasm:instantiated");
performance.measure("wasm init", "wasm:instantiated", "wasm:initialised");
performance.measure("wasm first call", "wasm:initialised", "wasm:first-call");

Marks appear in the DevTools Performance panel’s Timings track, aligned with network and compile activity.

Step 2 — separate download from compilation

Resource Timing gives the network part for the .wasm URL:

const [entry] = performance.getEntriesByName(new URL("./app_bg.wasm", import.meta.url).href);
console.table({
  queued: entry.requestStart - entry.startTime,
  ttfb: entry.responseStart - entry.requestStart,
  download: entry.responseEnd - entry.responseStart,
  transferKB: entry.transferSize / 1024,
  decodedKB: entry.decodedBodySize / 1024,
});

With streaming compilation, compile time beyond responseEnd is the tail that did not overlap the download: compiled − responseEnd (convert marks to the same time origin). A large tail means compilation is slower than the network — common on fast connections and slow devices — and points to size reduction or caching. A small tail means the network dominates.

Where startup time went on two configurations On a fast connection with a mid-range phone, compilation tail and initialisation dominate. On a slow connection with a laptop, download dominates and compilation overlaps it almost entirely. The same module needs different fixes for the two audiences. request queueing waiting for other requests download network + decompression compile tail compile beyond last byte instantiate + init memory, data, start code first call baseline-tier code, cold caches

Step 3 — record a trace for detail

In the Performance panel, record a cold load (cache disabled). The trace shows the network request, compile tasks on background threads (v8.wasm.compileStreaming, wasm.CompileLazy or similar names), instantiation on the main thread, and your marks. Look for gaps: time between the page starting and the request being issued (the loader started late), main-thread work during instantiation (large data segments or start-up code), and long first calls (lazy compilation of functions on first use). Repeat the recording with CPU throttling to approximate phones.

Step 4 — measure cold and warm separately

Startup has at least three states: cold (no cache), warm HTTP cache (bytes cached, compilation needed), and warm code cache (Chrome and others reuse compiled code for repeat loads of the same response). Measure each: the gap between cold and warm shows what caching delivers, and if warm loads are not much faster, check that the URL is stable, that Cache-Control allows caching, and that the module is loaded with streaming APIs, which are required for code caching in some engines. See setting cache-control headers for Wasm.

Step 5 — collect phase timings from real users

Lab measurements describe one device and one network; users span many. Send the phase measures to your analytics or real-user monitoring endpoint:

const phases = performance.getEntriesByType("measure").filter((m) => m.name.startsWith("wasm"));
navigator.sendBeacon("/rum", JSON.stringify(Object.fromEntries(phases.map((m) => [m.name, Math.round(m.duration)]))));

Report the 50th, 75th and 95th percentiles per phase. Real-user data often shows that the 95th percentile is dominated by one phase — download on slow networks, or initialisation on low-end devices — which tells you where the worst-off users wait.

Initialisation code is often the surprise

Teams usually expect download or compile to dominate, and are surprised by initialisation. Rust programs with large lazy_static or once_cell initialisers, C++ programs with static constructors, runtimes that parse embedded data at start (time-zone databases, Unicode tables, fonts), and interpreters that boot a standard library can spend hundreds of milliseconds after instantiation before the first useful call. The fix is to defer work until it is needed, precompute data at build time into a form that needs no parsing, or snapshot initialised memory with Wizer, which runs initialisation at build time and stores the resulting memory in the module.

The first call and tier-up

The first call into a function runs baseline-tier code in most engines, and some engines compile functions lazily on first call, so the first invocation pays compilation too. For features where the first interaction must feel instant, call the hot functions once with a small input during idle time after instantiation — a warm-up — so compilation and caches are ready when the user acts. Measure that the warm-up does not itself delay something more important.

Measuring in workers and frameworks

Many applications load WebAssembly inside a worker or through a framework’s glue, which hides the phases. In a worker, place the same marks inside the worker script — performance exists there with its own time origin — and post the measures back to the main thread with performance.timeOrigin, so both timelines can be aligned. Add marks for worker creation and for the first message round trip, since spinning up a worker and loading its script can cost tens of milliseconds before the Wasm request is even issued. For generated glue that calls instantiateStreaming internally (wasm-bindgen’s init, Emscripten’s module factory), mark before and after the glue’s promise, and use Resource Timing and the DevTools trace to separate the inner phases. Emscripten’s instantiateWasm hook and wasm-bindgen’s ability to accept a precompiled WebAssembly.Module both let you take over compilation so it can be measured, or moved, explicitly.

Turning measurements into a budget

Once the phases are known, set a budget for the total and for the phase most likely to regress — for example, “first useful call within 800 ms at the 75th percentile on mid-range Android, with initialisation under 150 ms”. Check the lab part in CI with a throttled profile and the field part in monitoring dashboards. A budget per phase makes regressions explainable: when the total grows, the dashboard already shows whether the module got bigger, the network got slower, or start-up code got heavier.

Expected output

A per-phase breakdown for the editor’s module: on a mid-range phone over Wi-Fi, 60 ms queueing, 250 ms download, 140 ms compile tail, 15 ms instantiate, 210 ms initialisation and 70 ms first call; real-user percentiles showing initialisation dominating at p95 on low-end devices; and a decision to snapshot initialisation with Wizer first.

Gotchas

  • Measuring instantiateStreaming as one number. Compile and instantiate are fused. Split them for measurement.
  • Ignoring queueing. A late request looks like slow download. Check when the request starts.
  • Testing warm loads only. Code caching hides compile time. Test cold too.
  • Development builds. Uncompressed and unoptimised timings mislead. Measure production builds.
  • Forgetting initialisation. Start-up code is often the largest phase.
  • Not aligning worker timelines. Worker marks use their own time origin. Convert with performance.timeOrigin.

Performance note

For the editor module, snapshotting initialisation with Wizer removed 190 ms on the mid-range phone, more than halving the gap between instantiation and first call; compression changes would have saved 60 ms at most on Wi-Fi.

Time from request to first useful call on a mid-range phone Milliseconds from request to the first completed call before and after snapshotting initialisation with Wizer, on Wi-Fi with a cold cache. ms to first useful call before 745 ms after Wizer snapshot 555 ms

Frequently Asked Questions

Does compileStreaming plus instantiate cost more than instantiateStreaming? No measurable difference; both stream compilation.

How do I measure compile time without streaming? Time WebAssembly.compile(bytes) after the bytes have arrived.

Do workers change the phases? Loading in a worker adds worker startup and message passing; mark those too.

Is the first call always slow? Usually somewhat slower; warm-up calls during idle time hide it.

How do I measure startup inside a worker? Mark phases inside the worker and post them back with the worker’s performance.timeOrigin so timelines can be aligned.

Which phase should I optimise first? The largest one at the percentile you care about — often initialisation on low-end devices and download on slow networks.

← Back to Wasm Performance Benchmarking