Measuring Wasm Startup Time End to End
This page answers one task: users wait for a WebAssembly feature to become usable, and you want to know exactly where that time goes — network, compilation, instantiation, start-up code or the first call — so you optimise the phase that matters instead of guessing.
Prerequisites
- [ ] Access to the code that loads the module (your loader or the generated glue).
- [ ] Chrome DevTools, and ideally a way to collect timings from real users.
- [ ] A production build served with production headers and compression.
The phases of Wasm startup
From the user’s point of view, startup is one wait. Underneath, it is a sequence: the request is issued (possibly delayed by other requests or by when the
page decides to load the module); bytes download; the engine compiles the module — with streaming compilation, mostly while bytes arrive; the module is
instantiated, which allocates memory, initialises data segments and links imports; initialisation code runs (_start, constructors, __wbindgen_start,
your own init function); and finally the first real call does its work, often slower than later calls because code is still in the baseline tier and
caches are cold.
Each phase has different fixes. Download time responds to compression and size; compile time to size and code caching; instantiation to memory size and data segments; initialisation to what start-up code does; first-call cost to warm-up strategy. Measuring the phases separately is what tells you which fix is worth doing.
Step 1 — mark each phase in the loader
Use performance.mark and performance.measure around each step. If you use instantiateStreaming, compile and instantiate are fused; split them with
compileStreaming plus instantiate for measurement:
performance.mark("wasm:request");
const response = fetch(new URL("./app_bg.wasm", import.meta.url));
const module = await WebAssembly.compileStreaming(response);
performance.mark("wasm:compiled");
const instance = await WebAssembly.instantiate(module, imports);
performance.mark("wasm:instantiated");
instance.exports.init();
performance.mark("wasm:initialised");
instance.exports.process(firstInput);
performance.mark("wasm:first-call");
performance.measure("wasm compile (incl. download)", "wasm:request", "wasm:compiled");
performance.measure("wasm instantiate", "wasm:compiled", "wasm:instantiated");
performance.measure("wasm init", "wasm:instantiated", "wasm:initialised");
performance.measure("wasm first call", "wasm:initialised", "wasm:first-call");
Marks appear in the DevTools Performance panel’s Timings track, aligned with network and compile activity.
Step 2 — separate download from compilation
Resource Timing gives the network part for the .wasm URL:
const [entry] = performance.getEntriesByName(new URL("./app_bg.wasm", import.meta.url).href);
console.table({
queued: entry.requestStart - entry.startTime,
ttfb: entry.responseStart - entry.requestStart,
download: entry.responseEnd - entry.responseStart,
transferKB: entry.transferSize / 1024,
decodedKB: entry.decodedBodySize / 1024,
});
With streaming compilation, compile time beyond responseEnd is the tail that did not overlap the download: compiled − responseEnd (convert marks to
the same time origin). A large tail means compilation is slower than the network — common on fast connections and slow devices — and points to size
reduction or caching. A small tail means the network dominates.
Step 3 — record a trace for detail
In the Performance panel, record a cold load (cache disabled). The trace shows the network request, compile tasks on background threads
(v8.wasm.compileStreaming, wasm.CompileLazy or similar names), instantiation on the main thread, and your marks. Look for gaps: time between the page
starting and the request being issued (the loader started late), main-thread work during instantiation (large data segments or start-up code), and
long first calls (lazy compilation of functions on first use). Repeat the recording with CPU throttling to approximate phones.
Step 4 — measure cold and warm separately
Startup has at least three states: cold (no cache), warm HTTP cache (bytes cached, compilation needed), and warm code cache (Chrome and others reuse
compiled code for repeat loads of the same response). Measure each: the gap between cold and warm shows what caching delivers, and if warm loads are not
much faster, check that the URL is stable, that Cache-Control allows caching, and that the module is loaded with streaming APIs, which are required for
code caching in some engines. See
setting cache-control headers for Wasm.
Step 5 — collect phase timings from real users
Lab measurements describe one device and one network; users span many. Send the phase measures to your analytics or real-user monitoring endpoint:
const phases = performance.getEntriesByType("measure").filter((m) => m.name.startsWith("wasm"));
navigator.sendBeacon("/rum", JSON.stringify(Object.fromEntries(phases.map((m) => [m.name, Math.round(m.duration)]))));
Report the 50th, 75th and 95th percentiles per phase. Real-user data often shows that the 95th percentile is dominated by one phase — download on slow networks, or initialisation on low-end devices — which tells you where the worst-off users wait.
Initialisation code is often the surprise
Teams usually expect download or compile to dominate, and are surprised by initialisation. Rust programs with large lazy_static or once_cell
initialisers, C++ programs with static constructors, runtimes that parse embedded data at start (time-zone databases, Unicode tables, fonts), and
interpreters that boot a standard library can spend hundreds of milliseconds after instantiation before the first useful call. The fix is to defer work
until it is needed, precompute data at build time into a form that needs no parsing, or snapshot initialised memory with Wizer, which runs initialisation
at build time and stores the resulting memory in the module.
The first call and tier-up
The first call into a function runs baseline-tier code in most engines, and some engines compile functions lazily on first call, so the first invocation pays compilation too. For features where the first interaction must feel instant, call the hot functions once with a small input during idle time after instantiation — a warm-up — so compilation and caches are ready when the user acts. Measure that the warm-up does not itself delay something more important.
Measuring in workers and frameworks
Many applications load WebAssembly inside a worker or through a framework’s glue, which hides the phases. In a worker, place the same marks inside the
worker script — performance exists there with its own time origin — and post the measures back to the main thread with performance.timeOrigin, so
both timelines can be aligned. Add marks for worker creation and for the first message round trip, since spinning up a worker and loading its script can
cost tens of milliseconds before the Wasm request is even issued. For generated glue that calls instantiateStreaming internally (wasm-bindgen’s
init, Emscripten’s module factory), mark before and after the glue’s promise, and use Resource Timing and the DevTools trace to separate the inner
phases. Emscripten’s instantiateWasm hook and wasm-bindgen’s ability to accept a precompiled WebAssembly.Module both let you take over compilation
so it can be measured, or moved, explicitly.
Turning measurements into a budget
Once the phases are known, set a budget for the total and for the phase most likely to regress — for example, “first useful call within 800 ms at the 75th percentile on mid-range Android, with initialisation under 150 ms”. Check the lab part in CI with a throttled profile and the field part in monitoring dashboards. A budget per phase makes regressions explainable: when the total grows, the dashboard already shows whether the module got bigger, the network got slower, or start-up code got heavier.
Expected output
A per-phase breakdown for the editor’s module: on a mid-range phone over Wi-Fi, 60 ms queueing, 250 ms download, 140 ms compile tail, 15 ms instantiate, 210 ms initialisation and 70 ms first call; real-user percentiles showing initialisation dominating at p95 on low-end devices; and a decision to snapshot initialisation with Wizer first.
Gotchas
- Measuring
instantiateStreamingas one number. Compile and instantiate are fused. Split them for measurement. - Ignoring queueing. A late request looks like slow download. Check when the request starts.
- Testing warm loads only. Code caching hides compile time. Test cold too.
- Development builds. Uncompressed and unoptimised timings mislead. Measure production builds.
- Forgetting initialisation. Start-up code is often the largest phase.
- Not aligning worker timelines. Worker marks use their own time origin. Convert with
performance.timeOrigin.
Performance note
For the editor module, snapshotting initialisation with Wizer removed 190 ms on the mid-range phone, more than halving the gap between instantiation and first call; compression changes would have saved 60 ms at most on Wi-Fi.
Frequently Asked Questions
Does compileStreaming plus instantiate cost more than instantiateStreaming?
No measurable difference; both stream compilation.
How do I measure compile time without streaming?
Time WebAssembly.compile(bytes) after the bytes have arrived.
Do workers change the phases? Loading in a worker adds worker startup and message passing; mark those too.
Is the first call always slow? Usually somewhat slower; warm-up calls during idle time hide it.
How do I measure startup inside a worker?
Mark phases inside the worker and post them back with the worker’s performance.timeOrigin so timelines can be aligned.
Which phase should I optimise first? The largest one at the percentile you care about — often initialisation on low-end devices and download on slow networks.
Related
- Reducing Wasm cold-start latency — fixes by phase.
- Simulating slow networks for Wasm loading — realistic conditions.
- Benchmarking Wasm on mobile devices — phone measurements.
- Measuring Wasm performance with real-user monitoring — field data.
← Back to Wasm Performance Benchmarking