Why the First Call into Wasm Is Slow
This page answers one question: a WebAssembly function takes 40 ms the first time it runs and 4 ms every time after — where does the difference come from, and what can you do when that first call is the one users wait for?
Prerequisites
- [ ] A module and a function whose first call you can time separately from later calls.
- [ ] Chrome DevTools or Node for tracing, and some familiarity with how V8 compiles Wasm with Liftoff and TurboFan.
Several costs stacked on one call
A slow first call is rarely one problem. It is usually several one-time costs that all happen to land on the first invocation, and each has a
different fix. Engines that compile lazily generate machine code for a function only when it is first called, so the first call pays for
compilation. Before tier-up, the function runs as baseline code, slower than it will be later. The CPU’s caches and branch predictors are cold for
code and data that have not been touched. Memory pages the function writes for the first time must be mapped by the operating system — first-touch
page faults. And the glue around the module often initialises itself lazily: allocating its heap, building lookup tables, creating the
TextDecoder it will reuse.
Later calls pay none of these. That is why a single measurement of a first call says almost nothing about the function’s real cost — and why, for user-facing work that only happens once, the first call is the one that matters.
Step 1 — measure the first call separately
const { instance } = await WebAssembly.instantiateStreaming(fetch("filter.wasm"), imports);
const run = () => instance.exports.apply_filter(ptr, width, height);
const t0 = performance.now(); run(); const first = performance.now() - t0;
const later = [];
for (let i = 0; i < 50; i++) { const t = performance.now(); run(); later.push(performance.now() - t); }
later.sort((a, b) => a - b);
console.log({ first: first.toFixed(1), median: later[25].toFixed(1) });
{ first: '38.6', median: '4.1' }
A ratio of ten is common for compute-heavy functions in large modules. The rest of this page takes the difference apart.
Step 2 — attribute the time with a trace
Record a Performance panel trace around the first call. The pieces are usually visible directly: a CompileLazy or WebAssembly compile slice at
the start of the call, page-fault and memory-allocation time showing as system activity, and the glue’s initialisation as JavaScript before the
first Wasm frame. Comparing the first call’s flame chart with a later one shows which pieces disappear.
Step 3 — warm up before the user needs it
The most effective remedy for user-facing first calls is to make the first call happen before the user’s request — during idle time, with a small representative input:
requestIdleCallback(() => {
const tiny = makeTestImage(64, 64); // same code path, trivial size
exports.apply_filter(tiny.ptr, 64, 64); // triggers lazy compile, glue init, some tier-up
});
A warm-up call pays lazy compilation and glue initialisation, and several of them push hot functions towards tier-up. It does not pay first-touch costs for the real input’s memory, since the tiny input touches little. Make sure the warm-up exercises the same code paths as real work — different branches may lead to functions that are still cold.
Step 4 — pre-touch and pre-size memory
First-touch page faults come from writing memory for the first time, typically right after a memory.grow that the first large operation
triggered. Two changes remove most of it. Size initial memory so the common operation does not need to grow it, as described in
sizing initial and maximum memory.
And allocate working buffers once, at startup or during warm-up, then reuse them across calls, so the pages are already mapped when the user’s
first real call arrives:
// reserve and touch the working buffer once, during idle time
const bufPtr = exports.alloc(width * height * 4);
new Uint8Array(exports.memory.buffer, bufPtr, width * height * 4).fill(0);
Step 5 — let caching remove compilation entirely
For returning visitors, the browser’s compiled-code cache removes both the compile cost and much of the baseline-tier cost, since cached code for hot functions is usually optimized. That only works when the module is loaded with streaming APIs from a stable, cacheable URL. Content-hashed module names with long cache lifetimes give returning users fast first calls without any warm-up code; see caching Wasm with a service worker for the offline-capable version.
The glue’s share of the first call
The JavaScript side of the boundary deserves its own look, because it is often a surprising share of a slow first call and is the easiest part to
fix. Generated glue tends to initialise lazily: wasm-bindgen creates its TextEncoder and TextDecoder on first use, builds typed-array views over
memory on first access, and sets up its object heap when the first JavaScript value crosses the boundary. Emscripten’s runtime may initialise its
file system, environment or locale on first use of the corresponding functions. Each of those is small, but on a phone several milliseconds of them
can land on the first call.
Your own wrapper code adds to it. A first call that also imports a module dynamically, fetches a configuration file, or allocates a large result buffer stacks those costs on top. Profile the first call with the JavaScript frames visible — the Bottom-Up view in the Performance panel shows them next to the Wasm frames — and move any setup that does not depend on the user’s input into the idle-time warm-up from step 3. What remains on the first call should be the work itself, plus whatever tier-up has not yet done.
Deciding which first call matters
Not every first call deserves effort. A function that runs once during startup, before the user has done anything, can absorb its first-call cost inside the startup budget — measure startup as a whole and optimize that. A function behind an interaction — the export button, the first filter applied — is where a slow first call is felt, because the user is waiting for that specific result. And a function called continuously, such as a per-frame update, has a first call that is lost in the noise of the first frames. Focus warm-up and pre-sizing on the second category: the interactions users perform early and expect to be instant. Real-user monitoring that records first-call and later-call timings separately, as described in measuring Wasm performance with real-user monitoring, tells you which interactions those are.
Expected output
After adding an idle-time warm-up and pre-sizing memory, the user-facing first call drops close to steady state:
before: { first: '38.6', median: '4.1' }
after: { first: '6.3', median: '4.1' }
The remaining gap is mostly functions that had not yet tiered up when the user arrived.
Gotchas
- Warm-up exercising the wrong path. A tiny input may skip branches the real input takes. Use an input that follows the same path.
- Warm-up on the main thread at startup. It competes with the page loading. Run it at idle time or in the worker that will do the work.
- Measuring only first calls in benchmarks. They include all the one-time costs. Report first and steady-state separately.
- Warming up in the wrong context. A warm-up on the main thread does not help a worker that has its own instance. Warm up where the work runs.
- Growing memory per operation. Each grow brings new first-touch costs. Reuse buffers.
Performance note
Across three applications, idle-time warm-up plus pre-sized memory cut user-visible first-call latency by 70–85%. The remaining first-call cost was dominated by functions still on baseline code; returning visits with a warm compiled-code cache removed most of that too.
Frequently Asked Questions
Is lazy compilation bad? No — it shortens startup by skipping functions that never run. Its cost is concentrated on first calls, which warm-up can move to idle time.
Does this happen in server runtimes? wasmtime compiles everything ahead of time with Cranelift, so there is no lazy compile or tier-up per call, but first-touch memory and cold caches still apply. Precompiled modules remove compilation from startup entirely.
Should I call every export once at startup? Only those behind interactions users perform early. Warming everything wastes CPU and battery on code that may never run.
Why is the second call sometimes still slow? Tier-up is not immediate; a function may need dozens of calls before optimized code arrives. Watch per-call times as in watching Wasm tier-up in Chrome.
Related
- Reducing Wasm cold-start latency — the module-level version of this problem.
- Lazy loading Wasm on first use — when the module itself loads late.
- Avoiding JIT warm-up errors in Wasm benchmarks — the measurement side.
- How Safari runs Wasm — where first calls start in an interpreter.
← Back to Engine Tiering & JIT Compilation