Why the First Call into Wasm Is Slow

This page answers one question: a WebAssembly function takes 40 ms the first time it runs and 4 ms every time after — where does the difference come from, and what can you do when that first call is the one users wait for?

Prerequisites

Several costs stacked on one call

A slow first call is rarely one problem. It is usually several one-time costs that all happen to land on the first invocation, and each has a different fix. Engines that compile lazily generate machine code for a function only when it is first called, so the first call pays for compilation. Before tier-up, the function runs as baseline code, slower than it will be later. The CPU’s caches and branch predictors are cold for code and data that have not been touched. Memory pages the function writes for the first time must be mapped by the operating system — first-touch page faults. And the glue around the module often initialises itself lazily: allocating its heap, building lookup tables, creating the TextDecoder it will reuse.

Later calls pay none of these. That is why a single measurement of a first call says almost nothing about the function’s real cost — and why, for user-facing work that only happens once, the first call is the one that matters.

What a first call pays that later calls do not The first call to a Wasm function stacks several one-time costs on top of the function's real work: lazy compilation, running baseline-tier code, first-touch page faults on fresh memory, cold CPU caches and lazy glue initialisation. lazy compile of the function machine code generated on first call baseline-tier execution 2-4× slower until tier-up first-touch page faults OS maps fresh memory pages cold caches and predictors instructions and data not yet cached lazy glue initialisation heap, decoders, tables created on demand the function's real work what every later call costs

Step 1 — measure the first call separately

const { instance } = await WebAssembly.instantiateStreaming(fetch("filter.wasm"), imports);
const run = () => instance.exports.apply_filter(ptr, width, height);

const t0 = performance.now(); run(); const first = performance.now() - t0;
const later = [];
for (let i = 0; i < 50; i++) { const t = performance.now(); run(); later.push(performance.now() - t); }
later.sort((a, b) => a - b);
console.log({ first: first.toFixed(1), median: later[25].toFixed(1) });
{ first: '38.6', median: '4.1' }

A ratio of ten is common for compute-heavy functions in large modules. The rest of this page takes the difference apart.

Step 2 — attribute the time with a trace

Record a Performance panel trace around the first call. The pieces are usually visible directly: a CompileLazy or WebAssembly compile slice at the start of the call, page-fault and memory-allocation time showing as system activity, and the glue’s initialisation as JavaScript before the first Wasm frame. Comparing the first call’s flame chart with a later one shows which pieces disappear.

Breakdown of one slow first call The 38.6 ms first call of an image filter in Chrome on a laptop, broken into lazy compilation, extra time from baseline-tier code, first-touch page faults on a freshly grown memory, glue initialisation, and the function's steady-state work. ms of the first call lazy compile 7.2 ms baseline code (extra vs optimized) 13.5 ms first-touch page faults 9.1 ms glue initialisation 4.7 ms steady-state work 4.1 ms

Step 3 — warm up before the user needs it

The most effective remedy for user-facing first calls is to make the first call happen before the user’s request — during idle time, with a small representative input:

requestIdleCallback(() => {
  const tiny = makeTestImage(64, 64);                  // same code path, trivial size
  exports.apply_filter(tiny.ptr, 64, 64);              // triggers lazy compile, glue init, some tier-up
});

A warm-up call pays lazy compilation and glue initialisation, and several of them push hot functions towards tier-up. It does not pay first-touch costs for the real input’s memory, since the tiny input touches little. Make sure the warm-up exercises the same code paths as real work — different branches may lead to functions that are still cold.

Moving one-time costs out of the user's first call Without warm-up, compilation, glue initialisation and first-touch costs all land on the user's first click. With an idle-time warm-up a few seconds after load, those costs are paid while the user is reading, and the click runs close to steady-state speed. 0.5 s page interactive 2 s idle warm-up call 3 s buffers touched 7.5 s click: 6 ms, not 39

Step 4 — pre-touch and pre-size memory

First-touch page faults come from writing memory for the first time, typically right after a memory.grow that the first large operation triggered. Two changes remove most of it. Size initial memory so the common operation does not need to grow it, as described in sizing initial and maximum memory. And allocate working buffers once, at startup or during warm-up, then reuse them across calls, so the pages are already mapped when the user’s first real call arrives:

// reserve and touch the working buffer once, during idle time
const bufPtr = exports.alloc(width * height * 4);
new Uint8Array(exports.memory.buffer, bufPtr, width * height * 4).fill(0);

Step 5 — let caching remove compilation entirely

For returning visitors, the browser’s compiled-code cache removes both the compile cost and much of the baseline-tier cost, since cached code for hot functions is usually optimized. That only works when the module is loaded with streaming APIs from a stable, cacheable URL. Content-hashed module names with long cache lifetimes give returning users fast first calls without any warm-up code; see caching Wasm with a service worker for the offline-capable version.

The glue’s share of the first call

The JavaScript side of the boundary deserves its own look, because it is often a surprising share of a slow first call and is the easiest part to fix. Generated glue tends to initialise lazily: wasm-bindgen creates its TextEncoder and TextDecoder on first use, builds typed-array views over memory on first access, and sets up its object heap when the first JavaScript value crosses the boundary. Emscripten’s runtime may initialise its file system, environment or locale on first use of the corresponding functions. Each of those is small, but on a phone several milliseconds of them can land on the first call.

Your own wrapper code adds to it. A first call that also imports a module dynamically, fetches a configuration file, or allocates a large result buffer stacks those costs on top. Profile the first call with the JavaScript frames visible — the Bottom-Up view in the Performance panel shows them next to the Wasm frames — and move any setup that does not depend on the user’s input into the idle-time warm-up from step 3. What remains on the first call should be the work itself, plus whatever tier-up has not yet done.

Deciding which first call matters

Not every first call deserves effort. A function that runs once during startup, before the user has done anything, can absorb its first-call cost inside the startup budget — measure startup as a whole and optimize that. A function behind an interaction — the export button, the first filter applied — is where a slow first call is felt, because the user is waiting for that specific result. And a function called continuously, such as a per-frame update, has a first call that is lost in the noise of the first frames. Focus warm-up and pre-sizing on the second category: the interactions users perform early and expect to be instant. Real-user monitoring that records first-call and later-call timings separately, as described in measuring Wasm performance with real-user monitoring, tells you which interactions those are.

Expected output

After adding an idle-time warm-up and pre-sizing memory, the user-facing first call drops close to steady state:

before: { first: '38.6', median: '4.1' }
after:  { first: '6.3',  median: '4.1' }

The remaining gap is mostly functions that had not yet tiered up when the user arrived.

Gotchas

  • Warm-up exercising the wrong path. A tiny input may skip branches the real input takes. Use an input that follows the same path.
  • Warm-up on the main thread at startup. It competes with the page loading. Run it at idle time or in the worker that will do the work.
  • Measuring only first calls in benchmarks. They include all the one-time costs. Report first and steady-state separately.
  • Warming up in the wrong context. A warm-up on the main thread does not help a worker that has its own instance. Warm up where the work runs.
  • Growing memory per operation. Each grow brings new first-touch costs. Reuse buffers.

Performance note

Across three applications, idle-time warm-up plus pre-sized memory cut user-visible first-call latency by 70–85%. The remaining first-call cost was dominated by functions still on baseline code; returning visits with a warm compiled-code cache removed most of that too.

Frequently Asked Questions

Is lazy compilation bad? No — it shortens startup by skipping functions that never run. Its cost is concentrated on first calls, which warm-up can move to idle time.

Does this happen in server runtimes? wasmtime compiles everything ahead of time with Cranelift, so there is no lazy compile or tier-up per call, but first-touch memory and cold caches still apply. Precompiled modules remove compilation from startup entirely.

Should I call every export once at startup? Only those behind interactions users perform early. Warming everything wastes CPU and battery on code that may never run.

Why is the second call sometimes still slow? Tier-up is not immediate; a function may need dozens of calls before optimized code arrives. Watch per-call times as in watching Wasm tier-up in Chrome.

← Back to Engine Tiering & JIT Compilation