Lazy Loading Wasm on First Use
This page answers one task: a page ships a WebAssembly module that only some users need, or only after some interaction — defer downloading and compiling it until it is needed, without making the first use feel slow.
Prerequisites
- [ ] A module behind a specific feature: an editor tool, an export format, a codec, a search index.
- [ ] Glue that can be imported dynamically (an ES module), as in
emitting ES modules from Emscripten
or wasm-pack’s
webandbundlertargets. - [ ] A bundler that supports dynamic
import()and splits chunks.
The cost of eager loading
A module loaded at startup competes with everything else the page needs. It takes bandwidth from the HTML, styles, fonts and the JavaScript that renders the first screen. Its compilation takes CPU — on a phone, often hundreds of milliseconds — on the main thread or a background thread that the page’s own scripts could have used. And its memory is allocated whether or not the user ever touches the feature. For a module that drives the page’s core function, that cost is necessary. For a module behind a button that one user in five presses, it is waste paid by everyone.
Lazy loading moves that cost to the moment of need. The risk is that the user then waits — a click that takes two seconds because a module is downloading feels broken. The technique that works is therefore two-part: defer the module off the critical path, and then use idle time and intent signals to load it before the user actually asks.
Step 1 — move the import behind the feature
Replace the static import with a dynamic one inside the code path that uses the module, and cache the promise so it loads only once:
// export.js
let pdfModule;
function loadPdf() {
pdfModule ??= import("./pkg/pdf_export.js").then(async (m) => {
await m.default(); // wasm-bindgen init: fetches and instantiates the .wasm
return m;
}).catch((err) => { pdfModule = undefined; throw err; }); // allow retry after a network error
return pdfModule;
}
export async function exportPdf(doc) {
const m = await loadPdf();
return m.render_pdf(doc.toBytes());
}
The bundler turns the dynamic import into a separate chunk containing the glue, and the glue’s import.meta.url reference makes the .wasm
a separate asset. Neither is requested at startup. Resetting the cached promise on failure matters on flaky connections — without it, one
failed download breaks the feature until reload.
Step 2 — prefetch during idle time
Once the page is interactive and idle, download the module at low priority so it is in the HTTP cache when needed:
function prefetch(url) {
const link = document.createElement("link");
link.rel = "prefetch";
link.href = url;
document.head.append(link);
}
requestIdleCallback(() => {
prefetch(new URL("./pkg/pdf_export_bg.wasm", import.meta.url).href);
}, { timeout: 4000 });
prefetch downloads with the lowest priority and does not execute or compile anything. On metered connections, check
navigator.connection?.saveData and skip prefetching large modules when it is set. The difference between prefetch and preload — which
is for resources needed for the current navigation — is covered in
preloading Wasm with link rel=preload.
Step 3 — load on intent, not only on click
Users signal intent before they click: hovering a button, opening a menu that contains it, focusing a field. Start loading at that signal and the module often arrives before the click:
const button = document.querySelector("#export-pdf");
const warm = () => { loadPdf(); }; // starts download + compile, returns immediately
button.addEventListener("pointerenter", warm, { once: true });
button.addEventListener("focus", warm, { once: true });
document.querySelector("#export-menu").addEventListener("toggle", warm, { once: true });
For modules large enough that even a head start is not enough, compile in a worker ahead of time — the compiled WebAssembly.Module can be
posted to the main thread or to the worker that will run it, as described in
compiling Wasm in a worker to free the main thread.
Step 4 — show progress for large modules
When a module is tens of megabytes — a model, a full application runtime — loading takes seconds even when started early, and the user deserves to see it. Download with a stream reader to report progress, then compile from the bytes:
async function fetchWithProgress(url, onProgress) {
const res = await fetch(url);
const total = Number(res.headers.get("Content-Length")) || 0;
const reader = res.body.getReader();
const chunks = []; let received = 0;
for (;;) {
const { done, value } = await reader.read();
if (done) break;
chunks.push(value); received += value.length;
onProgress(total ? received / total : NaN);
}
return new Blob(chunks).arrayBuffer();
}
Reading the stream yourself gives up streaming compilation, which is a fair trade when the progress bar is what the user needs. For a
middle ground, tee the response — one branch to compileStreaming, one to a progress counter. Note that Content-Length reflects the
compressed size when the response is compressed, so measure progress against compressed bytes.
Step 5 — handle failure as a normal state
Lazily loaded code fails in ways startup code does not: the user went offline, a deploy removed the old asset, the device ran out of memory compiling a large module. Treat failure as a state the UI can show — “PDF export is unavailable, retry” — not as an uncaught promise rejection. Keep the previous release’s assets deployed for a while so long-lived tabs can still load the module they reference, as discussed in versioning Wasm files with content hashes.
Measuring the result
Lazy loading is only a win if it improves what users experience, so measure both ends. Startup metrics — time to interactive, total blocking time, the size of the initial download — should improve. First-use latency for the deferred feature — from click to result — should not get noticeably worse, which is what the prefetching and intent signals are for. Record both in real-user monitoring, as described in measuring Wasm performance with real-user monitoring, and segment by device class. A change that helps desktops and hurts phones — because prefetching competes with a slow connection — shows up only when the data is split that way.
Expected output
On page load, the Network panel shows no request for the deferred module. A few seconds later, an idle-time prefetch appears with Lowest priority. When the user clicks, the module is served from the prefetch cache and the feature responds without a visible delay.
Gotchas
- The module still loads at startup. A static import elsewhere pulls the chunk into the main bundle. Search for the package name in static imports.
- Prefetch and the real request both download. The prefetched URL differs from the one the glue requests — a different hash or query string. Prefetch exactly the URL the glue resolves.
- A loading spinner on every click. The feature awaits the module even after it has loaded, adding a microtask and a re-render. Check the cached promise’s state, or keep a resolved reference, to make subsequent calls synchronous where possible.
- The first click double-loads. Two code paths each started their own import. Share one cached promise.
- Memory spikes on phones. Compiling a very large module on demand can fail on low-memory devices. Catch the error and offer the feature as unavailable rather than crashing.
Performance note
Deferring a 1.8 MB export module on a document editor cut time to interactive on a mid-range phone from 2.2 s to 0.9 s. With idle-time prefetch and hover warm-up, the median first export started 40 ms later than with eager loading, against a 1.3 s improvement at startup for every visit.
Frequently Asked Questions
Does lazy loading break the compiled-code cache? No. A lazily loaded module from a stable, hashed URL is cached and its compiled code reused just like an eagerly loaded one.
Does import() work in every browser I support?
Dynamic import is supported in all current browsers. Bundlers fall back to their own chunk loaders where needed.
Should I lazy load in a worker? If the module runs in a worker, load it there — create the worker on first use, and let the worker fetch and compile its module.
What if the feature is needed offline? Precache the module with a service worker even though the page loads it lazily; see caching Wasm with a service worker.
Can search engines or pre-renderers see content produced by a lazily loaded module? Only if the content does not depend on it. Anything that must appear in the initial HTML — for indexing or for users without JavaScript — should come from the server, with the module enhancing it afterwards, as in progressive enhancement with Wasm.
How small does a module need to be before lazy loading is pointless? For a few kilobytes, the extra request can cost more than it saves; inline or load such modules eagerly.
Related
- Splitting a Wasm module for lazy loading — making a separable module in the first place.
- Reducing Wasm cold-start latency — making the eventual load fast.
- Designing a promise-based API around a Wasm module — the loader pattern used above.
- Integrating Wasm into a React app — lazy loading inside a framework.
← Back to Module Caching & Startup Performance