Splitting a Wasm Module for Lazy Loading
This page answers one task: a single large WebAssembly module delays startup because everything — rarely used export formats, an admin tool, a full-text indexer — downloads and compiles before the first feature works. Split it so the critical code arrives first and the rest loads when it is needed.
Prerequisites
- [ ] A module large enough that startup cost is measurable: a few hundred kilobytes or more after compression.
- [ ] A clear view of which exports are used at startup and which are used later — a usage log or a profile helps.
- [ ] Binaryen 116+ if you want to try
wasm-split; plain cargo and wasm-bindgen for the crate-split approach.
Why splitting a Wasm module is harder than splitting JavaScript
JavaScript bundlers split code routinely: a dynamic import() becomes a separate chunk, loaded on demand, sharing the
same global scope. WebAssembly has no equivalent at the language level. A module is a closed unit — its functions,
tables, globals and memory are defined together and compiled together — and two modules can only share state that they
explicitly import from each other or from JavaScript.
There are two practical approaches. The first splits at the source level: build the rarely used functionality as a
separate crate and a separate module, with its own exports and a small interface to the main module. It is simple and
robust, and it costs some duplication, because each module carries its own copy of the allocator and any shared code. The
second splits at the binary level with Binaryen’s wasm-split, which takes one module, a profile of which functions ran
at startup, and produces a primary module plus a secondary one that is loaded the first time a moved function is called.
It avoids duplication but needs a profiling step and runtime support for the deferred load.
Step 1 — find what can wait
Before splitting anything, measure what startup actually uses. A simple approach is to instrument the JavaScript wrapper and record which exports are called in the first seconds of a session:
const used = new Set();
const exportsProxy = new Proxy(instance.exports, {
get(target, name) { used.add(name); return target[name]; },
});
setTimeout(() => console.table([...used]), 5000);
Combine that with a size breakdown from twiggy,
looking at the retained size of each export: twiggy dominators shows how much code is reachable only from a given
export. An export that is not used at startup and dominates a large share of the module is the candidate to move.
Step 2 — split by crate
Move the deferred feature into its own crate with its own #[wasm_bindgen] exports. Keep the interface between the two
small and data-oriented: the main module passes bytes or plain values, the secondary module returns bytes or plain values.
crates/
editor-core/ # startup: rendering, editing, file open
editor-export/ # deferred: PDF and DOCX export, fonts, layout engine
// main.js — startup path
import initCore, { open_document, render } from "./pkg/editor_core.js";
await initCore();
// export.js — loaded on first use
let exportModule;
export async function exportPdf(docBytes) {
exportModule ??= import("./pkg/editor_export.js").then(async (m) => { await m.default(); return m; });
const m = await exportModule;
return m.to_pdf(docBytes); // copies in, returns a Uint8Array
}
The dynamic import() makes the bundler emit the export package as a separate chunk, and the .wasm inside it is fetched
only when exportPdf is first called. The cost is a copy of the document across the boundary and a second allocator in the
second module — usually a few kilobytes, far less than the code moved out.
Step 3 — or split the binary with wasm-split
When the cold code is tangled with the hot code and does not separate cleanly into a crate, Binaryen can do the split. First build an instrumented module that records which functions run:
wasm-split --instrument app.wasm -o app.instrumented.wasm
Run the instrumented module through a representative startup — load the app, open a typical document — and save the profile the instrumentation writes (exposed through an exported function you call from JavaScript). Then split using that profile:
wasm-split app.wasm --profile=startup.prof \
-o1 app.primary.wasm -o2 app.secondary.wasm \
--export-prefix=% --placeholdermap=placeholders.txt
Functions that ran at startup stay in the primary module. Everything else moves to the secondary module, and the primary
gets placeholder functions in their table slots. When a placeholder is called, it calls a JavaScript import that loads and
instantiates the secondary module — sharing the primary’s memory and table — patches the table with the real functions,
and retries the call. Emscripten automates this loading with -sSPLIT_MODULE; for other toolchains you write that loader
yourself, which is the main reason the crate-level split is the more common choice.
Step 4 — prefetch before the first use
Lazy loading moves the cost from startup to first use; the user still waits, just later. Hide that wait by prefetching once the page is idle, so the module is usually in cache by the time the user clicks:
requestIdleCallback(() => {
const link = document.createElement("link");
link.rel = "prefetch";
link.href = new URL("./pkg/editor_export_bg.wasm", import.meta.url).href;
document.head.append(link);
}, { timeout: 5000 });
Prefetching downloads the bytes at low priority without compiling them. For an even faster first call, compile in a worker
ahead of time and keep the WebAssembly.Module ready, as in
compiling Wasm in a worker to free the main thread.
The broader pattern is covered in lazy loading Wasm on first use.
Step 5 — check that the split held
Splits erode: someone calls an export function from the startup path, and the deferred module is now loaded on every page view, plus the overhead of having two modules. Guard the split with a test that loads the app’s startup path and asserts the secondary module was not requested:
// e2e test with Playwright
const requested = [];
page.on("request", (r) => requested.push(r.url()));
await page.goto("/");
await page.waitForSelector("#editor[data-ready]");
expect(requested.some((u) => u.includes("editor_export_bg"))).toBe(false);
Expected output
For an editor with a large export feature, before and after a crate-level split:
before: app_bg.wasm 1,840 KB (612 KB brotli) loaded at startup
after: editor_core_bg.wasm 790 KB (268 KB brotli) loaded at startup
editor_export_bg.wasm 1,092 KB (361 KB brotli) loaded on first export
The two parts together are slightly larger than the original — the duplicated runtime — but startup downloads and compiles less than half.
Gotchas
- The split saves nothing at startup. Something in the startup path imports the deferred chunk statically. Check the bundler’s chunk graph for a static import of the secondary package.
- Two copies of a large shared dependency. Both crates depend on the same heavy crate. Move that code to one side, or accept a binary split instead.
- Version skew between the two parts. A cached startup module paired with a newly deployed secondary module from a different build. Content-hash both files and reference the secondary by its hashed name from the startup code.
- The first export is slow. Lazy loading moved the wait rather than removing it. Prefetch and precompile during idle time.
wasm-splitplaceholders trap. The loader import was not provided or did not patch the table. Test the deferred path in CI, not just startup.
Performance note
On a mid-range phone, the split cut time to interactive from 2.9 s to 1.6 s: the startup module compiled in 310 ms instead of 720 ms, and the download was 344 KB smaller. With idle-time prefetch and precompile, the first export took 140 ms longer than in the unsplit build — invisible next to the export’s own two-second run time.
Frequently Asked Questions
Can two modules share one linear memory?
Yes, if one exports its memory and the other imports it — that is what wasm-split does. Two independently built Rust
modules each expect to own their memory and allocator, so with separate crates you share data by copying.
Does this work with Emscripten?
Emscripten supports wasm-split directly with -sSPLIT_MODULE, including the loader. It also supports dynamic linking of
side modules, described in linking side modules at runtime.
Does splitting affect the browser’s compiled-code cache? Each module is cached separately under its own URL, which works in your favour: changing the deferred feature does not invalidate the cached startup module, so returning users recompile only what changed.
How small should the startup module be? There is no fixed number; measure compile time on a target phone. Below about 100 ms of compile time, further splitting rarely changes what users notice.
Is the component model a better way to split? Components compose separately compiled pieces with typed interfaces, which suits plugin-style splits. Browser support relies on transpilation today, so for page startup, plain module splitting is more practical.
Related
- Analyzing Wasm size with twiggy — finding what to move.
- Reducing Wasm cold-start latency — other startup levers.
- Structuring a monorepo with Rust Wasm and a JS app — managing two crates and packages.
- Exporting a table to JavaScript — the table patching wasm-split relies on.
← Back to Wasm Optimization Flags & Size Reduction