Reducing Instantiation Cost of Large Modules
This page answers one task: a large WebAssembly module compiles quickly (or comes from the code cache), yet WebAssembly.instantiate and the initialisation
that follows still take hundreds of milliseconds. You want to know what happens during instantiation, which parts are expensive for your module, and how to
shrink them.
Prerequisites
- [ ] A module whose startup you can measure with compile and instantiate split apart.
- [ ]
wasm-objdumporwasm-toolsto inspect data segments and the start function. - [ ] Access to the module’s build (Rust, C/C++ or another toolchain).
What instantiation does
Instantiation turns a compiled module into a running instance. The engine resolves imports, allocates linear memory at its initial size, allocates tables and
globals, copies active data segments into memory at their offsets, fills tables from active element segments, and finally runs the module’s start function
if it has one. After that, toolchain glue usually calls an initialisation export — Emscripten’s runtime init and static constructors, wasm-bindgen’s start
function, a _initialize export for WASI reactors — before your first real call.
Several of these steps scale with module content. Copying data segments is a memory copy proportional to their size; a module that embeds 30 MB of tables or assets copies 30 MB at every instantiation. Allocating a very large initial memory costs time and commits memory on some platforms. And start functions or static constructors can run arbitrary amounts of code — parsing embedded data, building lookup tables, initialising a runtime.
Step 1 — measure instantiation separately
const t0 = performance.now();
const module = await WebAssembly.compileStreaming(fetch(url));
const t1 = performance.now();
const instance = await WebAssembly.instantiate(module, imports);
const t2 = performance.now();
instance.exports._initialize?.(); // or the glue's init
const t3 = performance.now();
console.table({ compile: t1 - t0, instantiate: t2 - t1, init: t3 - t2 });
If instantiate is large, look at data segments and the start function; if init is large, look at constructors and runtime setup.
Step 2 — inspect data segments
wasm-objdump -h app.wasm | grep -i data # size of the Data section
wasm-objdump -x -j Data app.wasm | head # individual segments with offsets and sizes
Large segments usually come from embedded assets (fonts, models, dictionaries), large constant tables (Unicode data, lookup tables) or zero-filled arrays that the toolchain emitted as data instead of leaving in BSS. Each is copied at instantiation.
Step 3 — move big data out of active segments
For embedded assets, consider fetching them separately (cached independently, loaded when needed) instead of baking them into the module. For data that must
stay in the module, passive data segments (from the bulk-memory proposal, supported everywhere current) are not copied at instantiation; code copies them
into memory with memory.init when first needed, then drops them with data.drop. Some toolchains and wasm-opt passes can convert or split segments; in
C/C++, placing rarely used tables in a separate section that you load lazily achieves a similar effect.
Step 4 — make initialisation lazy
Static constructors in C++ and eager initialisation in Rust (lazy_static, once_cell forced at startup, building caches in main or the start function)
run before the first call. Move initialisation to first use: a OnceCell initialised the first time a feature needs it, a C++ function-local static instead of
a global object. If initialisation is unavoidable and deterministic — parsing a built-in dictionary, warming caches — run it at build time with Wizer, which
snapshots the initialised memory into the module; see
snapshotting initialized Wasm with Wizer.
Step 5 — right-size initial memory
A module declaring 512 MB of initial memory allocates it at every instantiation, even if the program uses 20 MB. Set initial memory close to the typical working
set and allow growth (-sINITIAL_MEMORY and -sALLOW_MEMORY_GROWTH in Emscripten; --initial-memory with wasm-ld). On mobile devices, oversized initial
memory can also fail outright.
Instantiating many times
Some designs instantiate a module repeatedly — per request on servers, per document or per worker in browsers. Instantiation cost then multiplies. Besides the
steps above, reuse instances where state allows, and in hosts like Wasmtime use instance pre-instantiation (InstancePre) and pooling allocators, which make
per-request instantiation take microseconds by doing import resolution and memory setup ahead of time.
Finding what fills the data section
Before moving data, find out what it is. Symbol maps from the linker (wasm-ld --Map, or -Wl,--print-map) list each symbol with its address and size, so the
largest static objects stand out. In Rust, large static arrays, include_bytes! assets, and dependencies with built-in tables (Unicode normalisation,
time-zone data, regular-expression tables, compression dictionaries) are typical sources; twiggy can show which data symbols are largest when names are kept.
In C and C++, large const arrays, embedded resources and initialised global structures dominate. Once you know the culprits, decide per item: drop it if
unused (a feature flag in a dependency can often remove built-in tables), fetch it separately if large and optional, load it lazily via a passive segment if it
must stay in the module, or keep it if it is small and always needed.
Instantiation on the main thread
Even though WebAssembly.instantiate returns a promise, the work of copying data segments and running start functions may happen on the thread that
instantiates — and that is often the main thread. A 300 ms instantiation there blocks input and rendering just like synchronous JavaScript. Instantiating in a
worker moves that cost off the main thread; for modules the page needs immediately, instantiate in a worker and expose the API through messages, or at least
schedule main-thread instantiation at a moment when a short freeze is acceptable (before the first paint of an app shell, rather than during interaction).
Chrome’s Performance panel shows instantiation as a task on the thread that performed it, which makes the effect easy to confirm.
Startup budgets
Set a budget for instantiation plus init on a target device and check it in CI with a headless browser on a throttled profile; growth in static data or new eager initialisers then shows up as a budget failure rather than a slow-start complaint.
Checking the effect in a trace
After each change, record a Performance trace of startup: the instantiate task should shrink, and initialisation work should move to the first use of the feature that needs it, not disappear unaccounted for.
Expected output
The module’s 34 MB of active data — mostly an embedded dictionary — moves to a separately fetched file loaded on first use; static constructors that built lookup tables become lazy; initial memory drops from 256 MB to 32 MB with growth enabled; and instantiation plus init falls from 410 ms to 45 ms on a laptop.
Gotchas
- Measuring only compile time. Instantiation and init can dominate. Split the measurements.
- Embedding large assets in the module. Copied every instantiation. Fetch separately or use passive segments.
- Eager initialisation in start functions. Delays the first call. Make it lazy or snapshot it.
- Oversized initial memory. Slow and may fail on phones. Size to the working set.
- Instantiating per request without pooling. Costs multiply. Reuse or pre-instantiate.
- Instantiating heavy modules on the main thread. Input freezes. Instantiate in a worker.
Performance note
Moving a 34 MB dictionary out of active data segments reduced instantiation from 160 ms to 12 ms; making table construction lazy removed a further 250 ms from initialisation.
Frequently Asked Questions
Does the code cache help instantiation? It removes compilation, not data copies or initialisation code.
Are zero-initialised arrays copied? Not if emitted as BSS (uninitialised memory); some toolchains emit zeros as data — check segment contents.
Can I see the start function?
wasm-objdump -x lists a Start section if present; toolchains often use exported init functions instead.
Does instantiation run in the background?
WebAssembly.instantiate is asynchronous, but the work still occupies the thread that eventually runs it in many engines; measure.
Does instantiation block the main thread? It can — data copying and start functions run on the instantiating thread; instantiate in a worker for heavy modules.
How do I find which static data is largest? Read the linker’s map file or use twiggy with names kept; large static arrays and embedded tables stand out.
Can dependency features remove built-in tables? Often — many crates offer features that drop Unicode, time-zone or regex tables you do not need.
Should instantiation time be budgeted in CI? Yes — measure instantiation plus init on a throttled profile so growth in static data or eager init fails the budget.
Where does lazy initialisation cost show up instead? On the first use of the feature that needs it — measure that interaction too.
Related
- Snapshotting initialized Wasm with Wizer — build-time init.
- Understanding the start function — start semantics.
- Reading the data section — segment details.
- Using bulk memory operations — passive segments.
← Back to Module Caching & Startup Performance