Reducing Instantiation Cost of Large Modules

This page answers one task: a large WebAssembly module compiles quickly (or comes from the code cache), yet WebAssembly.instantiate and the initialisation that follows still take hundreds of milliseconds. You want to know what happens during instantiation, which parts are expensive for your module, and how to shrink them.

Prerequisites

  • [ ] A module whose startup you can measure with compile and instantiate split apart.
  • [ ] wasm-objdump or wasm-tools to inspect data segments and the start function.
  • [ ] Access to the module’s build (Rust, C/C++ or another toolchain).

What instantiation does

Instantiation turns a compiled module into a running instance. The engine resolves imports, allocates linear memory at its initial size, allocates tables and globals, copies active data segments into memory at their offsets, fills tables from active element segments, and finally runs the module’s start function if it has one. After that, toolchain glue usually calls an initialisation export — Emscripten’s runtime init and static constructors, wasm-bindgen’s start function, a _initialize export for WASI reactors — before your first real call.

Several of these steps scale with module content. Copying data segments is a memory copy proportional to their size; a module that embeds 30 MB of tables or assets copies 30 MB at every instantiation. Allocating a very large initial memory costs time and commits memory on some platforms. And start functions or static constructors can run arbitrary amounts of code — parsing embedded data, building lookup tables, initialising a runtime.

Where instantiation time goes Instantiation resolves imports, allocates memory and tables, copies active data segments into memory, initialises tables from element segments, and runs the start function. Toolchain initialisation such as static constructors and runtime setup follows. Data copies and initialisation code usually dominate for large modules. resolve imports cheap allocate memory + tables grows with initial size copy active data segments ∝ static data size start function / constructors arbitrary code toolchain init export runtime setup

Step 1 — measure instantiation separately

const t0 = performance.now();
const module = await WebAssembly.compileStreaming(fetch(url));
const t1 = performance.now();
const instance = await WebAssembly.instantiate(module, imports);
const t2 = performance.now();
instance.exports._initialize?.();            // or the glue's init
const t3 = performance.now();
console.table({ compile: t1 - t0, instantiate: t2 - t1, init: t3 - t2 });

If instantiate is large, look at data segments and the start function; if init is large, look at constructors and runtime setup.

Step 2 — inspect data segments

wasm-objdump -h app.wasm | grep -i data       # size of the Data section
wasm-objdump -x -j Data app.wasm | head        # individual segments with offsets and sizes

Large segments usually come from embedded assets (fonts, models, dictionaries), large constant tables (Unicode data, lookup tables) or zero-filled arrays that the toolchain emitted as data instead of leaving in BSS. Each is copied at instantiation.

Step 3 — move big data out of active segments

For embedded assets, consider fetching them separately (cached independently, loaded when needed) instead of baking them into the module. For data that must stay in the module, passive data segments (from the bulk-memory proposal, supported everywhere current) are not copied at instantiation; code copies them into memory with memory.init when first needed, then drops them with data.drop. Some toolchains and wasm-opt passes can convert or split segments; in C/C++, placing rarely used tables in a separate section that you load lazily achieves a similar effect.

Active versus passive data segments Active data segments are copied into linear memory at every instantiation regardless of whether the data is used. Passive segments stay in the module until code copies them with memory.init when needed and drops them afterwards, so instantiation skips the copy for unused data. active segments copied at instantiation cost ∝ total static data simple, automatic small or always-used data passive segments copied on demand (memory.init) dropped after use instantiation skips them large, rarely used data

Step 4 — make initialisation lazy

Static constructors in C++ and eager initialisation in Rust (lazy_static, once_cell forced at startup, building caches in main or the start function) run before the first call. Move initialisation to first use: a OnceCell initialised the first time a feature needs it, a C++ function-local static instead of a global object. If initialisation is unavoidable and deterministic — parsing a built-in dictionary, warming caches — run it at build time with Wizer, which snapshots the initialised memory into the module; see snapshotting initialized Wasm with Wizer.

Step 5 — right-size initial memory

A module declaring 512 MB of initial memory allocates it at every instantiation, even if the program uses 20 MB. Set initial memory close to the typical working set and allow growth (-sINITIAL_MEMORY and -sALLOW_MEMORY_GROWTH in Emscripten; --initial-memory with wasm-ld). On mobile devices, oversized initial memory can also fail outright.

Instantiating many times

Some designs instantiate a module repeatedly — per request on servers, per document or per worker in browsers. Instantiation cost then multiplies. Besides the steps above, reuse instances where state allows, and in hosts like Wasmtime use instance pre-instantiation (InstancePre) and pooling allocators, which make per-request instantiation take microseconds by doing import resolution and memory setup ahead of time.

Finding what fills the data section

Before moving data, find out what it is. Symbol maps from the linker (wasm-ld --Map, or -Wl,--print-map) list each symbol with its address and size, so the largest static objects stand out. In Rust, large static arrays, include_bytes! assets, and dependencies with built-in tables (Unicode normalisation, time-zone data, regular-expression tables, compression dictionaries) are typical sources; twiggy can show which data symbols are largest when names are kept. In C and C++, large const arrays, embedded resources and initialised global structures dominate. Once you know the culprits, decide per item: drop it if unused (a feature flag in a dependency can often remove built-in tables), fetch it separately if large and optional, load it lazily via a passive segment if it must stay in the module, or keep it if it is small and always needed.

Instantiation on the main thread

Even though WebAssembly.instantiate returns a promise, the work of copying data segments and running start functions may happen on the thread that instantiates — and that is often the main thread. A 300 ms instantiation there blocks input and rendering just like synchronous JavaScript. Instantiating in a worker moves that cost off the main thread; for modules the page needs immediately, instantiate in a worker and expose the API through messages, or at least schedule main-thread instantiation at a moment when a short freeze is acceptable (before the first paint of an app shell, rather than during interaction). Chrome’s Performance panel shows instantiation as a task on the thread that performed it, which makes the effect easy to confirm.

Startup budgets

Set a budget for instantiation plus init on a target device and check it in CI with a headless browser on a throttled profile; growth in static data or new eager initialisers then shows up as a budget failure rather than a slow-start complaint.

Checking the effect in a trace

After each change, record a Performance trace of startup: the instantiate task should shrink, and initialisation work should move to the first use of the feature that needs it, not disappear unaccounted for.

Expected output

The module’s 34 MB of active data — mostly an embedded dictionary — moves to a separately fetched file loaded on first use; static constructors that built lookup tables become lazy; initial memory drops from 256 MB to 32 MB with growth enabled; and instantiation plus init falls from 410 ms to 45 ms on a laptop.

Gotchas

  • Measuring only compile time. Instantiation and init can dominate. Split the measurements.
  • Embedding large assets in the module. Copied every instantiation. Fetch separately or use passive segments.
  • Eager initialisation in start functions. Delays the first call. Make it lazy or snapshot it.
  • Oversized initial memory. Slow and may fail on phones. Size to the working set.
  • Instantiating per request without pooling. Costs multiply. Reuse or pre-instantiate.
  • Instantiating heavy modules on the main thread. Input freezes. Instantiate in a worker.

Performance note

Moving a 34 MB dictionary out of active data segments reduced instantiation from 160 ms to 12 ms; making table construction lazy removed a further 250 ms from initialisation.

Instantiation plus initialisation time after each change Milliseconds from compiled module to ready instance for a large module initially, after moving a 34 MB dictionary out of active data segments, and after making lookup-table construction lazy and reducing initial memory. ms to ready instance initial 410 ms dictionary fetched separately 262 ms + lazy tables, smaller memory 45 ms

Frequently Asked Questions

Does the code cache help instantiation? It removes compilation, not data copies or initialisation code.

Are zero-initialised arrays copied? Not if emitted as BSS (uninitialised memory); some toolchains emit zeros as data — check segment contents.

Can I see the start function? wasm-objdump -x lists a Start section if present; toolchains often use exported init functions instead.

Does instantiation run in the background? WebAssembly.instantiate is asynchronous, but the work still occupies the thread that eventually runs it in many engines; measure.

Does instantiation block the main thread? It can — data copying and start functions run on the instantiating thread; instantiate in a worker for heavy modules.

How do I find which static data is largest? Read the linker’s map file or use twiggy with names kept; large static arrays and embedded tables stand out.

Can dependency features remove built-in tables? Often — many crates offer features that drop Unicode, time-zone or regex tables you do not need.

Should instantiation time be budgeted in CI? Yes — measure instantiation plus init on a throttled profile so growth in static data or eager init fails the budget.

Where does lazy initialisation cost show up instead? On the first use of the feature that needs it — measure that interaction too.

← Back to Module Caching & Startup Performance