Snapshotting Initialized Wasm with Wizer

This page answers one task: a WebAssembly module spends a long time initialising before it can do useful work — parsing embedded data, building lookup tables, booting an interpreter, loading a standard library — and that work is the same on every start. You want to do it once, at build time, and ship a module that starts already initialised.

Prerequisites

  • [ ] Wizer (cargo install wizer --features env_logger,structopt, or the version bundled with your Wasmtime tooling).
  • [ ] A module whose initialisation is deterministic and does not depend on runtime inputs.
  • [ ] A WASI-compatible build (Wizer runs the module in a WASI environment).

How snapshotting works

Wizer instantiates your module at build time, calls an initialisation function you export, and then takes a snapshot of the instance’s state: the contents of linear memory and the values of globals. It writes a new module whose data segments contain that snapshot and whose globals start at the snapshotted values, removing the need to run the initialisation again. At runtime, instantiating the new module copies the snapshot into memory — a fast memory copy — and the instance is immediately in the post-initialisation state.

The technique trades initialisation work for initialisation data: the module gets larger by however much memory initialisation produced, but startup skips the computation. For interpreters (JavaScript, Python, Lua engines compiled to Wasm), it is transformative — the engine’s startup, which can take hundreds of milliseconds, becomes a data copy. ComponentizeJS and componentize-py use exactly this to make components start fast.

Build-time initialisation with Wizer Wizer instantiates the original module at build time and calls its exported initialisation function, which parses data and builds tables in linear memory. Wizer snapshots memory and globals and writes a new module with the snapshot as data segments. At runtime the new module instantiates directly into the initialised state. original module exports wizer.initialize Wizer runs init at build time snapshot memory + globals post-init state write new module snapshot as data runtime: instant start no init work

Step 1 — export an initialisation function

In Rust, export a function named wizer.initialize that performs the expensive setup into global state:

use std::sync::OnceLock;

static DICT: OnceLock<Dictionary> = OnceLock::new();

#[export_name = "wizer.initialize"]
pub extern "C" fn init() {
    DICT.set(Dictionary::parse(include_bytes!("dict.txt"))).ok();   // expensive, deterministic
}

#[no_mangle]
pub extern "C" fn lookup(ptr: *const u8, len: usize) -> u32 {
    let word = unsafe { std::slice::from_raw_parts(ptr, len) };
    DICT.get().expect("initialised by wizer").score(word)
}

In C, export a function with __attribute__((export_name("wizer.initialize"))) that fills global structures.

Step 2 — run Wizer

cargo build --release --target wasm32-wasip1
wizer target/wasm32-wasip1/release/dict.wasm -o dict.initialized.wasm --allow-wasi
wasm-opt -O3 dict.initialized.wasm -o dict.final.wasm          # optional, after snapshotting

--allow-wasi lets initialisation use WASI (reading preopened files, for example); with --dir, initialisation can read files from the build machine. The output module no longer needs initialisation — but the init function remains exported unless removed, so make sure runtime code does not call it again.

Initialising at startup versus snapshotting with Wizer Initialising at startup runs the parsing and table building every time an instance is created, keeping the module small. Snapshotting with Wizer runs it once at build time and stores the resulting memory in the module, making startup a memory copy at the cost of a larger file. initialise at startup work on every instantiation small module startup ∝ init work cheap init only Wizer snapshot work once at build time larger module (snapshot data) startup ≈ memory copy expensive, deterministic init

Step 3 — know what cannot be snapshotted

Snapshots capture linear memory and globals, nothing else. Initialisation must not depend on things that differ at runtime: the current time, random numbers (seeded generators get snapshotted with their seed — every instance would produce the same “random” sequence), environment variables, or host state. References to host objects (externref values, imported tables’ contents provided by the host) cannot be snapshotted. And because the snapshot is taken in Wizer’s environment, initialisation that calls imports other than WASI fails — browser-oriented modules that call JavaScript through wasm-bindgen during initialisation need restructuring so the snapshotted part is pure computation.

Step 4 — check size and compression

The snapshot adds the memory that initialisation populated. Wizer only writes non-zero regions, and snapshots of parsed data structures often compress well, so compare the compressed size before and after. If the snapshot is much larger than the original input data (a 2 MB dictionary becoming 30 MB of hash tables), consider a more compact in-memory representation, or snapshot only part of the initialisation.

Step 5 — measure the gain

Measure time from instantiation to the first useful call before and after snapshotting, on the target platform. For servers that instantiate per request, multiply by request rate; for browsers, compare startup on a mid-range phone, where initialisation work is slowest. Remember that a larger module also costs download and compile time; the net gain is what matters.

Using Wizer for browser modules

Wizer produces standard Wasm modules, which browsers can load. The constraint is the build-time environment: initialisation runs under WASI, not in a browser, so it cannot touch browser APIs. Modules built for wasm32-unknown-unknown with wasm-bindgen imports cannot be snapshotted directly if those imports are needed during initialisation. A workable pattern is a small pure-computation “core” module, built for WASI or with no imports at all, initialised by Wizer, and used from JavaScript through raw exports.

Server-side benefits

Snapshotting pays off most where instances are created constantly. Hosts that create a fresh instance per request — common for isolation in edge platforms and plugin systems — would otherwise run initialisation on every request. With a snapshot, per-request instantiation is a memory copy, and runtimes with copy-on-write memory initialisation (Wasmtime can map a module’s data segments so pages are copied only when written) make even that copy nearly free. The combination of Wizer snapshots and copy-on-write instantiation is what lets interpreter-based guests — a JavaScript engine, a Python interpreter — handle requests with microsecond-scale instantiation, which would be impossible if the interpreter booted per request.

Keeping snapshots correct across builds

A snapshot is a build artefact derived from both the code and the initialisation inputs. If the embedded dictionary changes but the build forgets to re-run Wizer, the module ships stale data. Make Wizer an unconditional step in the build pipeline after compilation, never a manual step, and add a test that checks behaviour that depends on initialised data (a lookup for a word added in the latest dictionary). Record the inputs’ hashes in a custom section or a build manifest, so it is possible to confirm which data a shipped module contains. When debugging, keep an un-snapshotted build available: if a bug disappears when initialisation runs at startup, the problem is in what was captured, often non-determinism that the snapshot froze.

Partial snapshots

Not all initialisation needs to be snapshotted. Split initialisation into a deterministic part (parse built-in data, build tables) run by Wizer, and a runtime part (read configuration, seed randomness, connect to host services) run at startup. Keeping the runtime part small preserves most of the gain while avoiding frozen state that should differ per instance.

Where the init function lives

Keep the wizer.initialize function next to the code it prepares and document what it may and may not do, so later changes do not add non-deterministic work to it by accident.

Expected output

The dictionary module’s startup falls from 220 ms of parsing to 6 ms of instantiation; the module grows from 2.4 MB to 9.1 MB uncompressed but only from 0.9 MB to 1.6 MB with Brotli; lookups return identical results before and after snapshotting; and the build pipeline runs Wizer after cargo build and before wasm-opt.

Gotchas

  • Non-deterministic initialisation. Time, randomness or environment get frozen. Keep init pure.
  • Seeded RNGs in snapshots. Every instance repeats the sequence. Reseed at runtime.
  • Calling non-WASI imports during init. Wizer cannot provide them. Restructure.
  • Ignoring size growth. Download and compile cost rise. Compare net startup.
  • Calling the init function again at runtime. Double initialisation. Guard or remove it.
  • Manual Wizer steps. Stale snapshots ship. Run Wizer unconditionally in the build.

Performance note

On a mid-range phone, the dictionary module reached its first lookup in 310 ms without snapshotting and in 28 ms with a Wizer snapshot, including the extra time to compile the larger module.

Time to first lookup on a mid-range phone Milliseconds from starting instantiation to the first dictionary lookup for the module initialised at startup and for the Wizer-snapshotted module, including compile time differences. ms to first lookup initialise at startup 310 ms Wizer snapshot 28 ms

Frequently Asked Questions

Does Wizer work with Emscripten output? With WASI-compatible builds that do not need JavaScript glue during initialisation; Emscripten’s browser builds generally do not fit.

Can components be snapshotted? Tools such as ComponentizeJS snapshot their engines; direct Wizer use targets core modules, with component support evolving.

Is the snapshot portable across engines? Yes — it is ordinary data segments in a standard module.

What about thread-local state? Snapshots capture the main instance’s memory; per-thread state must be initialised at runtime.

How do snapshots help per-request servers? Instantiation becomes a memory copy — nearly free with copy-on-write memory initialisation — instead of rerunning initialisation every request.

How do I know which data a shipped snapshot contains? Record input hashes in a custom section or build manifest, and test behaviour that depends on the latest data.

Should all initialisation be snapshotted? No — snapshot the deterministic part and keep per-instance setup such as seeding randomness at runtime.

How do I debug a snapshot problem? Compare with an un-snapshotted build; if the bug disappears, the snapshot captured non-deterministic state.

Does copy-on-write help in browsers? Browsers manage memory differently; the main browser gain is skipping initialisation work, not the copy itself.

← Back to Module Caching & Startup Performance