Testing & Verifying Wasm Builds

A WebAssembly module has more places to go wrong than an ordinary library: the logic can be correct while the compiled artifact is broken, the artifact can be valid while the glue that loads it is wrong, and everything can work while the module has quietly grown to three times the size it should be. A testing strategy that only covers the first of those leaves the ones that actually reach users uncovered.

Prerequisites

  • [ ] A build you can run from a single command — testing starts with reproducibility.
  • [ ] A headless browser available in CI, for the tests that genuinely need one.
  • [ ] The WebAssembly Binary Toolkit for wasm-validate and wasm-objdump.
  • [ ] A baseline to compare against: a committed size number and a set of fixtures.

Four layers, four kinds of bug

Think of it as a pyramid whose layers catch different failures and cost different amounts to run.

Native tests compile the same source for the host and run it with the language’s ordinary test runner. They are fast, they debug well, and they cover essentially all of the logic. This is where the majority of tests should live.

Module tests run the compiled .wasm under a standalone runtime such as wasmtime, which catches anything that differs between the host target and the WebAssembly target: integer width assumptions, alignment, endianness in serialisation code, and library functions that behave differently.

Browser tests load the module in a real browser through its real glue, and catch the wiring — a missing export, a bundler that did not emit the binary, an import the page does not provide.

Artifact checks inspect the binary itself: it validates, it imports only what you expect, it exports what the interface promises, and it has not grown.

Four layers, cheapest first Native tests are fastest and cover logic. Module tests under a standalone runtime catch target differences. Browser tests catch wiring. Artifact checks inspect the binary for validity, imports and size. native tests — milliseconds, most of your coverage catches: logic errors module tests under a standalone runtime catches: target-specific behaviour browser tests catches: wiring, glue, bundling artifact checks catches: validity, imports, size — seconds, and the layer teams most often skip entirely

Native first, always

The single most productive testing decision is structuring the code so most of it can be tested without WebAssembly at all. That means a core library with no bindings, no wasm_bindgen attributes and no browser assumptions, plus a thin wrapper that adapts it.

// crates/core/src/lib.rs — pure, testable anywhere
pub fn summarise(records: &[Record]) -> Summary { … }

#[cfg(test)]
mod tests {
    use super::*;
    #[test] fn empty_input_is_zero() { assert_eq!(summarise(&[]).count, 0); }
    #[test] fn weights_are_applied() { … }
}
cargo test                 # runs on the host, in milliseconds

Tests at this layer run in a fraction of a second, debug with an ordinary debugger, and do not need a browser or a runtime. A codebase where they cover the logic can afford to have very few tests at the slower layers, which is what keeps the suite fast enough that people run it.

Where the target genuinely differs

Some bugs only appear once compiled for WebAssembly, and they cluster in predictable places.

Pointer width is 32 bits, so code that assumed usize is 64 bits — a hash mixing constant, a bit shift, an offset calculation — behaves differently. Floating-point is IEEE 754 everywhere, but NaN bit patterns and some conversions are specified differently from what a native target may do. And any code that reads or writes a binary format has to agree about endianness, which WebAssembly fixes as little-endian.

Running the compiled module under wasmtime in CI catches all of these, and costs a second:

cargo build --release --target wasm32-wasip1
wasmtime run --dir=./fixtures target/wasm32-wasip1/release/engine-tests.wasm

Building a small test harness as a WASI executable — one that runs your fixtures and exits non-zero on failure — gives you the whole native suite executing as real WebAssembly, which is a remarkably good return for the effort.

Browser tests, kept few and meaningful

Browser tests are slow, flaky when written carelessly, and irreplaceable for what they cover. Keep them focused on the things only a browser can tell you.

// wasm-bindgen-test: runs in a real headless browser
use wasm_bindgen_test::*;
wasm_bindgen_test_configure!(run_in_browser);

#[wasm_bindgen_test]
fn exports_are_reachable() {
    assert_eq!(crate::add(2, 3), 5);
}

#[wasm_bindgen_test]
async fn instantiates_and_runs_a_fixture() {
    let result = crate::process(&FIXTURE).await;
    assert_eq!(result.count, 1024);
}
wasm-pack test --headless --chrome

The tests worth writing here are: the module loads at all, the exports the glue expects exist, a fixture produces the expected result end to end, and the page still works when the module fails to load. Four tests, each of which has caught a real production failure somewhere.

Testing the JavaScript side of the boundary

The module is half the system. The glue that loads it, marshals arguments and interprets results is the other half, and it is written in a dynamically typed language with no compiler to catch a mistake.

Three kinds of test are worth having for it. A contract test asserts that the loader calls the exports the module actually has — read them from the compiled binary rather than from a hand-written list, so the test fails when the module changes and the loader does not.

const bytes = await readFile('dist/engine.wasm');
const module = await WebAssembly.compile(bytes);
const exported = new Set(WebAssembly.Module.exports(module).map((e) => e.name));
for (const required of ['memory', 'alloc', 'process', 'result_len']) {
  expect(exported.has(required)).toBe(true);
}

A marshalling test round-trips representative values through the boundary and back, including the awkward ones: empty input, a value at the maximum length, a string with multi-byte characters, a float that is exactly representable and one that is not. Boundary marshalling bugs are almost always about edge values rather than typical ones.

And a lifetime test exercises the sequence that causes detached-buffer bugs: allocate, call something that grows memory, then read through a view built before the growth. If your loader rebuilds views correctly the test passes trivially; if it caches them, the test fails now rather than in production under load.

Checking the artifact itself

The binary is a build output like any other, and three checks on it belong in CI.

Validation confirms it is well-formed WebAssembly — usually true, and occasionally not after an aggressive post-processing step:

wasm-validate dist/engine.wasm || exit 1

An import allowlist confirms the module asks for nothing unexpected, which catches a dependency that quietly pulled in WASI or an Emscripten escape hatch:

wasm-objdump -x dist/engine.wasm | awk '/^Import\[/,/^Export\[/' | grep -oP '<\K[^>]+' | sort > imports.txt
diff -u expected-imports.txt imports.txt || { echo "unexpected imports"; exit 1; }

And a size check catches the regression that nobody notices:

SIZE=$(brotli -q 11 -c dist/engine.wasm | wc -c)
BASE=$(cat .size-baseline)
awk -v s="$SIZE" -v b="$BASE" 'BEGIN { exit (s <= b * 1.05) ? 0 : 1 }' \
  || { echo "size regression: $SIZE vs baseline $BASE"; exit 1; }

Fuzzing, for anything parsing untrusted input

A module that parses a file format, decodes a protocol or processes user-supplied bytes should be fuzzed. The sandbox means a crash cannot escape, but a panic is still a denial of service and an out-of-bounds read still returns wrong data.

cargo fuzz run parse_record -- -max_total_time=300

Fuzz the native build, because it is far faster and the logic is the same, then keep any crashing input as a regression fixture that runs in the ordinary suite. Five minutes of fuzzing per CI run finds the inputs no human would have thought to write.

What a pipeline looks like A build runs native tests first because they are fastest, then module tests under a runtime, then artifact validation and size checks, and finally the slower browser tests. Each stage fails fast. native tests 3 s module under wasmtime 8 s artifact checks 2 s browser tests 45 s Ordering matters: the cheap stages fail in seconds and the expensive one runs only on code that has already passed everything else. Total under a minute, which is the difference between a suite people run and one they route around.

Debugging a failure, layer by layer

When something fails, the layer that caught it tells you where to look, which is most of the value of having distinct layers at all.

A native test failure is an ordinary bug: attach a debugger, step through, fix it. Nothing about WebAssembly is involved.

A module test that fails while the native test passes points at a target difference. The usual suspects, in order of likelihood: pointer width, an assumption about usize, alignment in a struct layout, or a library function whose WebAssembly implementation differs. Dumping the module and reading the relevant function is often faster than guessing, and wasm-objdump -d produces readable output for a small function.

A browser test that fails while the module test passes is almost never a logic problem. Check, in order: did the .wasm load at all (network panel), did instantiation succeed (an exception in the console), does the glue call the export by the right name, and does the page supply every import the module declares. One of those four is the answer nearly every time.

An artifact check failure is the easiest of all, because the check reports exactly what changed — an unexpected import names itself, and a size regression gives you two numbers and a commit range.

# which commit grew it, when the baseline check finally fires
git bisect start HEAD HEAD~20
git bisect run sh -c 'cargo build --release --target wasm32-unknown-unknown &&
  test $(brotli -q 11 -c target/wasm32-unknown-unknown/release/engine.wasm | wc -c) -lt 20000'

Reproducible builds make everything else easier

A test suite is only meaningful against a build you can reproduce. Pin the toolchain version, pin the wasm-opt version, and check that two builds of the same commit produce identical bytes.

cargo build --release --target wasm32-unknown-unknown
sha256sum target/wasm32-unknown-unknown/release/engine.wasm > a.sha
cargo clean && cargo build --release --target wasm32-unknown-unknown
sha256sum target/wasm32-unknown-unknown/release/engine.wasm > b.sha
diff a.sha b.sha && echo "reproducible"

Non-reproducible output usually comes from an embedded path, a timestamp or a build identifier, all of which are worth removing: a deterministic artifact makes the size baseline meaningful, makes caching effective, and makes “is this the binary we tested” answerable.

Keeping fixtures honest

Fixtures are the backbone of every layer here, and they degrade in predictable ways.

Generate them once, from a source you trust, and commit them as files. A fixture produced by the code under test is a snapshot of current behaviour rather than a statement of correct behaviour, and it will happily lock in a bug. Where the correct answer comes from a reference implementation — a native library, a Python script, a published test vector — record where it came from in a comment next to the file.

Keep them small enough to read. A fixture nobody can inspect is one nobody will update correctly when the format changes, and a hundred-megabyte binary in the repository slows every clone for the life of the project. If a large input is genuinely needed, generate it deterministically from a seed at test time and commit the seed rather than the data.

Add a fixture for every bug you fix. That is the cheapest test to write — you already have the input that broke it — and it is the one that prevents the same regression twice. Over a couple of years these accumulate into a suite that reflects what actually goes wrong in your particular module, which no amount of speculative testing achieves.

Finally, review fixtures when the format changes. A stale fixture that still passes because the parser is lenient is worse than no fixture, because it creates confidence that nothing checked.

What each layer catches Validation catches a malformed binary, unit tests catch logic, browser tests catch host integration, and a size budget catches the regression nobody notices until the download gets slow. wasm-validate structural validity — a binary no engine would accept wasm-bindgen-test logic, run under Node or a real browser engine headless browser run host integration: workers, memory, DOM interaction size budget the slow regression nobody watches for Each layer is cheap enough to run on every commit; the expensive one is the browser matrix, so stage it. A build that passes all four can still be wrong — these catch classes of failure, not correctness.

Gotchas and failure modes

  • Testing only the native build. Misses everything that differs about the target.
  • Testing only in the browser. Slow, and debugging a failure is far harder than it needs to be.
  • No import allowlist. A dependency adds a WASI import and the module stops working on a host that does not provide it.
  • No size baseline. Payload regressions accumulate silently.
  • Fuzzing the WebAssembly build. Far slower than fuzzing natively for the same coverage.
  • Fixtures generated by the code under test. They will agree with each other forever, including when both are wrong.

Verifying it all works together

The final check is an end-to-end one that exercises the real artifact through the real loading path:

// playwright: the built page, the built module, no mocks
await page.goto('/app');
await page.waitForFunction(() => window.__engineReady === true, { timeout: 5000 });
const result = await page.evaluate(() => window.engine.process(FIXTURE));
expect(result.count).toBe(1024);

One such test per feature is enough. Its job is not coverage — the layers below provide that — but confirming that the pieces someone wired together are still wired together, which is precisely the thing unit tests cannot see.

Guides in this topic

Frequently Asked Questions

How much of the suite should be browser tests? As little as covers the wiring — typically under ten tests even for a substantial module. Everything else belongs at a layer that runs in milliseconds.

Do I need to test in every browser? For the module, rarely: WebAssembly semantics are specified and engines agree. For the glue and the loading path, yes — that is where browser differences actually appear.

What about testing performance? Separately, and with percentiles rather than a mean. A performance test in the correctness suite makes the suite flaky; a tracked benchmark with a threshold is the right shape, as described in building a reproducible benchmark harness.

Should fixtures live with the code or the tests? With the tests, committed as files rather than generated. A fixture generated at test time by the same code under test proves only that the code agrees with itself.

How do I test code that only runs in a worker? Give the worker a message-based interface and test that interface directly from the main thread in a browser test — the worker is then just an implementation detail. For the logic inside, keep it in the pure core and test it natively, where there is no worker at all.

What about testing the failure paths? Deliberately, and with fixtures that trigger them: a truncated input, an oversized allocation request, an input that trips every validation rule. Failure paths are the least exercised code in most modules and the most likely to be reached by an attacker or by a bad day.

How do I stop the suite rotting? Run every layer on every change rather than only the fast ones, and keep the total under a minute so that is practical. A layer that runs only on a nightly job stops being trusted within a month, and a layer nobody trusts is one nobody fixes.

Is snapshot testing useful here? For the summary a module produces, yes — a committed expected output that the suite compares against catches unintended behaviour changes cheaply. For the binary itself, no: a byte-level snapshot of the module changes on every toolchain update and teaches nothing, which is why the artifact checks above look at properties rather than at bytes.

Should the module’s version appear in its output? It is worth exporting, and worth logging. When a bug report arrives with a result attached, knowing which build produced it removes the first hour of every investigation, and a single exported integer costs nothing.

Does any of this change with the component model? The layers do not. What changes is that the artifact checks get better: a component declares its interface in a typed form, so verifying that a build still satisfies its contract becomes a tool’s job rather than a hand-written import comparison.

The strategy in one line: test the logic where it is cheapest, test the target where it differs, test the wiring where it breaks, and check the artifact because nothing else will.

← Back to Compilation Pipelines & Toolchain Setup