Property-Based Testing for Wasm Modules
This page answers one task: example-based tests for a WebAssembly module pass, but bugs keep appearing for inputs nobody thought to write down — edge values, odd lengths, unusual Unicode — and you want tests that generate inputs automatically and check that the module behaves correctly for all of them.
Prerequisites
- [ ] A WebAssembly module with a clear contract (a parser, codec, numeric kernel or data structure).
- [ ] A test runner for Rust (
cargo test,wasm-bindgen-test) or JavaScript (Vitest, Jest, Node’s test runner). - [ ] Ideally, a reference implementation — the original JavaScript, a native build, or a simpler slow version.
What property-based testing adds
An example-based test checks one input against one expected output. A property-based test states something that must hold for every input — “decoding
an encoded value returns the original”, “the Wasm result equals the reference result”, “output length never exceeds input length plus 16” — and lets a
generator produce hundreds or thousands of inputs, including nasty ones: empty buffers, maximum lengths, NaN and negative zero, surrogate pairs, values on
power-of-two boundaries. When a property fails, the framework shrinks the failing input to the smallest case that still fails, so the report shows
[0, 255] rather than a 4 KB random buffer.
WebAssembly modules are especially good candidates. Their boundary is narrow and typed, making properties easy to state; bugs tend to cluster in exactly
the places generators probe — length calculations, pointer arithmetic, integer overflow on 32-bit usize, encoding conversions between JavaScript strings
and UTF-8 — and a reference implementation often exists, because the module was ported from JavaScript or also builds natively.
Step 1 — choose properties
Good properties for Wasm modules fall into a few families:
- Equivalence with a reference. The Wasm result equals the JavaScript or native result for every input. The strongest property when a reference exists.
- Round trips.
decode(encode(x)) == x,parse(print(ast)) == ast,decompress(compress(b)) == b. - Invariants. Sorted output is sorted and a permutation of the input; a simplified path keeps its endpoints; a hash has a fixed length.
- No crashes. For any input, the function returns a result or a documented error — never a trap, never a hang.
Start with “no crashes” for every exported function — it needs no reference and finds out-of-bounds traps and panics immediately — then add equivalence or round-trip properties for the core behaviour.
Step 2 — test in Rust with proptest
For Rust modules, test the core logic with proptest natively, where runs are fast and debugging is easy:
use proptest::prelude::*;
proptest! {
#[test]
fn roundtrip(data in proptest::collection::vec(any::<u8>(), 0..4096)) {
let compressed = codec::compress(&data);
prop_assert_eq!(codec::decompress(&compressed).unwrap(), data);
}
#[test]
fn never_panics_on_garbage(data in proptest::collection::vec(any::<u8>(), 0..512)) {
let _ = codec::decompress(&data); // Ok or Err, but no panic
}
}
Run the same tests on a 32-bit Wasm target too (cargo test --target wasm32-wasip1 with a Wasmtime runner), because usize overflow bugs only appear
there. proptest stores failing seeds in proptest-regressions/; commit them so every future run retries past failures first.
Step 3 — test the built module from JavaScript with fast-check
Native tests cannot see the boundary — string conversion, memory growth, glue bugs. Test the built module from JavaScript with fast-check:
import fc from "fast-check";
import { test } from "vitest";
import { slugify as wasmSlugify } from "../pkg/text.js";
import { slugify as jsSlugify } from "../legacy/slugify.js";
test("Wasm slugify matches the JavaScript reference", () => {
fc.assert(
fc.property(fc.string({ unit: "grapheme", maxLength: 200 }), (s) => {
return wasmSlugify(s) === jsSlugify(s);
}),
{ numRuns: 2000 },
);
});
fc.string({ unit: "grapheme" }) produces combining characters, emoji and surrogate pairs — exactly what breaks UTF-16 to UTF-8 conversion. A failure
prints the shrunk input and a seed; rerun with { seed, path } to reproduce it exactly.
Step 4 — design generators for the domain
Uniform random bytes rarely produce valid structured input; a parser fed random bytes rejects almost everything at the first byte. Write generators that
produce mostly valid inputs with occasional corruption: generate an abstract value, serialise it, then optionally flip a byte or truncate. For numeric
kernels, bias generators toward boundaries — fc.double() already includes NaN, infinities and negative zero; add fc.constantFrom(0, 1, 2 ** 31 - 1)
for integers. Decide what equality means for floats: bit-identical results between Wasm and JavaScript are achievable for basic arithmetic but not for
transcendental functions, where comparing within a tolerance is the honest property.
Step 5 — run it in CI with a budget
Property tests are slower than example tests, so give them a budget: a few hundred runs per property on every push, many thousands in a nightly job. Log the seed of every CI run so a failure can be reproduced locally with the same inputs. Turn each shrunk failure into a regular example-based test when fixing it, so the regression is checked deterministically forever, independent of the generator.
Testing stateful modules with model-based tests
Modules that hold state — an allocator, a cache, a document model, a database — need sequences of operations, not single inputs. fast-check’s
fc.commands and proptest’s state-machine testing generate random sequences of calls, apply them to both the Wasm module and a simple model (a
JavaScript Map standing in for a Wasm hash table, an array standing in for a rope), and check after each step that observable state matches. When a
sequence fails, shrinking removes irrelevant commands until a minimal sequence remains — typically three or four calls that reveal, for example, that
deleting the last element after a resize corrupts the table. This finds bugs in memory management across calls, which single-call properties never
reach.
Interpreting failures at the boundary
When a property fails only for the built module and not for the native core, the bug is almost always at the boundary: string conversion (lone
surrogates become U+FFFD when encoded to UTF-8, so wasm(s) !== js(s) for strings containing them), integer conversion (a u32 returned as a negative
JavaScript number), or memory views detached by growth. These are worth knowing as expected differences rather than bugs: document lone-surrogate
behaviour, normalise both sides before comparison, or exclude such inputs from the generator explicitly so the remaining property is meaningful.
Sharing generators between Rust and JavaScript tests
A module with both a Rust core and a JavaScript wrapper ends up with two sets of generators that drift apart. Keep a small corpus of interesting inputs —
the shrunk failures, boundary values, real-world samples — in a shared directory of files, and have both proptest and fast-check seed their runs from it
before generating random cases. fast-check accepts explicit examples in fc.assert options, and proptest can combine a prop_oneof! over corpus values
with random generation. The corpus also feeds fuzzers, so an input that broke the module once is retried by every testing layer. Review the corpus when
the input format changes, and delete entries that no longer represent valid or meaningful cases, so it stays small enough to run on every push.
Keeping property tests readable
Properties are code that other developers must maintain, and a clever property that nobody understands gets deleted at the first false alarm. Name each
property after the guarantee it checks (roundtrip_preserves_bytes, matches_js_reference_for_ascii), keep generators in named helper functions, and
write a one-line comment stating why the property should hold. When a property needs exclusions — lone surrogates, NaN payloads — make them explicit
filters with a comment, rather than silently narrowing the generator, so future readers know what is deliberately untested.
Expected output
The codec’s Rust core passes round-trip and no-panic properties for 10,000 generated inputs on native and wasm32-wasip1 targets; the built module passes equivalence with the JavaScript reference for 2,000 grapheme strings per CI run; a stateful test of the cache runs 500 random command sequences; and two past failures live on as regression tests.
Gotchas
- Purely random bytes for structured formats. Everything is rejected early. Generate valid inputs and corrupt them.
- Exact float equality across implementations. Use tolerances for transcendental maths.
- Testing only the native core. Boundary bugs need tests against the built module.
- Lost seeds. Failures cannot be reproduced. Log and store seeds.
- Unbounded input sizes. Runs time out. Cap generator sizes and raise them in nightly jobs.
Performance note
Property tests found three bugs in a ported text module that 140 example tests missed: an off-by-one with combining marks, a panic on lone surrogates, and a length overflow above 64 KB. The CI property job added 9 s to the pipeline; the nightly job runs 50 times more cases in under 8 minutes.
Frequently Asked Questions
Do I need a reference implementation? No — round-trip, invariant and no-crash properties need none. A reference makes the strongest property, though.
How many runs are enough? Hundreds per push, thousands nightly. Coverage of input classes matters more than raw counts.
Is this the same as fuzzing? Related — fuzzing is coverage-guided and runs much longer; property tests run in normal test suites with typed generators.
Can I run fast-check in the browser? Yes; run it inside browser tests when the module depends on browser APIs.
What should I do when a property fails in CI but not locally? Rerun locally with the seed and path printed in the CI log; the same seed reproduces the same inputs.
Related
- Fuzzing a Wasm module — coverage-guided exploration.
- Differential testing Wasm against native builds — comparing builds.
- Keeping JavaScript and Wasm results identical — reference equivalence.
- Testing Wasm modules with Vitest — the runner.
← Back to Testing & Verifying Wasm Builds