Differential Testing Wasm Against Native Builds
This page answers one task: prove that the WebAssembly build of a library behaves the same as its native build, by running both on the same inputs and comparing the outputs — and find the inputs where they differ.
Prerequisites
- [ ] A Rust, C or C++ library that builds both natively and for a Wasm target.
- [ ] A WASI runtime such as
wasmtime, or Node forwasm32-unknown-unknownbuilds. - [ ] A corpus of inputs: test fixtures, a fuzzing corpus, or generated cases.
Why the same code can behave differently
It is tempting to assume that code which passes its tests natively is correct in WebAssembly too: same source, same compiler, same logic. Usually it is. But the WebAssembly target differs from a 64-bit desktop target in ways that occasionally change behaviour, and those differences are invisible to tests that only run natively.
Pointer width is the classic one. wasm32 has 32-bit pointers and a 32-bit usize, so arithmetic that silently relied
on 64-bit sizes can overflow, and a length that fits natively can wrap. In C, long is 32 bits on wasm32 but 64 bits
on 64-bit Linux, which changes the meaning of code that uses it for file offsets or hashes. Floating point is
deterministic in Wasm, but a native build may use fused multiply-add or x87 extended precision where Wasm does not, so
results can differ in the last bits. Unaligned access is legal in Wasm and trapping on some native platforms. And the
toolchain differs: wasm-opt runs only on the Wasm build, and its transformations are not exercised by native tests.
A differential test makes these differences visible without having to predict them. Run both builds on many inputs, compare outputs, and investigate every mismatch.
Step 1 — build both from the same command-line interface
Give the library a small harness binary that reads an input file and writes its result in a canonical format. The same source then compiles to a native executable and to a WASI module:
// src/bin/harness.rs
use std::io::{Read, Write};
fn main() {
let mut input = Vec::new();
std::io::stdin().read_to_end(&mut input).unwrap();
let result = mylib::process(&input);
let out = match result {
Ok(doc) => format!("ok {}\n{}", doc.checksum(), doc.summary()),
Err(e) => format!("err {}\n", e.code()), // compare error *kinds*, not messages
};
std::io::stdout().write_all(out.as_bytes()).unwrap();
}
cargo build --release --bin harness
cargo build --release --bin harness --target wasm32-wasip1
Writing the result in a stable, textual form — a checksum plus a summary, or error codes rather than messages — makes the comparison a plain text diff and keeps it immune to harmless formatting differences.
Step 2 — run both over a corpus and diff
#!/usr/bin/env bash
# scripts/differential.sh
set -uo pipefail
NATIVE=target/release/harness
WASM=target/wasm32-wasip1/release/harness.wasm
fails=0
for f in tests/corpus/*; do
a=$($NATIVE < "$f")
b=$(wasmtime run "$WASM" < "$f")
if [[ "$a" != "$b" ]]; then
echo "MISMATCH $f"; diff <(echo "$a") <(echo "$b") | head -5
fails=$((fails+1))
fi
done
echo "$fails mismatches across $(ls tests/corpus | wc -l) inputs"
exit $(( fails > 0 ))
Start the WASI runtime once per input for simplicity; for large corpora, a harness that loops over files inside one process is much faster. The corpus can be the fuzzing corpus from fuzzing a Wasm module, which is a ready-made collection of inputs that reach deep code paths.
Step 3 — compare floats with intent
If outputs contain floating-point values, decide what equality means before the first mismatch rather than after. For
code that must be bit-exact across platforms — a deterministic simulation, a lockstep game — compare bits and treat any
difference as a bug; then make the native build bit-exact too, for example by disabling FMA contraction with
-C target-feature=-fma or -ffp-contract=off. For numeric code where small differences are acceptable, compare with a
relative tolerance and report the largest error seen:
fn close(a: f64, b: f64) -> bool {
a == b || ((a - b).abs() / a.abs().max(b.abs())) <= 1e-12
}
Whatever you choose, make the harness print floats in a form that preserves all bits — {:?} in Rust prints the shortest
representation that round-trips, which is ideal.
Step 4 — fuzz differentially
The strongest version of this test generates inputs rather than reading a fixed corpus. A fuzzer can drive the native build while a differential oracle checks each input against the Wasm build — slow, because the Wasm side runs in a runtime, but very effective at finding the rare inputs where the builds diverge. A cheaper variant runs the fuzzer natively for a while, then replays the whole corpus it built through the differential script above.
For pointer-width bugs, generate large inputs deliberately. Many usize overflow bugs only appear when a length or an
offset exceeds 2³², which a 32-bit target cannot even represent — the Wasm build then fails with an allocation error or
a trap where the native build succeeds. Decide whether that is acceptable for your library and make the Wasm build fail
cleanly with an error rather than a trap.
Step 5 — turn every mismatch into a test
When the differential run finds a mismatch, minimise the input, understand the cause, fix it, and add the input to the regular test suite — on both targets. Over time the corpus of mismatches becomes a record of every way your code was target-dependent.
#[test]
fn large_record_count_does_not_overflow_on_wasm32() {
// found by differential testing: count * 24 overflowed usize on wasm32
let input = include_bytes!("fixtures/diff-0007-record-count.bin");
assert!(matches!(mylib::process(input), Err(e) if e.code() == "TooLarge"));
}
Where to run it
Differential tests are cheap enough for CI once the harness runs in a single process, and they belong on two kinds of change in particular. The first is any change to code that deals in sizes, offsets and lengths — parsers, allocators, binary formats — because that is where pointer-width bugs live. The second is a toolchain upgrade: a new rustc, LLVM or Binaryen release changes the Wasm build without touching the source, and a differential run across the full corpus is the most direct evidence that nothing changed in behaviour. Run it on a schedule as well, with whatever new inputs the fuzzer has found since the last run, so the corpus keeps growing and the comparison keeps finding new edge cases.
Expected output
A clean run across the corpus:
0 mismatches across 1842 inputs
A run that finds a real difference:
MISMATCH tests/corpus/rec-3f9a01
1c1
< ok 9c41e7a0 records=178956971
---
> err TooLarge
1 mismatches across 1842 inputs
Here the native build accepted an input claiming 178 million records because 64-bit arithmetic did not overflow; the Wasm build rejected it. Both should reject it — the native behaviour was the bug, and the Wasm build exposed it.
Gotchas
- Error messages differ and every error input mismatches. Messages include paths or platform text. Compare error codes or kinds instead.
- Hash-map iteration order differs. Output that iterates a
HashMapis nondeterministic on every target. Sort before printing, or use aBTreeMap. - The WASI run fails on file access. The harness reads stdin, so no preopens are needed; if it opens files, grant the directory as described in granting filesystem access with WASI preopens.
- The browser build differs from the WASI build. Rare for pure computation, possible for code using
target-specific paths. Add a Node-based harness for
wasm32-unknown-unknownif browser behaviour is what you ship.
Performance note
Running the 1,842-input corpus through both builds took 41 s with one wasmtime process per input, almost all of it
process startup. A single-process harness that compiled the module once and looped over the inputs inside it finished in
3.2 s, which made it cheap enough to run on every pull request.
Frequently Asked Questions
Is this worth it for pure Rust code?
Less than for C, because safe Rust prevents many target-dependent bugs. usize overflow and float differences still
apply, so for libraries that process untrusted or large inputs it remains worthwhile.
Can I compare against an older version instead of the native build? Yes — that is regression testing, and the same harness works: run the old and new Wasm builds side by side. It is a good way to validate a toolchain upgrade.
What about comparing two Wasm runtimes? Comparing wasmtime, Node and a browser on the same module catches engine bugs, which are rare but real. It is a cheap addition once the harness exists.
Does wasm-opt ever introduce differences? Very rarely, and when it does it is a Binaryen bug worth reporting. Running the differential test on the optimized module is how you would notice.
Related
- Fuzzing a Wasm module — a source of inputs for the comparison.
- Snapshot testing Wasm output — comparing against stored results instead.
- Understanding Wasm linear memory limits — why 32-bit addressing caps input sizes.
- Running Wasm modules with the wasmtime CLI — the runtime used for the Wasm side.
← Back to Testing & Verifying Wasm Builds