Fuzzing a Wasm Module

This guide answers one task: fuzz the code that goes into a WebAssembly module, so that inputs which panic, hang or read out of bounds are found by a machine rather than by a user.

Prerequisites

  • [ ] A Rust crate with a parsing or processing entry point that takes bytes.
  • [ ] cargo-fuzz and a nightly toolchain, which the fuzzer requires.
  • [ ] A few minutes of CPU per run; fuzzing rewards time.
  • [ ] Somewhere to keep a corpus, ideally in version control.

Fuzz the native build, not the module

The instinct is to fuzz the compiled .wasm, and it is the wrong one. A native fuzzing build runs hundreds of times faster, integrates with sanitizers that detect out-of-bounds access precisely, and uses coverage instrumentation to steer input generation. The code is the same code; only the target differs.

The exception is a bug that exists only in the WebAssembly build — a pointer-width assumption, an alignment issue — and those are better found by running the ordinary test suite under wasmtime as described in the topic overview than by fuzzing slowly.

Fuzz where it is fast, test where it differs A native fuzzing build runs with coverage instrumentation and sanitizers at high speed. Fuzzing the compiled module is far slower and lacks precise memory diagnostics, so target differences are better caught by running the test suite under a runtime. native fuzzing build coverage-guided, sanitizers on tens of thousands of cases per second exact diagnosis of a bad access use this fuzzing the .wasm no sanitizer, coarse coverage hundreds of cases per second a trap tells you little about why rarely worth it

A first fuzz target

cargo-fuzz generates a harness crate; the target is a function that takes arbitrary bytes and does something with them. It must not panic on valid input and must not crash on any input.

cargo install cargo-fuzz
cargo fuzz init
cargo fuzz add parse_record
// fuzz/fuzz_targets/parse_record.rs
#![no_main]
use libfuzzer_sys::fuzz_target;

fuzz_target!(|data: &[u8]| {
    // the same function the module exports, called directly
    let _ = my_crate::parse_records(data);
});
cargo fuzz run parse_record -- -max_total_time=60

The let _ = is deliberate: the function returns a Result and an error is a correct outcome for malformed input. What the fuzzer is looking for is a panic, an out-of-bounds access, or a hang — not a returned error.

Structured input, for formats with structure

Raw bytes find shallow bugs quickly and then plateau, because a random byte string rarely forms a valid header. Generating structured input lets the fuzzer spend its time on the interesting parts.

use arbitrary::Arbitrary;

#[derive(Arbitrary, Debug)]
struct Input {
    version: u8,
    record_count: u16,
    flags: u32,
    payload: Vec<u8>,
}

fuzz_target!(|input: Input| {
    let bytes = encode(&input);              // your own serialiser produces a plausible file
    let _ = my_crate::parse_records(&bytes);
});

That shifts the fuzzer from “can it survive garbage” to “can it survive plausible-but-wrong files”, which is where the bugs that reach production usually live: a record count that does not match the payload length, a version the parser half-supports, a flag combination nobody considered.

Keep a raw-bytes target as well. The two find different things and both are cheap.

Making the code fuzzable

A fuzz target is only as useful as the function underneath it, and a few design habits make the difference between a fuzzer that finds real bugs and one that spends its time on the wrong thing.

Separate parsing from processing. A function that takes bytes and produces a validated structure is trivially fuzzable; one that takes bytes, parses them, opens a file and writes a result is not. The boundary you want is the one where untrusted input becomes trusted data, and having it named as a function is worth the refactor on its own.

Bound every allocation derived from input. A length field read from the input and passed straight to Vec::with_capacity is an out-of-memory condition waiting for a four-byte value, and the fuzzer will find it in seconds — which is useful, but the fix is a check rather than a retry:

const MAX_RECORDS: usize = 1 << 20;

let count = read_u32(&data[4..8])? as usize;
if count > MAX_RECORDS {
    return Err(ParseError::TooManyRecords(count));    // not a panic, not an allocation
}
let mut records = Vec::with_capacity(count);

Avoid unbounded recursion on input structure. A nested format with no depth limit is a stack overflow waiting for a crafted file, and in WebAssembly a stack overflow is a trap rather than a catchable error. Track depth explicitly and reject past a sensible limit.

And keep the function deterministic. A fuzz target that consults a clock or a random source produces findings that cannot be reproduced, which wastes everyone’s time — including the fuzzer’s, since its coverage feedback becomes noise.

What a crash means in a sandbox

WebAssembly’s memory safety changes what a fuzzing finding implies, and it is worth being precise.

An out-of-bounds read or write inside linear memory is contained — it cannot touch the host — but it is still a bug: it reads or corrupts other data belonging to your own module, which for a multi-tenant host is a data leak between tenants and for a single user is silently wrong output.

A panic becomes a trap, which JavaScript catches as an exception. That is a denial of service for the feature, not for the page, provided the caller handles it. A caller that does not handle it propagates the failure upward.

A hang is the worst of the three, because nothing catches it: a loop that never terminates holds the thread until something terminates the worker. Fuzzers detect these with a timeout, and finding them before release is the entire argument for running a fuzzer on anything that parses untrusted input.

==12345== ERROR: libFuzzer: deadly signal
    #0 my_crate::parse_records::h3f8a
    ...
SUMMARY: libFuzzer: deadly signal
artifact_prefix='fuzz/artifacts/parse_record/'
Test unit written to fuzz/artifacts/parse_record/crash-8f3a91c2

Turning findings into fixtures

Every crash the fuzzer finds should become a permanent test, in the fast suite, so the same bug cannot return quietly.

# reproduce it
cargo fuzz run parse_record fuzz/artifacts/parse_record/crash-8f3a91c2

# minimise it to the smallest input that still crashes
cargo fuzz tmin parse_record fuzz/artifacts/parse_record/crash-8f3a91c2
// tests/regressions.rs — runs in milliseconds, forever
#[test]
fn issue_214_truncated_header_does_not_panic() {
    let bytes = include_bytes!("fixtures/crash-8f3a91c2.bin");
    assert!(my_crate::parse_records(bytes).is_err());
}

Minimising first matters: the raw artifact is often hundreds of bytes of noise around a two-byte trigger, and a minimised fixture documents the actual bug rather than the accident that found it.

Every finding becomes a permanent test The fuzzer produces a crashing input, minimisation reduces it to the smallest trigger, and the result is committed as a fixture in the fast test suite so the bug cannot return unnoticed. fuzzer finds a crash 412 bytes of input cargo fuzz tmin 6 bytes, same crash fix and commit fixture + test fast suite runs forever The corpus is also worth committing: it is a record of the inputs that reached interesting code paths, and it makes the next run start warm.

Running it in CI without slowing everything down

Fuzzing rewards hours and CI budgets minutes, so the pipeline run is a smoke test rather than a search.

- name: fuzz (short)
  run: |
    cargo fuzz run parse_record -- -max_total_time=120 -runs=0
    cargo fuzz run parse_record_structured -- -max_total_time=120 -runs=0

Two minutes per target on every change catches regressions against the committed corpus and finds shallow new bugs in freshly changed code. A longer run — an hour or overnight, on a schedule — does the actual searching, and any findings become fixtures the fast suite then covers.

Commit the corpus. A warm corpus means the two-minute run starts from inputs that already reach deep code paths rather than rediscovering the file header every time.

Expected output

A clean run reports coverage growth and no crashes:

cargo fuzz run parse_record -- -max_total_time=120
INFO: Seed: 2847193
INFO: 412 files found in fuzz/corpus/parse_record
#1024   NEW    cov: 1847 ft: 3219 corp: 418/62Kb exec/s: 14021
#8192   REDUCE cov: 1902 ft: 3410 corp: 431/58Kb exec/s: 13884
#65536  pulse  cov: 1902 ft: 3417 corp: 433/58Kb exec/s: 13901
Done 1642880 runs in 121 second(s)

cov is the number of code edges reached; a run where it stops growing early suggests the fuzzer cannot get past a validation check, which is a hint to generate structured input instead.

The fuzzing loop A seed corpus feeds a harness that calls the module's parsing entry point. A crash is saved, minimised to the smallest input that still reproduces, and promoted into the regression suite. seed corpus real inputs harness one entry point crash saved input recorded minimised case added to tests Fuzz the native build, not the Wasm one: the bugs are in the same code and the native tooling is far faster. A harness that does more than one thing produces crashes you cannot attribute; one entry point per target. Every minimised reproducer belongs in the unit-test suite, or the same bug returns in a month.

Gotchas

  • Fuzzing a function that allocates unboundedly from input. Every run exhausts memory and the fuzzer reports out-of-memory rather than finding real bugs. Cap allocations from untrusted lengths.
  • Treating a returned error as a crash. Errors are correct behaviour for bad input; only panics and memory violations are findings.
  • No corpus committed. Every run starts cold and rediscovers the basics.
  • Fuzzing only the raw-bytes target. Plateaus quickly on structured formats.
  • Long fuzz runs blocking a merge. Keep the pipeline run short; do the searching on a schedule.
  • Unminimised fixtures. A 400-byte artifact documents nothing; minimise before committing.

Performance note

On a laptop, the native fuzz target ran at roughly 14,000 executions per second and found three panics in the first two minutes against a fresh corpus — all of them in length handling of truncated input. The equivalent search against the compiled module under a runtime managed about 300 executions per second, which would have taken over an hour to reach the same coverage. That ratio is the whole argument for fuzzing natively.

Frequently Asked Questions

Is fuzzing worth it if the module only processes my own data? Less so, but “my own data” has a way of becoming user data. If the input can ever come from outside — an upload, a paste, a third-party feed — fuzz it.

What about fuzzing the JavaScript glue? Different tooling and usually lower value: the glue is small and its inputs are yours. Property-based tests over the marshalling functions catch more for less effort.

Should I fuzz continuously? If the module parses untrusted input at scale, yes — a scheduled job on a spare machine, with findings filed automatically. For most projects a nightly hour plus the short CI run is proportionate.

← Back to Testing & Verifying Wasm Builds