Snapshot Testing Wasm Output

This page answers one task: test a WebAssembly module whose output is too large or complex to assert on by hand — a decoded image, a rendered SVG, a generated report, a compressed file — by comparing it with a stored reference, and do it in a way that tolerates harmless differences and catches real ones.

Prerequisites

  • [ ] A module that produces deterministic output from fixed input: an encoder, decoder, renderer or transformer.
  • [ ] A test runner: Vitest, Node’s built-in node:test, or cargo test with the insta crate on the Rust side.
  • [ ] A directory in the repository for fixtures and snapshots.

Why snapshots suit Wasm workloads

The workloads people move into WebAssembly are often the ones with big, structured output: image and video codecs, document renderers, compilers, data transformers. Writing assertions for every pixel or byte of that output is impossible, and asserting on a handful of properties — width, height, the first pixel — misses most regressions. A snapshot test takes the opposite approach: run the module on a fixed input, store the full output as a reference file, and on every later run compare the new output with the reference. Any difference is either a bug or an intended change, and a human decides which.

Snapshots are especially valuable around compiler and toolchain upgrades. A new wasm-opt or LLVM release can change floating-point evaluation order, SIMD lowering, or the handling of edge cases. Unit tests written against a few known values may still pass; a snapshot of a full decoded image will not, and that is the signal you want.

The snapshot test cycle A fixed input is run through the module. The output is compared with the stored reference using an exact or tolerance-based comparison. A match passes; a mismatch fails and writes the new output and a diff for review. An intended change is accepted by updating the reference. fixed input fixtures/photo.jpg module output decoded RGBA compare exact or tolerance mismatch write .new + diff review accept or fix

Step 1 — snapshot text output with the runner’s built-in support

For textual output — SVG, JSON, CSV, a disassembly — Vitest’s snapshot support is enough. toMatchFileSnapshot stores each snapshot as a real file, which keeps diffs readable in code review:

import { it, expect } from "vitest";
import { render_svg } from "../pkg/chart.js";

it("renders the reference bar chart", async () => {
  const data = JSON.parse(await readFixture("bars.json"));
  const svg = render_svg(JSON.stringify(data), 640, 360);
  await expect(svg).toMatchFileSnapshot("./__snapshots__/bars.svg");
});

The first run writes the file; later runs compare against it. vitest -u updates snapshots after an intended change. Because the snapshot is an SVG, reviewers can open it and look — which is far more useful than reading a diff of coordinate numbers.

Step 2 — compare binary output by hash, then by content

For binary output that should be byte-exact — a compressed file, an encoded image with a deterministic encoder — store a hash and the bytes, and compare the hash first:

import { createHash } from "node:crypto";
import { readFile, writeFile } from "node:fs/promises";
import { compress } from "../pkg/zstd.js";

it("compresses the corpus file byte-identically", async () => {
  const input = await readFile("test/fixtures/corpus.txt");
  const out = compress(input, 19);
  const hash = createHash("sha256").update(out).digest("hex");
  const expected = (await readFile("test/__snapshots__/corpus.zst.sha256", "utf8")).trim();
  if (hash !== expected) await writeFile("test/__snapshots__/corpus.zst.new", out);
  expect(hash).toBe(expected);
});

Storing only the hash keeps the repository small; writing the actual output on failure gives you something to inspect. For decompressors and decoders, round-trip tests are a powerful complement — compress then decompress must return the input exactly, which needs no snapshot at all.

Step 3 — compare images with a tolerance

Decoded images, rendered frames and anything involving floating point should not be compared byte for byte. Different browsers, a different SIMD path, or a toolchain upgrade can legitimately change a few values by one. Compare per pixel with a threshold, and fail on the number of pixels that differ beyond it:

function diffRgba(a, b, { channelTolerance = 2, maxDiffPixels = 0 } = {}) {
  if (a.length !== b.length) return { ok: false, reason: "size differs" };
  let bad = 0, worst = 0;
  for (let i = 0; i < a.length; i += 4) {
    const d = Math.max(Math.abs(a[i] - b[i]), Math.abs(a[i+1] - b[i+1]),
                       Math.abs(a[i+2] - b[i+2]), Math.abs(a[i+3] - b[i+3]));
    worst = Math.max(worst, d);
    if (d > channelTolerance) bad++;
  }
  return { ok: bad <= maxDiffPixels, bad, worst };
}

it("decodes the reference JPEG within tolerance", async () => {
  const rgba = decode_jpeg(await readFile("test/fixtures/photo.jpg"));
  const ref = await readPng("test/__snapshots__/photo.rgba.png");
  const r = diffRgba(rgba, ref, { channelTolerance: 2, maxDiffPixels: 50 });
  expect(r, `${r.bad} pixels differ, worst channel delta ${r.worst}`).toMatchObject({ ok: true });
});

Store reference images as PNG, which is lossless and compresses well, rather than raw RGBA. On failure, write a diff image that highlights the differing pixels in a bright colour — it turns a failing test into something a person can judge in seconds.

Choosing a comparison for each kind of output Text output is compared exactly against a file snapshot, deterministic binary output by hash, decoded images by per-pixel tolerance, and floating-point numeric output by relative error. output kind comparison store text (SVG, JSON, WAT) exact, readable diff file snapshot deterministic binary sha256, write output on fail hash file decoded images / frames per-pixel tolerance + count lossless PNG float arrays relative error ≤ 1e-6 binary + tolerance

Step 4 — make snapshots deterministic

A snapshot test is only useful if the output is a function of the input. Anything else — timestamps, random seeds, iteration order over hash maps, thread scheduling — produces failures that are noise. Pin all of it:

// Rust side: deterministic options for test builds
pub struct EncodeOptions {
    pub threads: usize,      // 1 in tests: output must not depend on work splitting
    pub seed: u64,           // fixed seed for any randomised heuristics
    pub embed_timestamp: bool,
}

Threads deserve a specific mention. A parallel encoder that splits work differently depending on the number of cores can produce valid but different output on a CI runner and a laptop. Run snapshot tests single-threaded, or make the work split independent of the core count.

Step 5 — review and update deliberately

The weakness of snapshot testing is that updating snapshots is easy, and an update can bless a regression. Treat a snapshot change like a code change: it shows up in the diff, it gets reviewed, and the reviewer looks at the new output. A few habits help:

  • Update snapshots in a separate commit from the code change that caused them, with a message saying why.
  • Keep snapshots small. A 64×64 test image catches the same decoder bugs as a 4K one and keeps the repository light.
  • Name snapshots after what they test — jpeg-progressive-cmyk.png, not test3.png.
  • Fail CI when snapshots are missing rather than writing them silently (vitest --ci does this).

Expected output

On a mismatch, a good snapshot test tells you what changed and leaves the evidence:

 FAIL  test/decode.test.js > decodes the reference JPEG within tolerance
AssertionError: 1312 pixels differ, worst channel delta 37
  ❯ test/decode.test.js:21:5
wrote test/__snapshots__/photo.rgba.new.png and photo.rgba.diff.png

A worst-case delta of 37 is a real bug — a wrong IDCT coefficient, say. A worst case of 1 across a few hundred pixels is almost always rounding, and the tolerance should absorb it.

Gotchas

  • Snapshots pass on one machine and fail on another. The output depends on something outside the input — thread count, locale, a SIMD path chosen at runtime. Pin it in test builds.
  • Huge repository diffs after every change. Snapshots are too large or stored in a format that diffs badly. Use small fixtures and PNG for images.
  • A snapshot update that hid a regression. The reviewer approved the code diff and not the snapshot. Make snapshot changes visible in review, with images rendered in the pull request where possible.
  • Tolerance so loose it catches nothing. Set it from data: measure the differences between two browsers on a correct build, and set the threshold just above that.

Performance note

The suite of 40 image snapshots at 128×128 ran in 0.9 s in Node, most of it PNG decoding of the references. Raising the fixtures to full HD made the suite 30 times slower without catching a single additional bug in a year of history, which is the argument for small fixtures.

Snapshot suite time by fixture size Forty decoder snapshot tests run in Node at three fixture resolutions. Time grows with pixel count; the bugs caught did not. seconds for 40 image snapshots 128 × 128 fixtures 0.9 s 512 × 512 fixtures 6.1 s 1920 × 1080 fixtures 27.4 s

Frequently Asked Questions

Should snapshots be generated in Node or in a browser? Generate and compare in the environment where the code will run in production, if output can differ. For pure computation the engines agree to within float rounding, so Node is fine and faster.

Can I snapshot test on the Rust side instead? Yes — the insta crate does snapshot testing in cargo test, natively. That tests the algorithm; a JavaScript-side snapshot additionally tests the boundary and the Wasm build.

How is this different from differential testing? A snapshot compares against a stored past result. A differential test compares two implementations running now — for example the Wasm build against the native build; see differential testing Wasm against native builds.

Where should fixtures come from? Real inputs that once caused trouble are the best fixtures: a file a user reported, a crash minimised by the fuzzer, an image with an unusual colour profile. Add one each time a bug is fixed, and the snapshot suite becomes a record of every edge case the module has met — which is exactly what a toolchain upgrade needs to be checked against.

What about visual regression of the whole page? That is a browser screenshot test of the UI, a different tool. Snapshot testing here is about the module’s own output.

← Back to Testing & Verifying Wasm Builds