Recording Wasm Call Traces for Bug Reports
This page answers one task: users report bugs in a WebAssembly-powered feature — “the export came out wrong”, “the editor crashed after a few edits” — and you cannot reproduce them, because the sequence of inputs that triggered the problem is unknown. You want an opt-in recorder that captures the calls into the module and lets you replay them locally, deterministically, under a debugger.
Prerequisites
- [ ] A JavaScript wrapper through which all calls into the module go.
- [ ] A module whose behaviour is deterministic for the same sequence of inputs (or can be made so).
- [ ] A bug-report flow where users can attach a file, with explicit consent.
Why record-and-replay works well for Wasm
A WebAssembly module is a deterministic state machine as long as its inputs are: the same sequence of calls with the same arguments, starting from the same initial state, produces the same results and the same internal state. Non-determinism enters only through imports — the clock, random numbers, data fetched from JavaScript — and through the order of calls. If the wrapper records every call’s arguments and every value returned by non-deterministic imports, the whole session can be replayed in a test harness, with the same module build, and the bug reappears exactly. That is far more effective than asking users for steps to reproduce.
The cost is data: recordings can contain user content (documents, images), so they must be opt-in, scoped, and reviewed by the user before sending.
Step 1 — wrap calls with a recorder
const trace = { module: MODULE_HASH, startedAt: Date.now(), calls: [] };
let recording = false;
const MAX_CALLS = 5000;
function recorded(name, fn) {
return (...args) => {
if (!recording) return fn(...args);
const entry = { n: name, a: args.map(serialiseArg), t: performance.now() };
try {
const r = fn(...args);
entry.r = summarise(r); // hash or small value, not full outputs
return r;
} catch (e) {
entry.e = String(e?.message ?? e);
throw e;
} finally {
trace.calls.push(entry);
if (trace.calls.length > MAX_CALLS) trace.calls.shift(); // ring buffer
}
};
}
serialiseArg converts typed arrays to base64 (or stores them by content hash with the bytes kept separately), plain objects to JSON, and handles to stable
IDs. Results are summarised — a hash of the output — so the replay can detect divergence without storing everything.
Step 2 — record non-deterministic imports
Find every place the module receives values from outside its arguments: imported clock functions, random sources, callbacks returning data. Wrap those imports so their return values are logged in order:
const imports = {
env: {
now_ms: recordImport("now_ms", () => Date.now()),
random_u32: recordImport("random_u32", () => crypto.getRandomValues(new Uint32Array(1))[0]),
},
};
function recordImport(name, fn) {
return (...args) => { const v = fn(...args); if (recording) trace.calls.push({ imp: name, v }); return v; };
}
During replay, the same imports return logged values in sequence instead of calling the real functions. Modules built with wasm-bindgen often reach the clock or
randomness through js-sys; route those through your own imported functions in builds that support recording, or seed randomness from a recorded seed.
Step 3 — redact and let users review
Before upload, show the user what the trace contains — number of calls, size, which features — and offer to exclude content: replace document bytes with synthetic data of the same shape, or drop large arguments while keeping their sizes. Some bugs only reproduce with the real content; let users choose, and make it clear when the trace includes their data. Store traces with access controls and delete them after the bug is resolved.
Step 4 — package with version information
A trace is only replayable against the exact module build that produced it. Include the module’s content hash, the app version, the browser and the module’s feature flags (SIMD, threads). Keep released module builds (with names and debug info) archived by hash, so a trace from three versions ago can still be replayed against its build.
Step 5 — build a replay harness
The harness loads the archived module build, instantiates it with replay imports, and applies the recorded calls in order:
import { readFile } from "node:fs/promises";
const trace = JSON.parse(await readFile(process.argv[2], "utf8"));
const bytes = await readFile(`builds/${trace.module}.wasm`);
let impIdx = 0;
const replayImports = { env: new Proxy({}, { get: (_, name) => () => nextImport(name) }) };
const inst = await instantiateWithGlue(bytes, replayImports);
for (const c of trace.calls.filter((c) => c.n)) {
const r = callExport(inst, c.n, c.a.map(deserialiseArg));
if (c.r && summarise(r) !== c.r) { console.error("divergence at", c.n); break; }
}
Run it in Node with the inspector, or in a browser page with DevTools and DWARF debugging, and step into the call where the bug appears. A divergence report — “result differs from the recording at call 1,422” — points straight at the first wrong step.
Limits
Record-and-replay breaks down when the module depends on things you did not record: threads (interleavings differ between runs), shared memory modified by JavaScript, or host state not passed through imports. For threaded modules, record at a coarser level (whole jobs) or force single-threaded mode while recording. Very long sessions produce large traces; the ring buffer keeps the most recent calls, which is usually where the bug is.
Checkpoints for ring-buffer traces
A ring buffer keeps only the most recent calls, but replay must start from the state the module was in when the oldest kept call ran — not from a fresh
instance. Two techniques solve this. The module can expose a checkpoint function that serialises its logical state (the open document, settings, caches that
affect results) into bytes; the recorder takes a checkpoint periodically, say every 1,000 calls, and keeps the latest checkpoint plus every call after it. Replay
restores the checkpoint, then applies the calls. Alternatively, snapshot linear memory itself: copying memory.buffer and the values of mutable globals
captures the complete state for modules without external references, and restoring means instantiating the same build and writing the bytes back. Memory
snapshots are simple and exact but large (the whole memory) and fragile across builds; logical checkpoints are small and survive refactors but must be
maintained alongside the module. For most applications a logical checkpoint of “the document as saved” plus the call log since then is enough.
Using traces beyond single bugs
Once recording exists, traces become a source of realistic test data. With consent, collected redacted traces show which call sequences users actually produce; replaying them against new builds before release catches regressions that synthetic tests miss, and comparing timing across builds turns them into performance benchmarks grounded in real usage. Keep a curated library of traces in the repository, each tied to a fixed bug or an important workflow, and replay them in CI on every change to the module.
Expected output
A user enables “Record diagnostics”, reproduces the crash, and attaches a 2.4 MB redacted trace of the last 5,000 calls to a bug report after reviewing a summary; a developer replays it against the archived module build in Node, sees the result diverge at call 3,118, and steps into the failing function with the debugger.
Gotchas
- Recording without consent. Traces can contain user data. Make it opt-in and reviewable.
- Unrecorded non-determinism. Replays diverge. Record clock, randomness and callback values.
- Missing module builds. Traces cannot be replayed. Archive builds by hash.
- Unbounded traces. Memory and upload size grow. Use a ring buffer.
- Threads during recording. Interleavings differ. Record single-threaded or at job level.
- Ring buffers without checkpoints. Replay starts from the wrong state. Store periodic state checkpoints.
Performance note
Recording added about 3 µs per call for small arguments and the cost of hashing for large ones; for an editor making about 200 calls per second, overhead was under 1% of main-thread time.
Frequently Asked Questions
Can this record DOM interactions too? Session-replay tools record the UI; this records the module’s inputs, which is what reproduces Wasm bugs.
Does replay need a browser? Not if the module does not depend on browser APIs; Node is convenient for automated replays.
Can traces become regression tests? Yes — once fixed, keep the trace and assert the new results in CI.
What about server-side Wasm? The same wrapper pattern works in hosts; record per request with sampling.
How can a ring-buffer trace be replayed if early calls were dropped? Store a periodic checkpoint of the module’s state and replay calls from the latest checkpoint onward.
Should I snapshot linear memory or a logical checkpoint? Logical checkpoints are small and survive refactors; memory snapshots are exact but large and tied to one build.
How long should traces be kept? Only until the bug is resolved, under access controls; curated, redacted traces may live on as regression tests with consent.
Related
- Adding feature usage telemetry to Wasm modules — lightweight counts.
- Reporting Wasm crashes to an error tracker — crash reports.
- Symbolicating Wasm stack traces in production — archived symbols.
- Snapshot testing Wasm output — traces as tests.
← Back to Observability & Error Reporting