Recording Wasm Call Traces for Bug Reports

This page answers one task: users report bugs in a WebAssembly-powered feature — “the export came out wrong”, “the editor crashed after a few edits” — and you cannot reproduce them, because the sequence of inputs that triggered the problem is unknown. You want an opt-in recorder that captures the calls into the module and lets you replay them locally, deterministically, under a debugger.

Prerequisites

  • [ ] A JavaScript wrapper through which all calls into the module go.
  • [ ] A module whose behaviour is deterministic for the same sequence of inputs (or can be made so).
  • [ ] A bug-report flow where users can attach a file, with explicit consent.

Why record-and-replay works well for Wasm

A WebAssembly module is a deterministic state machine as long as its inputs are: the same sequence of calls with the same arguments, starting from the same initial state, produces the same results and the same internal state. Non-determinism enters only through imports — the clock, random numbers, data fetched from JavaScript — and through the order of calls. If the wrapper records every call’s arguments and every value returned by non-deterministic imports, the whole session can be replayed in a test harness, with the same module build, and the bug reappears exactly. That is far more effective than asking users for steps to reproduce.

The cost is data: recordings can contain user content (documents, images), so they must be opt-in, scoped, and reviewed by the user before sending.

Record, report and replay With the user's consent, the wrapper records each call's name, arguments and results, plus values returned by non-deterministic imports, into a ring buffer. When the user reports a bug, the trace is packaged with the module version and shown for review before upload. Developers replay the trace against the same module build in a harness and reproduce the bug under a debugger. user enables recording explicit consent wrapper logs calls args, results, imports ring buffer last N calls bug report + trace user reviews first replay in harness same module build

Step 1 — wrap calls with a recorder

const trace = { module: MODULE_HASH, startedAt: Date.now(), calls: [] };
let recording = false;
const MAX_CALLS = 5000;

function recorded(name, fn) {
  return (...args) => {
    if (!recording) return fn(...args);
    const entry = { n: name, a: args.map(serialiseArg), t: performance.now() };
    try {
      const r = fn(...args);
      entry.r = summarise(r);                          // hash or small value, not full outputs
      return r;
    } catch (e) {
      entry.e = String(e?.message ?? e);
      throw e;
    } finally {
      trace.calls.push(entry);
      if (trace.calls.length > MAX_CALLS) trace.calls.shift();   // ring buffer
    }
  };
}

serialiseArg converts typed arrays to base64 (or stores them by content hash with the bytes kept separately), plain objects to JSON, and handles to stable IDs. Results are summarised — a hash of the output — so the replay can detect divergence without storing everything.

Step 2 — record non-deterministic imports

Find every place the module receives values from outside its arguments: imported clock functions, random sources, callbacks returning data. Wrap those imports so their return values are logged in order:

const imports = {
  env: {
    now_ms: recordImport("now_ms", () => Date.now()),
    random_u32: recordImport("random_u32", () => crypto.getRandomValues(new Uint32Array(1))[0]),
  },
};
function recordImport(name, fn) {
  return (...args) => { const v = fn(...args); if (recording) trace.calls.push({ imp: name, v }); return v; };
}

During replay, the same imports return logged values in sequence instead of calling the real functions. Modules built with wasm-bindgen often reach the clock or randomness through js-sys; route those through your own imported functions in builds that support recording, or seed randomness from a recorded seed.

Step 3 — redact and let users review

Before upload, show the user what the trace contains — number of calls, size, which features — and offer to exclude content: replace document bytes with synthetic data of the same shape, or drop large arguments while keeping their sizes. Some bugs only reproduce with the real content; let users choose, and make it clear when the trace includes their data. Store traces with access controls and delete them after the bug is resolved.

Full traces versus redacted traces Full traces include real arguments such as document bytes and reproduce almost any bug exactly but contain user content and need explicit consent and careful handling. Redacted traces replace content with shape-preserving placeholders, protecting privacy, but reproduce only bugs that do not depend on the specific content. full trace real arguments included reproduces nearly everything needs consent + retention rules hard bugs, with consent redacted trace content replaced by shapes safer to share misses content-specific bugs default

Step 4 — package with version information

A trace is only replayable against the exact module build that produced it. Include the module’s content hash, the app version, the browser and the module’s feature flags (SIMD, threads). Keep released module builds (with names and debug info) archived by hash, so a trace from three versions ago can still be replayed against its build.

Step 5 — build a replay harness

The harness loads the archived module build, instantiates it with replay imports, and applies the recorded calls in order:

import { readFile } from "node:fs/promises";
const trace = JSON.parse(await readFile(process.argv[2], "utf8"));
const bytes = await readFile(`builds/${trace.module}.wasm`);
let impIdx = 0;
const replayImports = { env: new Proxy({}, { get: (_, name) => () => nextImport(name) }) };
const inst = await instantiateWithGlue(bytes, replayImports);
for (const c of trace.calls.filter((c) => c.n)) {
  const r = callExport(inst, c.n, c.a.map(deserialiseArg));
  if (c.r && summarise(r) !== c.r) { console.error("divergence at", c.n); break; }
}

Run it in Node with the inspector, or in a browser page with DevTools and DWARF debugging, and step into the call where the bug appears. A divergence report — “result differs from the recording at call 1,422” — points straight at the first wrong step.

Limits

Record-and-replay breaks down when the module depends on things you did not record: threads (interleavings differ between runs), shared memory modified by JavaScript, or host state not passed through imports. For threaded modules, record at a coarser level (whole jobs) or force single-threaded mode while recording. Very long sessions produce large traces; the ring buffer keeps the most recent calls, which is usually where the bug is.

Checkpoints for ring-buffer traces

A ring buffer keeps only the most recent calls, but replay must start from the state the module was in when the oldest kept call ran — not from a fresh instance. Two techniques solve this. The module can expose a checkpoint function that serialises its logical state (the open document, settings, caches that affect results) into bytes; the recorder takes a checkpoint periodically, say every 1,000 calls, and keeps the latest checkpoint plus every call after it. Replay restores the checkpoint, then applies the calls. Alternatively, snapshot linear memory itself: copying memory.buffer and the values of mutable globals captures the complete state for modules without external references, and restoring means instantiating the same build and writing the bytes back. Memory snapshots are simple and exact but large (the whole memory) and fragile across builds; logical checkpoints are small and survive refactors but must be maintained alongside the module. For most applications a logical checkpoint of “the document as saved” plus the call log since then is enough.

Using traces beyond single bugs

Once recording exists, traces become a source of realistic test data. With consent, collected redacted traces show which call sequences users actually produce; replaying them against new builds before release catches regressions that synthetic tests miss, and comparing timing across builds turns them into performance benchmarks grounded in real usage. Keep a curated library of traces in the repository, each tied to a fixed bug or an important workflow, and replay them in CI on every change to the module.

Expected output

A user enables “Record diagnostics”, reproduces the crash, and attaches a 2.4 MB redacted trace of the last 5,000 calls to a bug report after reviewing a summary; a developer replays it against the archived module build in Node, sees the result diverge at call 3,118, and steps into the failing function with the debugger.

Gotchas

  • Recording without consent. Traces can contain user data. Make it opt-in and reviewable.
  • Unrecorded non-determinism. Replays diverge. Record clock, randomness and callback values.
  • Missing module builds. Traces cannot be replayed. Archive builds by hash.
  • Unbounded traces. Memory and upload size grow. Use a ring buffer.
  • Threads during recording. Interleavings differ. Record single-threaded or at job level.
  • Ring buffers without checkpoints. Replay starts from the wrong state. Store periodic state checkpoints.

Performance note

Recording added about 3 µs per call for small arguments and the cost of hashing for large ones; for an editor making about 200 calls per second, overhead was under 1% of main-thread time.

Overhead of recording per call Microseconds added per call by the recorder for calls with small arguments and for calls whose 1 MB arguments are hashed and stored by content hash. µs per call small arguments 3 µs 1 MB argument (hashed) 450 µs

Frequently Asked Questions

Can this record DOM interactions too? Session-replay tools record the UI; this records the module’s inputs, which is what reproduces Wasm bugs.

Does replay need a browser? Not if the module does not depend on browser APIs; Node is convenient for automated replays.

Can traces become regression tests? Yes — once fixed, keep the trace and assert the new results in CI.

What about server-side Wasm? The same wrapper pattern works in hosts; record per request with sampling.

How can a ring-buffer trace be replayed if early calls were dropped? Store a periodic checkpoint of the module’s state and replay calls from the latest checkpoint onward.

Should I snapshot linear memory or a logical checkpoint? Logical checkpoints are small and survive refactors; memory snapshots are exact but large and tied to one build.

How long should traces be kept? Only until the bug is resolved, under access controls; curated, redacted traces may live on as regression tests with consent.

← Back to Observability & Error Reporting