Tracking Down Memory Corruption in Wasm

This page answers one task: a value in a WebAssembly module changes when nothing should have written it — a struct field turns into garbage, a string grows a strange suffix, results go wrong after hours of uptime — and you need to find which code is writing where it should not.

Prerequisites

  • [ ] A module built from C, C++, or Rust with unsafe code, where the corruption occurs.
  • [ ] Some reproducibility: an input or sequence of actions that triggers it, even occasionally.
  • [ ] Access to the JavaScript host, so you can inspect linear memory from outside.

Why corruption hides in WebAssembly

Linear memory is one flat array of bytes. The module’s static data, its shadow stack and its heap all live in it, side by side, and nothing inside the module enforces boundaries between them. A write through a dangling pointer, one element past the end of an array, or from a stack frame that has overflowed lands on whatever happens to be at that address — another object’s field, the allocator’s bookkeeping, a global. The WebAssembly sandbox protects the host from this; it does nothing to protect the module from itself.

Natively, many of these bugs crash immediately because memory pages around objects are unmapped or protected. In Wasm there are no such pages inside linear memory: every address from zero to the current size is readable and writable. The program keeps running with corrupted state, and the symptom appears later and elsewhere — which is what makes these bugs expensive. The approach that works is to narrow down when the corruption happens and what overwrote it, then catch the write in the act.

Narrowing a corruption bug down Start from the symptom, find the corrupted address, detect the moment it changes with a checksum or watch, identify the code running at that moment, and confirm with a sanitizer build or a guard value. symptom wrong value, odd crash which address? locate the field when does it change? checksums between calls who wrote it? watch + stack trace confirm + fix sanitizer, guard values

Step 1 — find the corrupted address

Start by locating the data that goes wrong. If the module exposes a pointer to the affected object — an exported getter, a handle the JavaScript side holds — read the bytes directly. Otherwise, export a small debug function that returns the address of the structure in question:

// debug build only
__attribute__((export_name("debug_addr_of_config")))
const void *debug_addr_of_config(void) { return &g_config; }
const addr = instance.exports.debug_addr_of_config();
const view = () => new Uint8Array(instance.exports.memory.buffer, addr, 64);
console.log("config bytes:", [...view()].map((b) => b.toString(16).padStart(2, "0")).join(" "));

Knowing the address lets you compare it with the memory layout: is it static data near the top of the stack region, a heap block, or something the allocator manages? That alone often suggests the culprit, as the layout in setting stack size and memory limits at link time shows — static data adjacent to a downward-growing stack is the classic victim of stack overflow.

Step 2 — find when it changes

Take a checksum of the region before and after every call into the module. The first call after which it differs is the call that corrupts it:

function hash(bytes) { let h = 2166136261; for (const b of bytes) h = Math.imul(h ^ b, 16777619); return h >>> 0; }

function guarded(name, fn) {
  return (...args) => {
    const before = hash(view());
    const result = fn(...args);
    const after = hash(view());
    if (before !== after) console.error(`config changed during ${name}`, args, new Error().stack);
    return result;
  };
}

for (const [name, fn] of Object.entries(instance.exports)) {
  if (typeof fn === "function" && !name.startsWith("debug_")) exports[name] = guarded(name, fn);
}

Wrapping every export is crude and effective. If the region is legitimately written by some calls, exclude them, or watch a narrower range. For corruption that appears only after many calls, log the call sequence and replay it in a test.

Step 3 — find the write inside the call

Once you know which call corrupts the data, find the instruction. In a debug build, set a breakpoint at the start of that export and step, watching the bytes — DevTools’ memory inspector can highlight a range and shows changes as you step, as in inspecting Wasm memory in Chrome DevTools. For long calls, bisect inside them instead: add a debug export that checksums the region, call it from the module at intermediate points, and narrow down between which two points the change happens.

// sprinkle in the suspect code path while hunting (debug build)
extern void debug_check(const char *where);   // imported from JS: compares the checksum
void process_batch(Batch *b) {
  debug_check("start");
  decode_headers(b);
  debug_check("after headers");
  copy_payloads(b);
  debug_check("after payloads");     // ← if this is the first to report a change, look in copy_payloads
}

Step 4 — suspect the usual culprits

A handful of causes account for nearly all corruption in Wasm modules, and checking them first saves time.

Stack overflow into static data. Deep recursion or large local arrays push the shadow stack past its region. The victim is static data just below the stack. Fix by raising the stack size and adding --stack-first, which turns future overflows into traps.

Heap buffer overflow. An off-by-one or a wrong length when copying into a heap block writes into the next block’s header or data. The victim is often allocator metadata, so the crash comes later, inside malloc or free.

Use after free. A pointer kept after free, used once the block has been reallocated to something else. The victim changes depending on allocation patterns, so the bug moves when you add logging.

Stale JavaScript views. JavaScript writes through a typed array created before memory.grow, or keeps a pointer to Wasm memory the module has since freed. The module’s code is innocent; the glue is the writer.

Corruption causes and their signatures The four most common causes of memory corruption in Wasm modules, what typically gets overwritten, how the symptom behaves, and the fix. cause victim behaviour fix stack overflow static data deep inputs only bigger stack, --stack-first heap overflow next block or header crash later in malloc/free bounds-check lengths use after free whatever reuses it moves when code changes ownership fix, ASan stale JS view any Wasm data after memory growth recreate views

Step 5 — confirm with a sanitizer build when possible

Once you have a suspect, confirm it. For C and C++, an AddressSanitizer build catches heap overflows and use-after-free at the faulting instruction with a full report, as described in catching memory bugs with Emscripten sanitizers. For Rust, corruption almost always comes from unsafe blocks or from JavaScript writing into module memory; Miri on the native build checks the unsafe code. Where a sanitizer is impractical — a large application, a performance-sensitive repro — guard values work: place a known pattern after a suspect buffer and check it after each operation, which catches overflows at the operation that caused them.

Making corruption cheaper to find next time

The investigation above is expensive; the habits that make the next one cheaper are not. Build debug configurations with stack overflow checks enabled. Keep a sanitizer build in CI running the test suite, so heap overflows are caught when introduced rather than months later. Validate every length that crosses the JavaScript boundary inside the module, because a wrong length from the caller is a common way for a correct module to overflow its buffers. And keep JavaScript-side views short-lived — create them for one operation and discard them — so they cannot outlive a memory growth. None of these eliminates corruption, but together they move most of it from “mysterious production bug” to “test failure with a stack trace”.

Expected output

The guarded wrapper reports the first call that changes the region:

config changed during process_batch [ 1048576, 4096 ]
Error
    at Object.process_batch (debug.js:14:42)
    at handleUpload (app.js:88:20)

and the in-module checks narrow it to one function:

debug_check: unchanged at "start"
debug_check: unchanged at "after headers"
debug_check: CHANGED at "after payloads"

Gotchas

  • The bug disappears when you add logging. Logging changes allocation patterns or stack usage. That points at use-after-free or stack overflow; prefer checksums and guard values, which disturb less.
  • The corrupted value changes but nothing writes it. JavaScript writes through a view. Wrap JavaScript-side writes in the same checks.
  • A debug build cannot reproduce it. Optimization changes layout and inlining. Reproduce with an optimized build that keeps names and add the checks there.
  • Checksums on a moving region. If the object is reallocated, its address changes. Track the object, not a fixed address.

Performance note

The guarded wrapper with a 64-byte checksum added about 0.2 µs per call — fine for hunting, not for production. A full AddressSanitizer build ran the reproduction 2.8 times slower and found the culprit, an off-by-one in a payload copy, on the first run.

Cost of each detection technique on the reproduction run Run time of the corruption reproduction with no checks, with export-level checksums, with in-module checkpoints, and as an AddressSanitizer build. seconds to reproduce no checks 4.1 s checksums around exports 4.6 s + in-module checkpoints 5.3 s AddressSanitizer build 11.5 s

Frequently Asked Questions

Can a debugger break when an address is written? Native debuggers have hardware watchpoints; browser debuggers do not yet offer watchpoints on Wasm memory. Checksums and stepping with the memory inspector are the substitutes.

Is Rust immune to this? Safe Rust is. unsafe code, FFI to C, and JavaScript writing into module memory are not, and those are where Rust corruption bugs come from.

Could the engine be at fault? It is possible but rare. Reproduce in a second engine before suspecting it; if both show the corruption, the module is almost certainly responsible.

Does Memory64 change anything? The causes are the same. Larger memories make stray writes less likely to hit something important, which can make bugs rarer and harder to reproduce.

← Back to Debugging & Profiling Wasm Modules