Recovering a Module After a Trap
This page answers one task: a WebAssembly call trapped — a panic, an out-of-bounds access, unreachable — and the application should keep working
afterwards, without the next call failing mysteriously or returning wrong results.
Prerequisites
- [ ] A module whose traps you already catch, as in catching Wasm traps in JavaScript.
- [ ] Control of how the module is loaded, so it can be instantiated again on demand.
A trap stops the call, not the damage
When a WebAssembly instruction traps, the engine abandons the current call immediately and throws a WebAssembly.RuntimeError into JavaScript. No
code in the module runs after the trapping instruction: no destructors, no finally blocks in the source language, no unlocking of mutexes, no
restoring of invariants. The instance itself survives — its exports can be called again — and its linear memory, globals and tables are exactly as
they were at the moment of the trap.
That is the problem. A trap in the middle of an allocator call can leave a free list half-linked. A trap while a Rust RefCell is borrowed leaves it
borrowed forever, so the next call panics with “already borrowed”. The stack pointer global may not be restored, so later calls run with a shifted
stack. A static mut counter may be half-incremented. None of this is visible until a later call fails in a confusing way — or, worse, succeeds with
corrupted data.
Step 1 — decide that a trapped instance is disposable
The robust rule is simple: after any trap, stop using that instance. Recovering in place would require knowing exactly what the trapping function had touched, which is precisely what a trap prevents you from knowing. Instances are cheap to create when the compiled module is cached, so treating them as disposable costs a few milliseconds at worst.
There are narrow exceptions — a pure function with no global state, written so it cannot leave anything half-done — but rules like “this export is safe after a trap” are hard to keep true as code evolves. Make re-instantiation the default and opt out only with evidence.
Step 2 — keep the compiled module, not just the instance
Recovery is fast only if compilation is not repeated. Hold on to the WebAssembly.Module and create instances from it:
let compiled; // WebAssembly.Module, compiled once
let instance;
async function load() {
compiled ??= await WebAssembly.compileStreaming(fetch(new URL("./engine.wasm", import.meta.url)));
instance = await WebAssembly.instantiate(compiled, makeImports());
return instance;
}
function current() { return instance ?? load(); }
WebAssembly.instantiate with a Module skips compilation entirely; for a multi-megabyte module it takes single-digit milliseconds instead of
hundreds. Glue generators support the same: wasm-bindgen’s initSync({ module }) and Emscripten’s instantiateWasm hook both accept a precompiled
module, as described in
instantiating one module many times.
Step 3 — catch, discard and re-create in the wrapper
Wrap every export call so a trap discards the instance and the error still reaches the caller:
async function call(name, ...args) {
const inst = await current();
try {
return inst.exports[name](...args);
} catch (err) {
if (err instanceof WebAssembly.RuntimeError) {
instance = undefined; // next call gets a fresh instance
reportTrap(name, err); // the trap is still a bug: record it
}
throw err;
}
}
Note that the failing call itself is not retried automatically. A trap usually means a bug or an input the module cannot handle; retrying the same input on a fresh instance will most likely trap again. Let the caller decide — show an error, skip that item — while later calls, with other inputs, proceed normally.
Step 4 — restore the state the application needs
A fresh instance starts empty. If the old one held state the application depends on — a loaded document, a configured parser, a warm cache — that state has to be rebuilt. Two approaches work. Keep the authoritative state in JavaScript and replay it into each new instance: configuration objects, the current document bytes. Or snapshot the state at known-good points — serialise it out of the module after each successful operation — and restore the last snapshot after a re-instantiation. Either way, design the state so it can be reconstructed, because the trapped instance’s copy cannot be trusted.
Step 5 — handle workers and threads
If the module runs in a worker, the same logic applies inside the worker, or the main thread can simply terminate the worker and start a new one. Threaded modules with shared memory are harder: a trap in one thread leaves the shared memory inconsistent for all threads, and other threads may be blocked on a lock the trapping thread held. For threaded modules, recovery means tearing down the whole pool and its memory, then starting over.
Panics in Rust: abort, not unwind
Rust compiled to wasm32-unknown-unknown uses panic = "abort" semantics by default: a panic prints a message through the panic hook and executes
unreachable, which traps. Destructors along the stack do not run, so locks and borrow flags stay held, and the “already borrowed” or “already mutably
borrowed” panic on the next call is the classic symptom of an instance that should have been discarded. Installing
console_error_panic_hook makes the original panic message visible, which is essential for finding the real bug — the second panic is only a
consequence. Experimental unwinding support with Wasm exception handling changes some of this, but even with unwinding, state touched mid-operation can
be inconsistent, so the disposable-instance rule still applies.
Making the module easier to recover
Some design choices in the module make recovery cheaper and safer. Keep long-lived state small and explicit — a configuration struct and a few
handles — rather than spread across many globals, so it is easy to serialise and replay. Avoid lazily initialised statics that do expensive work on
first use; after re-instantiation, that work repeats on the first call and shows up as a latency spike long after the trap. Prefer per-operation
allocations that are freed when the operation finishes, so a fresh instance needs no warm-up to reach a steady state. And expose a cheap health()
export that exercises the allocator and returns a known value: calling it immediately after re-instantiation confirms the new instance is good before
real work is routed to it. None of this prevents traps, but together they turn recovery from a special case into the same code path as startup — which
is the path that gets the most testing.
Expected output
After a deliberately trapping call, the wrapper logs one trap report, the caller receives the RuntimeError, and the next call succeeds on a new instance
in under 5 ms of extra latency. A test that alternates trapping and valid inputs a thousand times shows no growth in failures and no “already borrowed”
panics.
Gotchas
- Continuing to use the trapped instance. Later calls fail with unrelated-looking errors. Discard it.
- Recompiling on every recovery. Cache the
WebAssembly.Module; instantiate from it. - Auto-retrying the failing input. It will usually trap again. Report and move on.
- Forgetting state lived in the old instance. Replay configuration and data into the new one.
- Recovering in a loop. If every new instance traps on its first call, the problem is the replayed state. Cap re-instantiations per minute.
- Memory held by the old instance. Drop every reference — exports, memory views, closures — so it can be garbage-collected.
Performance note
For a 3.1 MB module, compiling from bytes took 240 ms on a mid-range laptop; instantiating from the cached Module took 3.8 ms. Recovery after a trap is
therefore close to free when the compiled module is kept, and a visible stall when it is not.
Frequently Asked Questions
Is the memory of a trapped instance corrupted? Not in the sense of random bytes — it is exactly what the program had written. But the program’s data structures may be inconsistent.
Can I reset an instance without re-instantiating? Only by restoring a snapshot of all its memory and globals, which is more work and more error-prone than creating a new instance.
Does a trap in one instance affect others? No, unless they share a memory. Separate instances have separate memory and globals.
Should I recover from stack overflow traps? Yes, the same way. Overflow traps leave state just as inconsistent; see fixing stack overflow in Wasm.
How do I know which state to replay? List what the application has told the module since startup — configuration, loaded documents, registered callbacks — and keep that list in JavaScript. Anything not on the list is derived state the new instance can rebuild itself.
Related
- Debugging unreachable-executed traps — finding the underlying bug.
- Reporting Wasm crashes to an error tracker — making the trap visible in production.
- Handling out-of-memory in Wasm — a trap you can sometimes avoid.
- Cancelling long-running Wasm work — another reason to replace an instance.
← Back to Errors & Traps Across the Boundary