Avoiding Data Races in Shared Wasm Memory
This page answers one task: a WebAssembly module shares memory between workers — or between a worker and JavaScript on the main thread — and you need to be sure that no two threads touch the same data in a way that produces torn values, lost updates or results that depend on timing.
Prerequisites
- [ ] A threaded module (Rayon, pthreads, or a hand-built pool) or JavaScript code reading shared Wasm memory.
- [ ] Familiarity with using Atomics for Wasm thread synchronization.
What a data race is, and why it hides
A data race happens when two threads access the same memory location at the same time, at least one of them writes, and the accesses are not
synchronised — neither atomic nor ordered by a lock or other synchronisation. The outcome is not merely “one of the two values”. A 64-bit value written
by one thread can be read half-old, half-new by another. An increment from two threads (load, add, store) can lose one of the updates. A flag set
“after” writing data can be observed before the data, because neither the compiler nor the CPU is obliged to keep unsynchronised writes in program
order.
The WebAssembly and JavaScript memory models give racy programs some defined behaviour — unlike C, a race does not make the whole program undefined — but the values observed are unpredictable. Worse, races are timing-dependent. On a developer’s laptop with a light workload, the problematic interleaving may never happen; under load on a user’s eight-core phone, it happens once in ten thousand frames. Testing finds races poorly. Design and tooling find them well.
Step 1 — classify every shared location
The most effective defence is knowing, for each piece of shared memory, which of three categories it belongs to:
- Immutable after publication. Written once before other threads can see it, then only read. Needs no atomics, only a correct hand-off.
- Owned by one thread at a time. Only the owner reads or writes; ownership passes through synchronisation (a lock, a queue). Plain accesses are fine.
- Concurrently accessed. Several threads read and write. Every access must be atomic, or protected by the same lock.
Most races come from data that was assumed to be in the first or second category but is not — a “read-only” table that is lazily filled in, a buffer “owned” by a worker that the main thread also reads for progress. Write the category down next to each shared structure.
Step 2 — publish with release, consume with acquire
The hand-off is where “immutable” data goes wrong. A producer fills a buffer, then sets a flag; a consumer sees the flag and reads the buffer. If the flag is a plain variable, the consumer may see the flag set but stale buffer contents. Make the flag atomic, and store it after the data:
use std::sync::atomic::{AtomicU32, Ordering::{Acquire, Release}};
static READY: AtomicU32 = AtomicU32::new(0);
static mut RESULT: [f32; 1024] = [0.0; 1024];
fn produce() {
unsafe { fill(&mut RESULT); }
READY.store(1, Release); // everything before is visible to an Acquire load that sees 1
}
fn consume() -> Option<&'static [f32; 1024]> {
if READY.load(Acquire) == 1 { Some(unsafe { &RESULT }) } else { None }
}
In JavaScript, every Atomics operation is sequentially consistent, which is stronger than acquire/release and gives the same guarantee: write the data,
then Atomics.store the flag; Atomics.load the flag, then read the data. WebAssembly’s atomic instructions are likewise sequentially consistent, so
Rust’s Release/Acquire compile to them.
Step 3 — make read-modify-write atomic
Counters, indices and accumulators updated by several threads must use atomic read-modify-write operations — fetch_add, compare_exchange — never a
load followed by a store:
static PROCESSED: AtomicU32 = AtomicU32::new(0);
PROCESSED.fetch_add(1, Relaxed); // correct: one indivisible operation
// let n = PROCESSED.load(Relaxed); PROCESSED.store(n + 1, Relaxed); // races: lost updates
Relaxed is fine for a statistics counter whose value is not used to order other memory accesses. When the counter guards other data — a job index that
decides which slot a worker reads — use the stronger orderings.
Step 4 — let the language prevent races
Rust’s type system rules out data races in safe code: shared mutable state must be wrapped in Mutex, RwLock or atomics, and Send/Sync bounds stop
non-thread-safe types from crossing threads. In a threaded Wasm build, keep unsafe and static mut out of shared paths — each one bypasses those
guarantees — and prefer channels, Rayon’s data-parallel iterators (which split data into disjoint chunks per thread) or Mutex-protected state. In C and
C++, use _Atomic/std::atomic for every concurrently accessed variable and follow a lock discipline. ThreadSanitizer is not available for Wasm targets,
so run the same C or Rust code natively with -fsanitize=thread (or RUSTFLAGS=-Zsanitizer=thread) as part of testing; most races are in portable logic
that reproduces natively.
Step 5 — treat JavaScript as another thread
JavaScript reading shared Wasm memory is a thread like any other. A main-thread progress bar reading a counter, a renderer reading a frame buffer a worker
is writing, an audio callback reading samples — all are concurrent accesses. Use Atomics.load for values that change while being read, and for buffers
use ownership transfer: double or triple buffering, where the worker writes one buffer while JavaScript reads another and the index of the readable
buffer is published atomically. Never read a multi-byte value with a plain typed-array access while a worker may be writing it; a Float64Array element
can be torn on some platforms.
Patterns that make races unlikely by design
The surest way to avoid races is to share less. Partition data so each thread works on its own slice and results are combined at the end — the map-reduce shape that Rayon encourages. Pass messages that move ownership of buffers between threads rather than letting several threads touch one buffer. Keep shared mutable state small and concentrated in a few well-reviewed structures — a queue, a set of counters, a result table — each with a documented synchronisation rule. Prefer immutable snapshots for configuration: build a new configuration object and publish a pointer to it atomically, rather than updating fields in place while workers read them. These designs make the category of every location obvious, which turns race-freedom from a property you hope for into one you can read off the code. When a design requires fine-grained shared mutation, isolate it behind a small API and test that API under stress — many threads, random delays, millions of operations — to give rare interleavings a chance to appear.
Expected output
A stress test running 8 workers for 10 million operations against the shared job queue and counters produces exact, repeatable totals; the native ThreadSanitizer run of the same logic reports no races; and the main thread’s progress display reads only atomically published values.
Gotchas
- Plain flag for hand-off. Consumers see the flag before the data. Use an atomic store after writing.
- Load-then-store increments. Updates are lost under contention. Use
fetch_addorAtomics.add. static mutin threaded Rust. Bypasses the type system’s race protection. Use atomics or locks.- JavaScript reading while a worker writes. Torn reads and stale data. Use atomics or buffer ownership.
- Trusting tests on a quiet machine. Races need load to appear. Stress-test and use native sanitizers.
Performance note
Uncontended atomic operations in Wasm cost a few nanoseconds; contended ones on a shared cache line can cost 50–100 ns. In a histogram benchmark, eight threads incrementing one shared atomic array ran slower than one thread; giving each thread a private histogram and summing at the end ran 6.8× faster than single-threaded.
Frequently Asked Questions
Are plain loads and stores of i32 atomic in Wasm?
Aligned accesses are not torn in practice, but they carry no ordering guarantees. Use atomic instructions for anything shared.
Does Relaxed ordering exist in WebAssembly?
Wasm atomics are sequentially consistent; Rust’s weaker orderings compile to the same instructions, so they are correct but not cheaper.
Can races corrupt the engine or escape the sandbox? No. Races produce unpredictable values within linear memory but never violate memory safety outside it.
How do I find a race I suspect exists? Reproduce the logic natively with ThreadSanitizer, or add assertions on invariants and run a stress test with many threads.
Is a Mutex in threaded Wasm expensive?
Uncontended, it is a couple of atomic operations. Contended, waiting workers sleep in memory.atomic.wait32, which is efficient but adds wake-up latency.
Related
- Implementing a lock-free ring buffer in shared memory — a race-free hand-off structure.
- Building a Wasm thread pool — a shared queue done right.
- Aligning data in linear memory — atomic alignment rules.
- Differential testing Wasm against native builds — running the same logic natively.
← Back to SharedArrayBuffer, Atomics & Threading