Avoiding Data Races in Shared Wasm Memory

This page answers one task: a WebAssembly module shares memory between workers — or between a worker and JavaScript on the main thread — and you need to be sure that no two threads touch the same data in a way that produces torn values, lost updates or results that depend on timing.

Prerequisites

What a data race is, and why it hides

A data race happens when two threads access the same memory location at the same time, at least one of them writes, and the accesses are not synchronised — neither atomic nor ordered by a lock or other synchronisation. The outcome is not merely “one of the two values”. A 64-bit value written by one thread can be read half-old, half-new by another. An increment from two threads (load, add, store) can lose one of the updates. A flag set “after” writing data can be observed before the data, because neither the compiler nor the CPU is obliged to keep unsynchronised writes in program order.

The WebAssembly and JavaScript memory models give racy programs some defined behaviour — unlike C, a race does not make the whole program undefined — but the values observed are unpredictable. Worse, races are timing-dependent. On a developer’s laptop with a light workload, the problematic interleaving may never happen; under load on a user’s eight-core phone, it happens once in ten thousand frames. Testing finds races poorly. Design and tooling find them well.

A lost update from an unsynchronised increment Two workers both read a counter value of 41, both add one, and both store 42. One increment is lost. With an atomic add, the read-modify-write is indivisible and the counter reaches 43. worker A counter in memory worker B load → 41 load → 41 store 42 store 42 (update lost)

Step 1 — classify every shared location

The most effective defence is knowing, for each piece of shared memory, which of three categories it belongs to:

  • Immutable after publication. Written once before other threads can see it, then only read. Needs no atomics, only a correct hand-off.
  • Owned by one thread at a time. Only the owner reads or writes; ownership passes through synchronisation (a lock, a queue). Plain accesses are fine.
  • Concurrently accessed. Several threads read and write. Every access must be atomic, or protected by the same lock.

Most races come from data that was assumed to be in the first or second category but is not — a “read-only” table that is lazily filled in, a buffer “owned” by a worker that the main thread also reads for progress. Write the category down next to each shared structure.

Shared-memory access rules by category Data that is immutable after publication needs only a synchronised hand-off. Data owned by one thread at a time needs ownership transfer through a lock or queue. Concurrently accessed data needs atomics or a lock on every access. category plain access allowed? synchronisation immutable after publish yes, after hand-off publish with release / atomic store owned by one thread by the owner only lock or queue transfers ownership concurrently accessed no atomics or one lock for all access

Step 2 — publish with release, consume with acquire

The hand-off is where “immutable” data goes wrong. A producer fills a buffer, then sets a flag; a consumer sees the flag and reads the buffer. If the flag is a plain variable, the consumer may see the flag set but stale buffer contents. Make the flag atomic, and store it after the data:

use std::sync::atomic::{AtomicU32, Ordering::{Acquire, Release}};

static READY: AtomicU32 = AtomicU32::new(0);
static mut RESULT: [f32; 1024] = [0.0; 1024];

fn produce() {
    unsafe { fill(&mut RESULT); }
    READY.store(1, Release);                 // everything before is visible to an Acquire load that sees 1
}

fn consume() -> Option<&'static [f32; 1024]> {
    if READY.load(Acquire) == 1 { Some(unsafe { &RESULT }) } else { None }
}

In JavaScript, every Atomics operation is sequentially consistent, which is stronger than acquire/release and gives the same guarantee: write the data, then Atomics.store the flag; Atomics.load the flag, then read the data. WebAssembly’s atomic instructions are likewise sequentially consistent, so Rust’s Release/Acquire compile to them.

Step 3 — make read-modify-write atomic

Counters, indices and accumulators updated by several threads must use atomic read-modify-write operations — fetch_add, compare_exchange — never a load followed by a store:

static PROCESSED: AtomicU32 = AtomicU32::new(0);
PROCESSED.fetch_add(1, Relaxed);             // correct: one indivisible operation
// let n = PROCESSED.load(Relaxed); PROCESSED.store(n + 1, Relaxed);   // races: lost updates

Relaxed is fine for a statistics counter whose value is not used to order other memory accesses. When the counter guards other data — a job index that decides which slot a worker reads — use the stronger orderings.

Step 4 — let the language prevent races

Rust’s type system rules out data races in safe code: shared mutable state must be wrapped in Mutex, RwLock or atomics, and Send/Sync bounds stop non-thread-safe types from crossing threads. In a threaded Wasm build, keep unsafe and static mut out of shared paths — each one bypasses those guarantees — and prefer channels, Rayon’s data-parallel iterators (which split data into disjoint chunks per thread) or Mutex-protected state. In C and C++, use _Atomic/std::atomic for every concurrently accessed variable and follow a lock discipline. ThreadSanitizer is not available for Wasm targets, so run the same C or Rust code natively with -fsanitize=thread (or RUSTFLAGS=-Zsanitizer=thread) as part of testing; most races are in portable logic that reproduces natively.

Step 5 — treat JavaScript as another thread

JavaScript reading shared Wasm memory is a thread like any other. A main-thread progress bar reading a counter, a renderer reading a frame buffer a worker is writing, an audio callback reading samples — all are concurrent accesses. Use Atomics.load for values that change while being read, and for buffers use ownership transfer: double or triple buffering, where the worker writes one buffer while JavaScript reads another and the index of the readable buffer is published atomically. Never read a multi-byte value with a plain typed-array access while a worker may be writing it; a Float64Array element can be torn on some platforms.

Patterns that make races unlikely by design

The surest way to avoid races is to share less. Partition data so each thread works on its own slice and results are combined at the end — the map-reduce shape that Rayon encourages. Pass messages that move ownership of buffers between threads rather than letting several threads touch one buffer. Keep shared mutable state small and concentrated in a few well-reviewed structures — a queue, a set of counters, a result table — each with a documented synchronisation rule. Prefer immutable snapshots for configuration: build a new configuration object and publish a pointer to it atomically, rather than updating fields in place while workers read them. These designs make the category of every location obvious, which turns race-freedom from a property you hope for into one you can read off the code. When a design requires fine-grained shared mutation, isolate it behind a small API and test that API under stress — many threads, random delays, millions of operations — to give rare interleavings a chance to appear.

Expected output

A stress test running 8 workers for 10 million operations against the shared job queue and counters produces exact, repeatable totals; the native ThreadSanitizer run of the same logic reports no races; and the main thread’s progress display reads only atomically published values.

Gotchas

  • Plain flag for hand-off. Consumers see the flag before the data. Use an atomic store after writing.
  • Load-then-store increments. Updates are lost under contention. Use fetch_add or Atomics.add.
  • static mut in threaded Rust. Bypasses the type system’s race protection. Use atomics or locks.
  • JavaScript reading while a worker writes. Torn reads and stale data. Use atomics or buffer ownership.
  • Trusting tests on a quiet machine. Races need load to appear. Stress-test and use native sanitizers.

Performance note

Uncontended atomic operations in Wasm cost a few nanoseconds; contended ones on a shared cache line can cost 50–100 ns. In a histogram benchmark, eight threads incrementing one shared atomic array ran slower than one thread; giving each thread a private histogram and summing at the end ran 6.8× faster than single-threaded.

Building a histogram of 100 million values Milliseconds to build a 256-bin histogram with one thread, with eight threads using atomic increments on a shared histogram, and with eight threads using private histograms merged at the end. ms per run 1 thread 210 ms 8 threads, shared atomic bins 340 ms 8 threads, private bins + merge 31 ms

Frequently Asked Questions

Are plain loads and stores of i32 atomic in Wasm? Aligned accesses are not torn in practice, but they carry no ordering guarantees. Use atomic instructions for anything shared.

Does Relaxed ordering exist in WebAssembly? Wasm atomics are sequentially consistent; Rust’s weaker orderings compile to the same instructions, so they are correct but not cheaper.

Can races corrupt the engine or escape the sandbox? No. Races produce unpredictable values within linear memory but never violate memory safety outside it.

How do I find a race I suspect exists? Reproduce the logic natively with ThreadSanitizer, or add assertions on invariants and run a stress test with many threads.

Is a Mutex in threaded Wasm expensive? Uncontended, it is a couple of atomic operations. Contended, waiting workers sleep in memory.atomic.wait32, which is efficient but adds wake-up latency.

← Back to SharedArrayBuffer, Atomics & Threading