Implementing a Mutex with Atomics
This page answers one task: several workers share a WebAssembly memory and must take turns updating a shared structure, and you want to understand — and if
necessary write — a mutex built from Atomics operations, including how waiting works, why it is efficient, and where it falls short.
Prerequisites
- [ ] A cross-origin isolated page (
crossOriginIsolated === true) soSharedArrayBufferis available. - [ ] Workers sharing one
WebAssembly.Memory({ shared: true })orSharedArrayBuffer. - [ ] Familiarity with data races and why plain reads and writes are unsafe across threads.
What a mutex needs from the platform
A mutex needs two things: an atomic way to claim it, so two threads cannot both believe they hold it, and a way for a thread that fails to claim it to
sleep until it is released, rather than spinning and burning CPU. Atomics.compareExchange provides the first: it changes a value from expected to new
only if it still equals expected, returning the old value, all as one indivisible step. Atomics.wait and Atomics.notify provide the second: wait
puts the calling thread to sleep as long as a memory location still holds a given value, and notify wakes sleepers on that location. The pair is a
futex (fast userspace mutex) — the same primitive Linux locks are built on — and WebAssembly has the equivalent instructions
(memory.atomic.wait32, memory.atomic.notify) for code compiled to Wasm.
A futex-based mutex is fast when uncontended: locking is a single compare-exchange, with no system involvement. Only when threads actually collide does anyone sleep or wake.
Step 1 — the three-state lock
The classic design uses three states: 0 = unlocked, 1 = locked with no waiters, 2 = locked with possible waiters. Tracking waiters lets unlock skip
the notify call when nobody is waiting.
// lock word at index i of an Int32Array over shared memory
export function lock(i32, i) {
let c = Atomics.compareExchange(i32, i, 0, 1);
if (c === 0) return; // fast path: acquired
if (c !== 2) c = Atomics.exchange(i32, i, 2); // mark contended
while (c !== 0) {
Atomics.wait(i32, i, 2); // sleep while contended
c = Atomics.exchange(i32, i, 2); // try to take it, keeping "contended"
}
}
export function unlock(i32, i) {
if (Atomics.exchange(i32, i, 0) === 2) { // there may be waiters
Atomics.notify(i32, i, 1); // wake one
}
}
A thread that wakes sets the word to 2 even if it was the only waiter, because it cannot know whether others are sleeping; the cost is an occasional
unnecessary notify, which is cheap.
Step 2 — the same lock inside WebAssembly
Locks are usually needed by the Wasm code itself. In Rust with the atomics target feature, std::sync::Mutex on wasm32-unknown-unknown uses exactly
this futex design with memory.atomic.wait32/notify under the hood, so you rarely write it by hand:
use std::sync::Mutex;
static QUEUE: Mutex<Vec<Job>> = Mutex::new(Vec::new());
pub fn push_job(job: Job) { QUEUE.lock().unwrap().push(job); }
In C with Emscripten pthreads, pthread_mutex_t is implemented the same way. Writing your own is useful when you need a lock shared between JavaScript
and Wasm code — both operating on the same word in shared memory — or a lock with special behaviour.
Step 3 — do not block the main thread
Atomics.wait throws a TypeError on the browser’s main thread: blocking it would freeze the page. Any code path that might contend for a lock must not
run on the main thread, or must use a non-blocking alternative there. Atomics.waitAsync returns a promise instead of blocking, which lets the main thread
wait for a lock asynchronously:
async function lockAsync(i32, i) {
let c = Atomics.compareExchange(i32, i, 0, 1);
if (c === 0) return;
if (c !== 2) c = Atomics.exchange(i32, i, 2);
while (c !== 0) {
const r = Atomics.waitAsync(i32, i, 2);
if (r.async) await r.value; // resolves on notify (or timeout)
c = Atomics.exchange(i32, i, 2);
}
}
Code compiled to Wasm that blocks in memory.atomic.wait32 traps on the main thread for the same reason. Emscripten works around this for some pthread
operations by busy-waiting on the main thread, which is why its documentation recommends running pthread-heavy code off the main thread (PROXY_TO_PTHREAD).
See using Atomics.waitAsync on the main thread.
Step 4 — keep critical sections short
A mutex serialises the code it protects. Hold it only for the minimal update — push to a queue, swap a pointer — never across I/O, postMessage,
callbacks into JavaScript, or long computations. A worker holding a lock while it waits for something else is the usual start of a deadlock. Where
possible, avoid locks entirely: lock-free single-producer/single-consumer ring buffers, per-thread data merged at the end, or message passing.
Step 5 — understand fairness and starvation
This simple futex mutex is not fair. When the lock is released, a newly arriving thread can grab it with the fast path before the woken waiter retries, and a thread that keeps re-locking in a loop can starve others indefinitely. For most workloads that is fine — and faster than strict fairness. Where fairness matters (a UI-critical thread competing with background workers), use a ticket lock (each thread takes a number and waits for its turn) or redesign to avoid contention. Measure contention before optimising: a counter of how often the slow path runs tells you whether the lock matters at all.
Timeouts and robustness
Atomics.wait accepts a timeout in milliseconds and returns "timed-out" when it expires; use it to detect locks held for implausibly long, log, and fail
rather than hang forever. A worker that crashes or is terminated while holding a lock leaves it locked; nothing releases it automatically. Robust designs
either restart the whole worker pool on a worker failure or store the holder’s ID with the lock so a supervisor can detect and recover from an abandoned
lock.
Testing a hand-written lock
Concurrency bugs in locks appear rarely and under load, so tests must create load. The standard stress test runs several workers that each acquire the lock, perform a non-atomic read-modify-write on shared data (read a counter, compute, write it back), and release — millions of times. If the lock is broken, lost updates appear as a final count below the expected total. Run it with more workers than cores to force preemption inside critical sections, and vary the work inside the critical section so timing differs between runs. Add a second check for mutual exclusion directly: an “owner” slot that each thread sets on entry and verifies on exit; any mismatch means two threads were inside at once. Repeat the stress test in every engine you support, since scheduling differs, and keep it in CI with a modest iteration count plus a longer nightly run.
Reader-writer locks and condition variables
A mutex is the building block for other primitives. A reader-writer lock allows many readers or one writer, useful for shared lookup tables read far more
than written; it can be built from a state word counting readers plus a writer flag, with waiting through Atomics.wait on that word. Condition
variables let a thread sleep until some condition on shared data becomes true — “the queue is not empty” — without polling: the waiter holds the mutex,
checks the condition, and if false atomically releases the mutex and waits on a separate sequence word that signallers increment and notify. These are
subtle to get right; prefer the implementations in Rust’s standard library, parking_lot, or pthreads in Emscripten unless you specifically need a lock
shared with JavaScript.
Expected output
Four workers increment a shared counter one million times each under the lock and the final value is exactly four million; uncontended lock/unlock costs
about 10 ns; contended runs show the slow path in 3% of acquisitions; the main thread acquires the same lock with lockAsync without blocking; and a lock
held for over two seconds is reported as stuck.
Gotchas
Atomics.waiton the main thread. It throws. UsewaitAsyncor move the work.- Plain reads of the lock word. Not atomic. Use
Atomics.load/exchange. - Holding locks across I/O or messages. Deadlocks and stalls. Keep critical sections tiny.
- Assuming fairness. Threads can starve. Measure; use a ticket lock if needed.
- Workers dying with the lock held. Nothing releases it. Restart the pool or track holders.
Performance note
Uncontended, a lock/unlock pair cost about 10 ns; with four workers contending continuously, throughput fell to about 6 million acquisitions per second in total — a reminder that a heavily contended lock serialises the work it protects.
Frequently Asked Questions
Is std::sync::Mutex in Rust Wasm the same design?
Yes — with the atomics feature it uses a futex built on Wasm wait and notify instructions.
Can JavaScript and Wasm share one lock? Yes, if both operate on the same word in shared memory with atomic operations.
What is the difference from a spinlock? A spinlock loops while waiting; a futex sleeps, saving CPU when waits are long.
Do I need a mutex for a single counter?
No — Atomics.add updates it atomically without a lock.
How do I test that my lock really excludes other threads? Stress it with more workers than cores doing non-atomic read-modify-writes, and check an owner slot on entry and exit.
Related
- Using Atomics for Wasm thread synchronization — the primitives.
- Debugging deadlocks in threaded Wasm — when locks go wrong.
- Implementing a lock-free ring buffer in shared memory — avoiding locks.
- Avoiding data races in shared Wasm memory — the underlying rules.
← Back to SharedArrayBuffer, Atomics & Threading