Implementing a Mutex with Atomics

This page answers one task: several workers share a WebAssembly memory and must take turns updating a shared structure, and you want to understand — and if necessary write — a mutex built from Atomics operations, including how waiting works, why it is efficient, and where it falls short.

Prerequisites

  • [ ] A cross-origin isolated page (crossOriginIsolated === true) so SharedArrayBuffer is available.
  • [ ] Workers sharing one WebAssembly.Memory({ shared: true }) or SharedArrayBuffer.
  • [ ] Familiarity with data races and why plain reads and writes are unsafe across threads.

What a mutex needs from the platform

A mutex needs two things: an atomic way to claim it, so two threads cannot both believe they hold it, and a way for a thread that fails to claim it to sleep until it is released, rather than spinning and burning CPU. Atomics.compareExchange provides the first: it changes a value from expected to new only if it still equals expected, returning the old value, all as one indivisible step. Atomics.wait and Atomics.notify provide the second: wait puts the calling thread to sleep as long as a memory location still holds a given value, and notify wakes sleepers on that location. The pair is a futex (fast userspace mutex) — the same primitive Linux locks are built on — and WebAssembly has the equivalent instructions (memory.atomic.wait32, memory.atomic.notify) for code compiled to Wasm.

A futex-based mutex is fast when uncontended: locking is a single compare-exchange, with no system involvement. Only when threads actually collide does anyone sleep or wake.

Locking and unlocking a futex mutex A thread tries to change the lock word from 0 to 1. If that succeeds it holds the lock. If not, it marks the lock as contended with 2 and waits while the word is 2. On unlock, the holder sets the word to 0 and, if it was contended, notifies one waiter, which wakes and tries again. CAS 0 → 1 uncontended fast path else set 2 mark contended wait while == 2 thread sleeps unlock: swap to 0 was it 2? notify one waiter it retries

Step 1 — the three-state lock

The classic design uses three states: 0 = unlocked, 1 = locked with no waiters, 2 = locked with possible waiters. Tracking waiters lets unlock skip the notify call when nobody is waiting.

// lock word at index i of an Int32Array over shared memory
export function lock(i32, i) {
  let c = Atomics.compareExchange(i32, i, 0, 1);
  if (c === 0) return;                                 // fast path: acquired
  if (c !== 2) c = Atomics.exchange(i32, i, 2);        // mark contended
  while (c !== 0) {
    Atomics.wait(i32, i, 2);                           // sleep while contended
    c = Atomics.exchange(i32, i, 2);                   // try to take it, keeping "contended"
  }
}

export function unlock(i32, i) {
  if (Atomics.exchange(i32, i, 0) === 2) {             // there may be waiters
    Atomics.notify(i32, i, 1);                         // wake one
  }
}

A thread that wakes sets the word to 2 even if it was the only waiter, because it cannot know whether others are sleeping; the cost is an occasional unnecessary notify, which is cheap.

Step 2 — the same lock inside WebAssembly

Locks are usually needed by the Wasm code itself. In Rust with the atomics target feature, std::sync::Mutex on wasm32-unknown-unknown uses exactly this futex design with memory.atomic.wait32/notify under the hood, so you rarely write it by hand:

use std::sync::Mutex;
static QUEUE: Mutex<Vec<Job>> = Mutex::new(Vec::new());

pub fn push_job(job: Job) { QUEUE.lock().unwrap().push(job); }

In C with Emscripten pthreads, pthread_mutex_t is implemented the same way. Writing your own is useful when you need a lock shared between JavaScript and Wasm code — both operating on the same word in shared memory — or a lock with special behaviour.

Step 3 — do not block the main thread

Atomics.wait throws a TypeError on the browser’s main thread: blocking it would freeze the page. Any code path that might contend for a lock must not run on the main thread, or must use a non-blocking alternative there. Atomics.waitAsync returns a promise instead of blocking, which lets the main thread wait for a lock asynchronously:

async function lockAsync(i32, i) {
  let c = Atomics.compareExchange(i32, i, 0, 1);
  if (c === 0) return;
  if (c !== 2) c = Atomics.exchange(i32, i, 2);
  while (c !== 0) {
    const r = Atomics.waitAsync(i32, i, 2);
    if (r.async) await r.value;                        // resolves on notify (or timeout)
    c = Atomics.exchange(i32, i, 2);
  }
}

Code compiled to Wasm that blocks in memory.atomic.wait32 traps on the main thread for the same reason. Emscripten works around this for some pthread operations by busy-waiting on the main thread, which is why its documentation recommends running pthread-heavy code off the main thread (PROXY_TO_PTHREAD). See using Atomics.waitAsync on the main thread.

Atomics.wait versus Atomics.waitAsync Atomics.wait blocks the calling thread until notified, which is efficient in workers but throws on the browser main thread. Atomics.waitAsync returns a promise that resolves on notify, letting the main thread wait without blocking, at the cost of async control flow. Atomics.wait blocks the thread workers only in browsers simple synchronous code inside workers Atomics.waitAsync returns a promise allowed on the main thread async control flow main thread

Step 4 — keep critical sections short

A mutex serialises the code it protects. Hold it only for the minimal update — push to a queue, swap a pointer — never across I/O, postMessage, callbacks into JavaScript, or long computations. A worker holding a lock while it waits for something else is the usual start of a deadlock. Where possible, avoid locks entirely: lock-free single-producer/single-consumer ring buffers, per-thread data merged at the end, or message passing.

Step 5 — understand fairness and starvation

This simple futex mutex is not fair. When the lock is released, a newly arriving thread can grab it with the fast path before the woken waiter retries, and a thread that keeps re-locking in a loop can starve others indefinitely. For most workloads that is fine — and faster than strict fairness. Where fairness matters (a UI-critical thread competing with background workers), use a ticket lock (each thread takes a number and waits for its turn) or redesign to avoid contention. Measure contention before optimising: a counter of how often the slow path runs tells you whether the lock matters at all.

Timeouts and robustness

Atomics.wait accepts a timeout in milliseconds and returns "timed-out" when it expires; use it to detect locks held for implausibly long, log, and fail rather than hang forever. A worker that crashes or is terminated while holding a lock leaves it locked; nothing releases it automatically. Robust designs either restart the whole worker pool on a worker failure or store the holder’s ID with the lock so a supervisor can detect and recover from an abandoned lock.

Testing a hand-written lock

Concurrency bugs in locks appear rarely and under load, so tests must create load. The standard stress test runs several workers that each acquire the lock, perform a non-atomic read-modify-write on shared data (read a counter, compute, write it back), and release — millions of times. If the lock is broken, lost updates appear as a final count below the expected total. Run it with more workers than cores to force preemption inside critical sections, and vary the work inside the critical section so timing differs between runs. Add a second check for mutual exclusion directly: an “owner” slot that each thread sets on entry and verifies on exit; any mismatch means two threads were inside at once. Repeat the stress test in every engine you support, since scheduling differs, and keep it in CI with a modest iteration count plus a longer nightly run.

Reader-writer locks and condition variables

A mutex is the building block for other primitives. A reader-writer lock allows many readers or one writer, useful for shared lookup tables read far more than written; it can be built from a state word counting readers plus a writer flag, with waiting through Atomics.wait on that word. Condition variables let a thread sleep until some condition on shared data becomes true — “the queue is not empty” — without polling: the waiter holds the mutex, checks the condition, and if false atomically releases the mutex and waits on a separate sequence word that signallers increment and notify. These are subtle to get right; prefer the implementations in Rust’s standard library, parking_lot, or pthreads in Emscripten unless you specifically need a lock shared with JavaScript.

Expected output

Four workers increment a shared counter one million times each under the lock and the final value is exactly four million; uncontended lock/unlock costs about 10 ns; contended runs show the slow path in 3% of acquisitions; the main thread acquires the same lock with lockAsync without blocking; and a lock held for over two seconds is reported as stuck.

Gotchas

  • Atomics.wait on the main thread. It throws. Use waitAsync or move the work.
  • Plain reads of the lock word. Not atomic. Use Atomics.load/exchange.
  • Holding locks across I/O or messages. Deadlocks and stalls. Keep critical sections tiny.
  • Assuming fairness. Threads can starve. Measure; use a ticket lock if needed.
  • Workers dying with the lock held. Nothing releases it. Restart the pool or track holders.

Performance note

Uncontended, a lock/unlock pair cost about 10 ns; with four workers contending continuously, throughput fell to about 6 million acquisitions per second in total — a reminder that a heavily contended lock serialises the work it protects.

Lock acquisitions per second Millions of lock and unlock pairs per second for a single uncontended thread and for four workers contending for the same lock continuously. million acquisitions per second 1 thread, uncontended 95 M/s 4 workers, contended 6 M/s

Frequently Asked Questions

Is std::sync::Mutex in Rust Wasm the same design? Yes — with the atomics feature it uses a futex built on Wasm wait and notify instructions.

Can JavaScript and Wasm share one lock? Yes, if both operate on the same word in shared memory with atomic operations.

What is the difference from a spinlock? A spinlock loops while waiting; a futex sleeps, saving CPU when waits are long.

Do I need a mutex for a single counter? No — Atomics.add updates it atomically without a lock.

How do I test that my lock really excludes other threads? Stress it with more workers than cores doing non-atomic read-modify-writes, and check an owner slot on entry and exit.

← Back to SharedArrayBuffer, Atomics & Threading