Implementing a Lock-Free Ring Buffer in Shared Memory

This page answers one task: a WebAssembly worker produces a continuous stream of data — audio samples, sensor readings, log records — and another thread consumes it, and the hand-off must be fast, allocation-free and never block either side.

Prerequisites

  • [ ] A cross-origin isolated page, so SharedArrayBuffer is available.
  • [ ] A producer and a consumer on different threads: a worker and the main thread, or a worker and an AudioWorkletProcessor.
  • [ ] Familiarity with using Atomics for Wasm thread synchronization.

Why a ring buffer, and why lock-free

postMessage is the default way to move data between threads, but every message allocates, is queued, and is delivered only when the receiving thread’s event loop turns. For a real-time consumer such as an audio thread, which must produce 128 samples every 2.7 ms without fail, that is unacceptable: allocation and delivery delays cause glitches. A ring buffer in shared memory avoids all of it. The producer writes into a fixed array and advances a write index; the consumer reads from the same array and advances a read index; both indices wrap around at the end.

With exactly one producer and one consumer (SPSC), no locks are needed. Each index is written by only one side: the producer owns the write index, the consumer owns the read index. Each side reads the other’s index to know how much space or data is available. Atomic loads and stores of the two indices are all the synchronisation required, so neither thread ever waits on a lock held by the other — essential when one of them is a real-time audio thread that must never block.

A ring buffer in a SharedArrayBuffer The first eight bytes hold the read (R) and write (W) indices as 32-bit integers. The rest of the buffer is the data region. The producer writes after the write index and the consumer reads from the read index, with both wrapping around at the end of the data region. SharedArrayBuffer — header then data region R W consumed (free) readable data free for producer 0 8 8 + capacity

Step 1 — lay out the shared buffer

Use one SharedArrayBuffer with a small header for the two indices and a data region. Using a capacity that is a power of two lets indices wrap with a bit mask rather than a modulo:

export function createRing(capacity) {                     // capacity: power of two, in elements
  const sab = new SharedArrayBuffer(8 + capacity * 4);
  return { sab, capacity };
}

export function attach({ sab, capacity }) {
  return {
    idx: new Int32Array(sab, 0, 2),                         // [read, write]
    data: new Float32Array(sab, 8, capacity),
    mask: capacity - 1,
  };
}

The trick that avoids ambiguity between “full” and “empty” is to let the indices count up freely (wrapping at 2³²) and mask only when indexing the array. Then write - read is the number of readable items, and the buffer is full when that equals the capacity.

Step 2 — write from the producer

export function push(ring, samples) {
  const { idx, data, mask } = ring;
  const read = Atomics.load(idx, 0);                       // consumer's progress
  const write = idx[1];                                     // our own index: plain read is fine
  const free = data.length - ((write - read) | 0);
  const n = Math.min(free, samples.length);
  for (let i = 0; i < n; i++) data[(write + i) & mask] = samples[i];
  Atomics.store(idx, 1, (write + n) | 0);                   // publish: data before index
  return n;                                                 // may be < samples.length if full
}

The order matters. The data is written first, then the write index is published with Atomics.store. Atomic operations in JavaScript and WebAssembly are sequentially consistent, which guarantees that a consumer that sees the new index also sees the data written before it. Publishing the index before the data would let the consumer read stale samples.

Step 3 — read from the consumer

export function pull(ring, out) {
  const { idx, data, mask } = ring;
  const write = Atomics.load(idx, 1);                      // producer's progress
  const read = idx[0];
  const available = (write - read) | 0;
  const n = Math.min(available, out.length);
  for (let i = 0; i < n; i++) out[i] = data[(read + i) & mask];
  Atomics.store(idx, 0, (read + n) | 0);                    // release the space
  return n;                                                 // may be < out.length if empty
}

The consumer mirrors the producer: load the other side’s index atomically, copy, then publish its own index. In an AudioWorkletProcessor, pull runs inside process(); if fewer samples are available than needed, fill the rest with silence and count an underrun rather than waiting.

One push and one pull through the ring The producer loads the read index to compute free space, writes samples into the data region, then stores the new write index. The consumer loads the write index to compute available samples, copies them out, then stores the new read index. producer (worker) shared buffer consumer (audio) Atomics.load(read) write samples at write & mask Atomics.store(write + n) Atomics.load(write) copy out, Atomics.store(read + n)

Step 4 — let Wasm write directly into the ring

The JavaScript producer above copies samples from a Wasm buffer into the ring. If the producer module is itself threaded, with a shared linear memory, the ring can live inside that memory and the module can write straight into it, removing the copy. Allocate the ring inside the module, export its pointer, and have the consumer construct its views over the module’s memory.buffer at that offset. In Rust, the producer side uses AtomicU32 for the indices and writes samples into a slice; the JavaScript consumer uses Atomics on an Int32Array over the same bytes. Both sides must agree on the layout exactly, as in aligning data in linear memory.

Step 5 — handle full and empty without blocking

A real-time consumer must never block, so pull returns fewer items when the ring is empty. The producer, usually not real-time, has a choice when the ring is full: drop data, retry later, or wait. Waiting can use Atomics.wait on the read index in a worker — the consumer calls Atomics.notify after releasing space — which sleeps the producer efficiently instead of spinning. Never call Atomics.wait on the main thread (it throws) or in an audio worklet (it would block audio). Size the ring for the worst-case jitter of the producer: for audio, 100–200 ms of samples usually absorbs garbage collection pauses and scheduling delays in the producing worker.

Generalising to messages of varying size

Samples are fixed-size, which makes the ring simple. Variable-size messages — log records, events with payloads — need framing. The common technique is to prefix each message with its length in the ring’s byte stream, and to let messages wrap around the end of the buffer, copying in two parts when they do. The producer checks that the free space can hold the length prefix and the full message before writing either, and publishes the write index only after the whole message is in place, so the consumer never sees a partial message. A slightly simpler variant inserts a padding marker when a message would wrap, and starts the message at the beginning of the buffer instead; that wastes a little space but lets the consumer read each message as one contiguous slice, which is convenient for TextDecoder and for zero-copy views. Both variants keep the SPSC rule: with several producers, either give each its own ring or protect the write side with a lock, because two writers advancing one index will corrupt it.

Expected output

An audio worklet consumes 48,000 samples per second from a Wasm synthesiser running in a worker through a 8,192-sample ring; the underrun counter stays at zero during normal use, and no allocations occur in the audio thread.

Gotchas

  • Publishing the index before the data. The consumer may read stale values. Write data, then Atomics.store the index.
  • Using % with non-power-of-two capacity. Free-running indices wrap at 2³², which breaks modulo. Use a power-of-two capacity and a mask.
  • More than one producer or consumer. SPSC logic is not safe for multiple writers. Use one ring per producer or a lock.
  • Atomics.wait on the main thread or in audio. It throws or blocks real-time work. Wait only in workers.
  • Ring too small. Producer jitter causes underruns. Size for the worst-case delay.

Performance note

Moving 128-sample blocks through the ring cost about 0.4 µs per block in Chrome, against 18–60 µs for postMessage with a transferred buffer, plus the latter’s allocation and scheduling jitter. In a 10-minute test, the ring produced zero underruns; postMessage produced 14 audible glitches under load.

Handing one 128-sample block to the audio thread Microseconds per block to move audio from a worker to an AudioWorklet, using the shared-memory ring buffer and using postMessage with a transferred buffer. microseconds per block SPSC ring buffer 0.4 µs postMessage (transfer) 35 µs

Frequently Asked Questions

Is Atomics.load necessary for my own index? No — only the owning thread writes it, so a plain read is fine. The other side’s index must be read atomically.

Can the ring hold structs instead of floats? Yes — use a byte ring with fixed-size records and DataView or typed-array views per record.

Does this need cross-origin isolation? Yes, for SharedArrayBuffer. Without it, fall back to postMessage.

What about Atomics.waitAsync? The main thread can use it to wake when data arrives without blocking; see using Atomics.waitAsync on the main thread.

How do I test a lock-free ring? Run producer and consumer in two workers for millions of items with a counting pattern, and have the consumer assert every value arrives once, in order. Vary the chunk sizes randomly to exercise wrap-around.

Can I monitor how full the ring is? Yes — any thread can compute write - read from atomic loads. Sampling it shows whether the ring is sized well.

← Back to SharedArrayBuffer, Atomics & Threading