Implementing a Lock-Free Ring Buffer in Shared Memory
This page answers one task: a WebAssembly worker produces a continuous stream of data — audio samples, sensor readings, log records — and another thread consumes it, and the hand-off must be fast, allocation-free and never block either side.
Prerequisites
- [ ] A cross-origin isolated page, so
SharedArrayBufferis available. - [ ] A producer and a consumer on different threads: a worker and the main thread, or a worker and an
AudioWorkletProcessor. - [ ] Familiarity with using Atomics for Wasm thread synchronization.
Why a ring buffer, and why lock-free
postMessage is the default way to move data between threads, but every message allocates, is queued, and is delivered only when the receiving
thread’s event loop turns. For a real-time consumer such as an audio thread, which must produce 128 samples every 2.7 ms without fail, that is
unacceptable: allocation and delivery delays cause glitches. A ring buffer in shared memory avoids all of it. The producer writes into a fixed array
and advances a write index; the consumer reads from the same array and advances a read index; both indices wrap around at the end.
With exactly one producer and one consumer (SPSC), no locks are needed. Each index is written by only one side: the producer owns the write index, the consumer owns the read index. Each side reads the other’s index to know how much space or data is available. Atomic loads and stores of the two indices are all the synchronisation required, so neither thread ever waits on a lock held by the other — essential when one of them is a real-time audio thread that must never block.
Step 1 — lay out the shared buffer
Use one SharedArrayBuffer with a small header for the two indices and a data region. Using a capacity that is a power of two lets indices wrap with a
bit mask rather than a modulo:
export function createRing(capacity) { // capacity: power of two, in elements
const sab = new SharedArrayBuffer(8 + capacity * 4);
return { sab, capacity };
}
export function attach({ sab, capacity }) {
return {
idx: new Int32Array(sab, 0, 2), // [read, write]
data: new Float32Array(sab, 8, capacity),
mask: capacity - 1,
};
}
The trick that avoids ambiguity between “full” and “empty” is to let the indices count up freely (wrapping at 2³²) and mask only when indexing the
array. Then write - read is the number of readable items, and the buffer is full when that equals the capacity.
Step 2 — write from the producer
export function push(ring, samples) {
const { idx, data, mask } = ring;
const read = Atomics.load(idx, 0); // consumer's progress
const write = idx[1]; // our own index: plain read is fine
const free = data.length - ((write - read) | 0);
const n = Math.min(free, samples.length);
for (let i = 0; i < n; i++) data[(write + i) & mask] = samples[i];
Atomics.store(idx, 1, (write + n) | 0); // publish: data before index
return n; // may be < samples.length if full
}
The order matters. The data is written first, then the write index is published with Atomics.store. Atomic operations in JavaScript and WebAssembly
are sequentially consistent, which guarantees that a consumer that sees the new index also sees the data written before it. Publishing the index before
the data would let the consumer read stale samples.
Step 3 — read from the consumer
export function pull(ring, out) {
const { idx, data, mask } = ring;
const write = Atomics.load(idx, 1); // producer's progress
const read = idx[0];
const available = (write - read) | 0;
const n = Math.min(available, out.length);
for (let i = 0; i < n; i++) out[i] = data[(read + i) & mask];
Atomics.store(idx, 0, (read + n) | 0); // release the space
return n; // may be < out.length if empty
}
The consumer mirrors the producer: load the other side’s index atomically, copy, then publish its own index. In an AudioWorkletProcessor, pull
runs inside process(); if fewer samples are available than needed, fill the rest with silence and count an underrun rather than waiting.
Step 4 — let Wasm write directly into the ring
The JavaScript producer above copies samples from a Wasm buffer into the ring. If the producer module is itself threaded, with a shared linear memory,
the ring can live inside that memory and the module can write straight into it, removing the copy. Allocate the ring inside the module, export its
pointer, and have the consumer construct its views over the module’s memory.buffer at that offset. In Rust, the producer side uses AtomicU32
for the indices and writes samples into a slice; the JavaScript consumer uses Atomics on an Int32Array over the same bytes. Both sides must agree on
the layout exactly, as in
aligning data in linear memory.
Step 5 — handle full and empty without blocking
A real-time consumer must never block, so pull returns fewer items when the ring is empty. The producer, usually not real-time, has a choice when the
ring is full: drop data, retry later, or wait. Waiting can use Atomics.wait on the read index in a worker — the consumer calls Atomics.notify after
releasing space — which sleeps the producer efficiently instead of spinning. Never call Atomics.wait on the main thread (it throws) or in an audio
worklet (it would block audio). Size the ring for the worst-case jitter of the producer: for audio, 100–200 ms of samples usually absorbs garbage
collection pauses and scheduling delays in the producing worker.
Generalising to messages of varying size
Samples are fixed-size, which makes the ring simple. Variable-size messages — log records, events with payloads — need framing. The common technique
is to prefix each message with its length in the ring’s byte stream, and to let messages wrap around the end of the buffer, copying in two parts when
they do. The producer checks that the free space can hold the length prefix and the full message before writing either, and publishes the write index
only after the whole message is in place, so the consumer never sees a partial message. A slightly simpler variant inserts a padding marker when a
message would wrap, and starts the message at the beginning of the buffer instead; that wastes a little space but lets the consumer read each message as
one contiguous slice, which is convenient for TextDecoder and for zero-copy views. Both variants keep the SPSC rule: with several producers, either
give each its own ring or protect the write side with a lock, because two writers advancing one index will corrupt it.
Expected output
An audio worklet consumes 48,000 samples per second from a Wasm synthesiser running in a worker through a 8,192-sample ring; the underrun counter stays at zero during normal use, and no allocations occur in the audio thread.
Gotchas
- Publishing the index before the data. The consumer may read stale values. Write data, then
Atomics.storethe index. - Using
%with non-power-of-two capacity. Free-running indices wrap at 2³², which breaks modulo. Use a power-of-two capacity and a mask. - More than one producer or consumer. SPSC logic is not safe for multiple writers. Use one ring per producer or a lock.
Atomics.waiton the main thread or in audio. It throws or blocks real-time work. Wait only in workers.- Ring too small. Producer jitter causes underruns. Size for the worst-case delay.
Performance note
Moving 128-sample blocks through the ring cost about 0.4 µs per block in Chrome, against 18–60 µs for postMessage with a transferred buffer, plus the
latter’s allocation and scheduling jitter. In a 10-minute test, the ring produced zero underruns; postMessage produced 14 audible glitches under load.
Frequently Asked Questions
Is Atomics.load necessary for my own index?
No — only the owning thread writes it, so a plain read is fine. The other side’s index must be read atomically.
Can the ring hold structs instead of floats?
Yes — use a byte ring with fixed-size records and DataView or typed-array views per record.
Does this need cross-origin isolation?
Yes, for SharedArrayBuffer. Without it, fall back to postMessage.
What about Atomics.waitAsync?
The main thread can use it to wake when data arrives without blocking; see
using Atomics.waitAsync on the main thread.
How do I test a lock-free ring? Run producer and consumer in two workers for millions of items with a counting pattern, and have the consumer assert every value arrives once, in order. Vary the chunk sizes randomly to exercise wrap-around.
Can I monitor how full the ring is?
Yes — any thread can compute write - read from atomic loads. Sampling it shows whether the ring is sized well.
Related
- Passing audio samples without copying — the audio side in depth.
- Avoiding data races in shared Wasm memory — memory ordering rules.
- Running a DSP kernel in an AudioWorklet — the consumer.
- Reporting progress from Wasm to the UI — a simpler shared counter.
← Back to SharedArrayBuffer, Atomics & Threading