Double Buffering Data Between JavaScript and Wasm

This page answers one task: a WebAssembly module produces data continuously — simulation states, audio blocks, decoded frames — and JavaScript consumes it for rendering or output. With a single buffer, one side must wait for the other or copy the data. You want both sides to work concurrently on separate buffers and exchange them by swapping, not copying.

Prerequisites

  • [ ] A producer in Wasm and a consumer in JavaScript (or the reverse), exchanging fixed-size data repeatedly.
  • [ ] Data that fits twice (or three times) in linear memory.
  • [ ] For the threaded version, shared memory and cross-origin isolation.

The idea

Double buffering keeps two buffers of the same size. At any moment one is the back buffer, which the producer writes the next state into, and the other is the front buffer, which the consumer reads the last complete state from. When the producer finishes, the roles swap: the freshly written buffer becomes the front, and the old front becomes the back for the next write. Swapping is changing an index — no data moves. The consumer never sees a half-written state (no torn reads), and the producer never waits for the consumer to finish copying.

In a single-threaded page the benefit is mainly zero copies and a clean separation: Wasm writes the next state while JavaScript’s render code reads the previous one through a view. With a worker producing on shared memory, the benefit is concurrency: the worker computes frame N+1 while the main thread renders frame N.

A double-buffered frame loop The producer writes the next state into the back buffer while the consumer reads the front buffer. When the producer finishes, it publishes by swapping the index, so the back becomes the front. The consumer picks up the new front on its next frame, and the producer starts writing into the other buffer. producer writes back buffer B consumer reads front buffer A producer done swap index front = B consumer reads new state producer writes A next state

Step 1 — allocate two buffers in the module

use wasm_bindgen::prelude::*;

const N: usize = 100_000;

#[wasm_bindgen]
pub struct Sim { bufs: [Vec<f32>; 2], front: usize }

#[wasm_bindgen]
impl Sim {
    #[wasm_bindgen(constructor)]
    pub fn new() -> Sim { Sim { bufs: [vec![0.0; N * 2], vec![0.0; N * 2]], front: 0 } }

    /// Computes the next state into the back buffer, then swaps.
    pub fn step(&mut self, dt: f32) {
        let (front, back) = if self.front == 0 { let (a, b) = self.bufs.split_at_mut(1); (&a[0], &mut b[0]) }
                            else { let (a, b) = self.bufs.split_at_mut(1); (&b[0], &mut a[0]) };
        advance(front, back, dt);                   // read previous state, write next
        self.front ^= 1;                            // publish
    }
    pub fn front_ptr(&self) -> *const f32 { self.bufs[self.front].as_ptr() }
    pub fn len(&self) -> usize { N * 2 }
}

The simulation reads the previous state from the front buffer and writes the next state into the back buffer — natural for simulations that need the previous state anyway — then flips the index.

Step 2 — read the front buffer from JavaScript

const sim = new Sim();
function frame(t) {
  sim.step(1 / 60);
  const positions = new Float32Array(memory.buffer, sim.front_ptr(), sim.len());   // a view, no copy
  draw(positions);                                                                // or writeBuffer to WebGPU
  requestAnimationFrame(frame);
}
requestAnimationFrame(frame);

Single-threaded, step and draw alternate, so there is no race. The value of the two buffers here is that the simulation needs both states, and JavaScript always reads a complete state without a copy.

Step 3 — run the producer in a worker with shared memory

With a worker, producer and consumer truly overlap. The swap index and a sequence counter live in shared memory and are accessed atomically:

// shared control: [frontIndex, sequence]
const ctl = new Int32Array(sharedControlBuffer);

// worker (producer)
for (;;) {
  const back = Atomics.load(ctl, 0) ^ 1;
  wasm.step_into(back, dt);                       // writes buffer `back`
  Atomics.store(ctl, 0, back);                    // publish: back becomes front
  Atomics.add(ctl, 1, 1);                         // new frame available
  Atomics.notify(ctl, 1);
  throttleToTargetRate();
}

// main thread (consumer)
function frame() {
  const front = Atomics.load(ctl, 0);
  const view = new Float32Array(sharedMemory.buffer, bufPtr[front], len);
  draw(view);
  requestAnimationFrame(frame);
}

There is a subtle race in plain double buffering with threads: if the producer finishes two frames while the consumer is still reading the front buffer, it starts writing into the buffer the consumer is reading. Either the producer must wait until the consumer has released the front buffer (an atomic “in use” flag), or you use three buffers.

Double versus triple buffering with a worker producer Double buffering needs the producer to wait if it finishes twice while the consumer still reads, or reads can tear. Triple buffering gives the producer a spare buffer, so it never waits and the consumer always gets the latest complete frame, at the cost of one more buffer of memory. double buffering two buffers producer may wait on consumer simple handshake matched rates triple buffering three buffers producer never waits consumer gets latest frame uneven rates

Step 4 — use triple buffering for uneven rates

When the producer and consumer run at different rates — a simulation at 120 Hz, a renderer at 60 Hz or variable — triple buffering avoids both waiting and tearing. The three roles are front (being read), ready (latest complete, not yet picked up) and back (being written). The producer writes into back, then atomically swaps back and ready. The consumer, at the start of each frame, atomically swaps front and ready if a new frame is available. Each side only ever touches its own buffer, and the swaps are single atomic exchanges on a packed state word. Frames the consumer never picked up are simply overwritten — the consumer always sees the most recent complete state.

Step 5 — measure latency and memory

Double and triple buffering add latency: the consumer shows a state that is one (or up to two) frames old. For rendering, that is normally invisible; for input-driven interactions, measure end-to-end latency from input to display. Memory cost is two or three copies of the data in linear memory — trivial for small states, significant for large frames, where reducing resolution or using regions may matter more.

Audio and other real-time consumers

Audio worklets consume fixed-size blocks at a hard real-time rate and must never block. A ring of buffers (or a lock-free ring buffer of samples) between a Wasm producer and the worklet follows the same principle as triple buffering: the consumer never waits and never reads a block being written. When the producer falls behind, the consumer outputs silence or repeats data rather than blocking — dropouts are preferable to glitches from waiting. See implementing a lock-free ring buffer in shared memory.

When not to bother

If data is small and copying it costs microseconds, a copy per frame is simpler and perfectly fast. Double buffering earns its complexity when data is large, when the producer and consumer must overlap in time, or when torn reads would be visible.

Testing for torn reads

A torn read — the consumer seeing part of one state and part of the next — is rare and timing-dependent, so tests must make it likely and detectable. Have the producer write a generation number into the first and last element of each buffer, and fill the rest with values derived from it; the consumer checks that the first and last generation numbers match and that every value agrees. Run producer and consumer as fast as possible, with no throttling, for millions of frames, in every browser you support. Any mismatch is a torn read and means the handshake is wrong. Keep this test permanently; a refactor that changes the order of the swap and the write, or replaces an atomic store with a plain one, will reintroduce the race silently otherwise.

Expected output

The simulation of 100,000 particles produces states in a worker at 120 Hz while the main thread renders at 60 Hz; with triple buffering neither side waits; the renderer reads positions through a view with no copies; no frame shows a mix of two states; and added latency is measured at one frame on average.

Gotchas

  • Plain double buffering with a fast producer. It can overwrite the buffer being read. Add a handshake or use three buffers.
  • Non-atomic index updates across threads. Use Atomics.load and store or exchange.
  • Views created before memory growth. They detach in non-shared memory. Allocate buffers up front.
  • Ignoring added latency. Measure input-to-display time for interactive features.
  • Double buffering tiny data. A copy is simpler.
  • No torn-read test. A refactor can reintroduce the race silently. Keep a generation-number stress test.

Performance note

For 100,000 particles (800 KB per state), copying the state out of Wasm each frame cost 0.3 ms; the double-buffered view cost nothing. Moving the simulation to a worker with triple buffering cut main-thread frame time from 14 ms to 3 ms.

Main-thread frame time for a 100,000-particle simulation Milliseconds per frame on the main thread with the simulation running there and a copy per frame, with double-buffered views on the main thread, and with the simulation in a worker using triple buffering. main-thread ms per frame sim on main + copy 14.3 ms sim on main + views 14 ms worker + triple buffering 3 ms

Frequently Asked Questions

Is this the same as GPU double buffering? Same idea, applied to CPU-side data shared between JavaScript and Wasm.

Can the consumer write and the producer read? Yes — the pattern is symmetric; swap roles for JavaScript-to-Wasm flows.

Do I need shared memory? Only when producer and consumer run on different threads.

What about very large frames? Two or three copies may be too much memory; process regions or reduce resolution.

How do I detect torn reads in tests? Write a generation number at both ends of each buffer and assert they match on every read, under maximum producer and consumer speed.

Does triple buffering always show the newest state? Yes — the consumer picks up the most recently completed buffer; older unread frames are overwritten.

← Back to Zero-Copy Data Transfer Patterns