Double Buffering Data Between JavaScript and Wasm
This page answers one task: a WebAssembly module produces data continuously — simulation states, audio blocks, decoded frames — and JavaScript consumes it for rendering or output. With a single buffer, one side must wait for the other or copy the data. You want both sides to work concurrently on separate buffers and exchange them by swapping, not copying.
Prerequisites
- [ ] A producer in Wasm and a consumer in JavaScript (or the reverse), exchanging fixed-size data repeatedly.
- [ ] Data that fits twice (or three times) in linear memory.
- [ ] For the threaded version, shared memory and cross-origin isolation.
The idea
Double buffering keeps two buffers of the same size. At any moment one is the back buffer, which the producer writes the next state into, and the other is the front buffer, which the consumer reads the last complete state from. When the producer finishes, the roles swap: the freshly written buffer becomes the front, and the old front becomes the back for the next write. Swapping is changing an index — no data moves. The consumer never sees a half-written state (no torn reads), and the producer never waits for the consumer to finish copying.
In a single-threaded page the benefit is mainly zero copies and a clean separation: Wasm writes the next state while JavaScript’s render code reads the previous one through a view. With a worker producing on shared memory, the benefit is concurrency: the worker computes frame N+1 while the main thread renders frame N.
Step 1 — allocate two buffers in the module
use wasm_bindgen::prelude::*;
const N: usize = 100_000;
#[wasm_bindgen]
pub struct Sim { bufs: [Vec<f32>; 2], front: usize }
#[wasm_bindgen]
impl Sim {
#[wasm_bindgen(constructor)]
pub fn new() -> Sim { Sim { bufs: [vec![0.0; N * 2], vec![0.0; N * 2]], front: 0 } }
/// Computes the next state into the back buffer, then swaps.
pub fn step(&mut self, dt: f32) {
let (front, back) = if self.front == 0 { let (a, b) = self.bufs.split_at_mut(1); (&a[0], &mut b[0]) }
else { let (a, b) = self.bufs.split_at_mut(1); (&b[0], &mut a[0]) };
advance(front, back, dt); // read previous state, write next
self.front ^= 1; // publish
}
pub fn front_ptr(&self) -> *const f32 { self.bufs[self.front].as_ptr() }
pub fn len(&self) -> usize { N * 2 }
}
The simulation reads the previous state from the front buffer and writes the next state into the back buffer — natural for simulations that need the previous state anyway — then flips the index.
Step 2 — read the front buffer from JavaScript
const sim = new Sim();
function frame(t) {
sim.step(1 / 60);
const positions = new Float32Array(memory.buffer, sim.front_ptr(), sim.len()); // a view, no copy
draw(positions); // or writeBuffer to WebGPU
requestAnimationFrame(frame);
}
requestAnimationFrame(frame);
Single-threaded, step and draw alternate, so there is no race. The value of the two buffers here is that the simulation needs both states, and
JavaScript always reads a complete state without a copy.
Step 3 — run the producer in a worker with shared memory
With a worker, producer and consumer truly overlap. The swap index and a sequence counter live in shared memory and are accessed atomically:
// shared control: [frontIndex, sequence]
const ctl = new Int32Array(sharedControlBuffer);
// worker (producer)
for (;;) {
const back = Atomics.load(ctl, 0) ^ 1;
wasm.step_into(back, dt); // writes buffer `back`
Atomics.store(ctl, 0, back); // publish: back becomes front
Atomics.add(ctl, 1, 1); // new frame available
Atomics.notify(ctl, 1);
throttleToTargetRate();
}
// main thread (consumer)
function frame() {
const front = Atomics.load(ctl, 0);
const view = new Float32Array(sharedMemory.buffer, bufPtr[front], len);
draw(view);
requestAnimationFrame(frame);
}
There is a subtle race in plain double buffering with threads: if the producer finishes two frames while the consumer is still reading the front buffer, it starts writing into the buffer the consumer is reading. Either the producer must wait until the consumer has released the front buffer (an atomic “in use” flag), or you use three buffers.
Step 4 — use triple buffering for uneven rates
When the producer and consumer run at different rates — a simulation at 120 Hz, a renderer at 60 Hz or variable — triple buffering avoids both waiting and tearing. The three roles are front (being read), ready (latest complete, not yet picked up) and back (being written). The producer writes into back, then atomically swaps back and ready. The consumer, at the start of each frame, atomically swaps front and ready if a new frame is available. Each side only ever touches its own buffer, and the swaps are single atomic exchanges on a packed state word. Frames the consumer never picked up are simply overwritten — the consumer always sees the most recent complete state.
Step 5 — measure latency and memory
Double and triple buffering add latency: the consumer shows a state that is one (or up to two) frames old. For rendering, that is normally invisible; for input-driven interactions, measure end-to-end latency from input to display. Memory cost is two or three copies of the data in linear memory — trivial for small states, significant for large frames, where reducing resolution or using regions may matter more.
Audio and other real-time consumers
Audio worklets consume fixed-size blocks at a hard real-time rate and must never block. A ring of buffers (or a lock-free ring buffer of samples) between a Wasm producer and the worklet follows the same principle as triple buffering: the consumer never waits and never reads a block being written. When the producer falls behind, the consumer outputs silence or repeats data rather than blocking — dropouts are preferable to glitches from waiting. See implementing a lock-free ring buffer in shared memory.
When not to bother
If data is small and copying it costs microseconds, a copy per frame is simpler and perfectly fast. Double buffering earns its complexity when data is large, when the producer and consumer must overlap in time, or when torn reads would be visible.
Testing for torn reads
A torn read — the consumer seeing part of one state and part of the next — is rare and timing-dependent, so tests must make it likely and detectable. Have the producer write a generation number into the first and last element of each buffer, and fill the rest with values derived from it; the consumer checks that the first and last generation numbers match and that every value agrees. Run producer and consumer as fast as possible, with no throttling, for millions of frames, in every browser you support. Any mismatch is a torn read and means the handshake is wrong. Keep this test permanently; a refactor that changes the order of the swap and the write, or replaces an atomic store with a plain one, will reintroduce the race silently otherwise.
Expected output
The simulation of 100,000 particles produces states in a worker at 120 Hz while the main thread renders at 60 Hz; with triple buffering neither side waits; the renderer reads positions through a view with no copies; no frame shows a mix of two states; and added latency is measured at one frame on average.
Gotchas
- Plain double buffering with a fast producer. It can overwrite the buffer being read. Add a handshake or use three buffers.
- Non-atomic index updates across threads. Use
Atomics.loadandstoreorexchange. - Views created before memory growth. They detach in non-shared memory. Allocate buffers up front.
- Ignoring added latency. Measure input-to-display time for interactive features.
- Double buffering tiny data. A copy is simpler.
- No torn-read test. A refactor can reintroduce the race silently. Keep a generation-number stress test.
Performance note
For 100,000 particles (800 KB per state), copying the state out of Wasm each frame cost 0.3 ms; the double-buffered view cost nothing. Moving the simulation to a worker with triple buffering cut main-thread frame time from 14 ms to 3 ms.
Frequently Asked Questions
Is this the same as GPU double buffering? Same idea, applied to CPU-side data shared between JavaScript and Wasm.
Can the consumer write and the producer read? Yes — the pattern is symmetric; swap roles for JavaScript-to-Wasm flows.
Do I need shared memory? Only when producer and consumer run on different threads.
What about very large frames? Two or three copies may be too much memory; process regions or reduce resolution.
How do I detect torn reads in tests? Write a generation number at both ends of each buffer and assert they match on every read, under maximum producer and consumer speed.
Does triple buffering always show the newest state? Yes — the consumer picks up the most recently completed buffer; older unread frames are overwritten.
Related
- Implementing a lock-free ring buffer in shared memory — streams of blocks.
- Sharing buffers with WebGPU from Wasm — uploading the front buffer.
- Creating views into Wasm memory safely — view lifetimes.
- Using Atomics for Wasm thread synchronization — the swap primitives.
← Back to Zero-Copy Data Transfer Patterns