Passing Audio Samples Without Copying
This page answers one task: a WebAssembly DSP module runs inside an AudioWorkletProcessor, and every 128-sample block must move between the Web Audio
graph and the module with no allocation and as little copying as possible, so audio never glitches.
Prerequisites
- [ ] A Wasm module with a DSP function that processes a block of samples in place or from input to output buffers.
- [ ] An
AudioWorkletProcessorthat instantiates the module, as in running a DSP kernel in an AudioWorklet.
The constraints of the audio thread
An AudioWorkletProcessor’s process() method is called on the real-time audio thread for every render quantum: 128 frames, about 2.7 ms at 48 kHz.
It must return on time, every time. Anything that can pause the thread unpredictably — allocating JavaScript objects that the garbage collector later has
to reclaim, growing Wasm memory, waiting on locks — risks a dropout, heard as a click.
The samples arrive as Float32Arrays owned by the audio engine (inputs[i][channel]), and the output is written into Float32Arrays it provides
(outputs[i][channel]). These arrays are not in Wasm memory, so samples must be copied in and out. The goal is to do exactly one copy in and one copy out
per channel per block, into buffers that were allocated once at startup, and to create no new JavaScript objects in process() at all.
Step 1 — allocate buffers once, at construction
In the processor’s constructor, after instantiating the module, ask it to allocate input and output buffers for the maximum channel count, and create the views that point at them:
class WasmDsp extends AudioWorkletProcessor {
constructor(options) {
super();
const { module } = options.processorOptions; // compiled on the main thread, posted in
this.instance = new WebAssembly.Instance(module, {});
this.ex = this.instance.exports;
this.channels = 2;
this.inPtr = this.ex.alloc_f32(128 * this.channels);
this.outPtr = this.ex.alloc_f32(128 * this.channels);
this.refreshViews();
}
refreshViews() {
const buf = this.ex.memory.buffer;
this.inViews = [0, 1].map((c) => new Float32Array(buf, this.inPtr + c * 512, 128));
this.outViews = [0, 1].map((c) => new Float32Array(buf, this.outPtr + c * 512, 128));
this.buf = buf;
}
The module must not allocate during processing, so its memory never grows and the cached views stay valid. Still, a cheap identity check guards against growth caused by, for example, a parameter change that allocates.
Step 2 — copy in, process, copy out
process(inputs, outputs) {
if (this.ex.memory.buffer !== this.buf) this.refreshViews(); // rare: memory grew
const input = inputs[0], output = outputs[0];
const n = output[0].length; // 128
for (let c = 0; c < output.length; c++) {
if (input[c]) this.inViews[c].set(input[c]); else this.inViews[c].fill(0);
}
this.ex.process_block(this.inPtr, this.outPtr, n, output.length);
for (let c = 0; c < output.length; c++) output[c].set(this.outViews[c]);
return true;
}
}
registerProcessor("wasm-dsp", WasmDsp);
TypedArray.prototype.set between two Float32Arrays is a straight memory copy — 512 bytes per channel, a fraction of a microsecond. No arrays are
created, no closures are allocated, and the Wasm call takes only numbers. Missing inputs (an unconnected input has zero channels) are handled by zeroing
the buffer rather than allocating a silent array.
Step 3 — keep the Wasm side allocation-free
The module’s process_block must also avoid allocating: no Vec::new(), no String formatting, no Box per block. Filter states, delay lines and
scratch buffers are allocated at construction and reused. In Rust, a struct holding all state, created once, with a process(&mut self, input: &[f32], output: &mut [f32]) method that only reads and writes slices, is the right shape:
#[no_mangle]
pub unsafe extern "C" fn process_block(inp: *const f32, out: *mut f32, frames: usize, channels: usize) {
let dsp = &mut *DSP; // created once in init
for c in 0..channels {
let i = core::slice::from_raw_parts(inp.add(c * 128), frames);
let o = core::slice::from_raw_parts_mut(out.add(c * 128), frames);
dsp.channel(c).process(i, o);
}
}
Using raw extern "C" exports rather than wasm-bindgen glue inside the worklet also avoids glue that may allocate when converting arguments; wasm-bindgen
can be used in worklets but needs care to load without fetch, which AudioWorkletGlobalScope lacks.
Step 4 — compile on the main thread, instantiate in the worklet
The worklet scope cannot fetch, and compiling a module there would block the audio thread. Compile on the main thread and pass the compiled module in
processorOptions:
const module = await WebAssembly.compileStreaming(fetch("dsp.wasm"));
await ctx.audioWorklet.addModule("wasm-dsp-processor.js");
const node = new AudioWorkletNode(ctx, "wasm-dsp", { processorOptions: { module }, outputChannelCount: [2] });
Synchronous new WebAssembly.Instance in the constructor is acceptable because it happens once, before audio flows, and instantiating a precompiled module
is fast.
Step 5 — measure and watch for glitches
Check that process() allocates nothing: in Chrome DevTools, record a performance profile with the audio running and confirm there are no minor GCs
attributed to the audio worklet thread. Measure the time per block against the 2.7 ms budget, leaving ample headroom — a block that takes 1.5 ms on a fast
laptop can take 4 ms on a low-end phone. Count underruns by comparing currentTime progress with the expected rate, or listen for glitches while loading
the main thread heavily.
Parameters without allocation
Audio processors also receive parameters — gain, cutoff frequency, wet/dry mix — and those need the same allocation-free treatment. AudioParams
declared through parameterDescriptors arrive in process() as Float32Arrays of length 1 (constant for the block) or 128 (automated per sample).
Pass them to the module by writing them into another preallocated buffer, or, for a few scalar parameters, as plain numeric arguments to
process_block; both avoid allocation. For parameters changed from the main thread through the node’s port, do not process messages inside the hot
path in a way that allocates: let the onmessage handler write the new value into a preallocated Float32Array or directly into the module’s state with a
setter export, and have process() read it. Smoothing parameter changes inside the module — ramping to a new gain over a few milliseconds — avoids
zipper noise and keeps all the arithmetic in Wasm. With parameters handled this way, the only objects process() touches are the ones the engine already
created for it.
Going fully zero-copy with shared memory
The two copies per block are tiny, but some designs remove them anyway — usually not for speed, but because the audio data is produced elsewhere. When a
worker generates audio (a synthesiser, a decoder) and the worklet only plays it, the data can flow through a ring buffer in a SharedArrayBuffer, and the
worklet copies from the ring directly into the output arrays — one copy instead of two, and no message passing at all. If the producer is a threaded Wasm
module whose memory is shared, the ring can live inside its linear memory, so the producer writes samples directly into the place the worklet reads from.
That design is described in
implementing a lock-free ring buffer in shared memory.
It requires cross-origin isolation, and the worklet must never block on the producer: when the ring runs dry, output silence and count an underrun.
Expected output
The processor runs a stereo filter chain at 48 kHz with process() taking about 18 µs per block; the profiler shows zero allocations on the audio thread;
and a ten-minute run with heavy main-thread load produces no audible glitches.
Gotchas
- Creating arrays in
process().new Float32Array,slice, array literals — each allocates. Preallocate everything. - Growing memory during processing. It detaches views and takes time. Allocate all buffers at startup.
- Compiling in the worklet. It blocks the audio thread and cannot fetch. Compile on the main thread.
- Assuming 128 frames forever. The render quantum may become configurable. Use the array lengths you are given.
- Unconnected inputs.
inputs[0]may have zero channels. Zero the buffer instead of skipping the copy.
Performance note
Copying two channels in and out cost about 0.4 µs per block in Chrome; the filter chain itself took 17 µs. An earlier version that created a new
Float32Array view per channel per block triggered a minor GC on the audio thread roughly every 9 seconds, each causing an audible click.
Frequently Asked Questions
Can the engine’s input arrays be in Wasm memory directly? No. The audio engine owns them. One copy in and one out per block is the minimum without shared-memory producers.
Is the copy worth avoiding? Rarely. Half a microsecond per block is negligible; allocation and GC are the real enemies.
Can I use wasm-bindgen in an AudioWorklet?
Yes, with the web target and initSync given a precompiled module, but keep glue out of process().
What about Emscripten’s Wasm Audio Worklets? Emscripten offers an API that runs C/C++ audio callbacks in a worklet with shared memory; it follows the same allocation-free rules.
How many channels can one processor handle? As many as the budget allows. Allocate buffers for the maximum channel count at construction and process only the channels present.
Related
- Reading Wasm linear memory with typed arrays — the views used here.
- Creating views into Wasm memory safely — cache validation.
- Encoding audio in the browser with Wasm — the offline counterpart.
- Writing v128 SIMD intrinsics in Rust — speeding up the DSP.
← Back to Zero-Copy Data Transfer Patterns