Running a DSP Kernel in an AudioWorklet

This guide answers one task: run a compiled DSP kernel — a filter, a reverb, a pitch shifter — inside an AudioWorkletProcessor, where the callback must complete within a render quantum every single time or the user hears a click.

Prerequisites

  • [ ] A kernel compiled with no runtime dependencies: wasm-pack build --target web or emcc with -sSTANDALONE_WASM for the simplest case.
  • [ ] A secure context — AudioWorklet is unavailable over plain HTTP except on localhost.
  • [ ] Understanding that the worklet thread has no fetch, no WebAssembly.instantiateStreaming, and no DOM.
  • [ ] Headphones. Audio glitches are invisible on a waveform and obvious in your ears.

The constraint that shapes everything

process() is called on a dedicated real-time audio thread once per render quantum: 128 frames, which at 48 kHz is 2.67 milliseconds. If your callback takes longer than that, the audio graph underruns and the user hears a click or a dropout. There is no recovery, no backpressure and no queue — the deadline is hard.

That single fact rules out several things you would do anywhere else. You cannot allocate: a new Float32Array(n) inside the callback may trigger a garbage collection whose pause is longer than the budget. You cannot await anything. You cannot post a message and wait for a reply. And you cannot grow linear memory, because growth reallocates and copies the entire heap. Every buffer the kernel will ever need must exist before the first callback runs.

The 2.67 ms render quantum Each process callback must finish within one quantum. A kernel that computes in a fixed arena finishes with margin, while one that allocates can trigger a collection pause that overruns the deadline and produces an audible click. preallocated kernel — fits filter 128 frames · 0.9 ms idle margin deadline allocating kernel — overruns filter · 0.9 ms allocation + collection pause click The margin is not spare capacity — it absorbs scheduling jitter on a loaded machine. Aim to use under half the quantum. Measure with performance.now() around the kernel call and report the maximum, not the average; the maximum is what people hear.

Get the module into the worklet scope

The worklet thread cannot fetch. The standard pattern is to fetch the bytes on the main thread, pass them through port.postMessage, and instantiate synchronously inside the processor — synchronous instantiation is permitted for buffers under a few megabytes, and a DSP kernel is far below that limit.

// main thread
const ctx = new AudioContext({ latencyHint: 'interactive' });
await ctx.audioWorklet.addModule('/dsp-processor.js');
const bytes = await (await fetch('/dsp.wasm')).arrayBuffer();
const node = new AudioWorkletNode(ctx, 'dsp', { processorOptions: { bytes } });
node.connect(ctx.destination);
// dsp-processor.js — runs on the audio thread
class DspProcessor extends AudioWorkletProcessor {
  constructor({ processorOptions }) {
    super();
    const mod = new WebAssembly.Module(processorOptions.bytes);        // synchronous, once
    this.inst = new WebAssembly.Instance(mod, {});
    const { memory, init, in_ptr, out_ptr } = this.inst.exports;
    init(128);                                                          // preallocate everything
    this.inView  = new Float32Array(memory.buffer, in_ptr(),  128);
    this.outView = new Float32Array(memory.buffer, out_ptr(), 128);
  }
  process(inputs, outputs) {
    const inCh = inputs[0][0], outCh = outputs[0][0];
    if (!inCh) return true;
    this.inView.set(inCh);
    this.inst.exports.render(128);
    outCh.set(this.outView);
    return true;
  }
}
registerProcessor('dsp', DspProcessor);

Passing the bytes through processorOptions rather than a later postMessage means the instance exists before the first process() call, which removes an entire class of “the first half second is silence” bugs. The two views are built once in the constructor and are valid for the life of the node — but only because init() preallocated and the kernel never grows memory afterwards.

Planar channels, not interleaved

The Web Audio API hands you planar buffers: inputs[0] is an array of channels, each a separate Float32Array of 128 samples. Most C DSP code expects the same layout, but a lot of older code expects interleaved stereo. Converting per quantum is cheap, but doing it in JavaScript is unnecessary work on the audio thread — do it inside the module where it is a tight loop, or better, compile the kernel to accept planar input directly.

process(inputs, outputs) {
  const input = inputs[0];
  for (let c = 0; c < input.length; c++) this.chanViews[c].set(input[c]);
  this.inst.exports.render_planar(input.length, 128);
  const output = outputs[0];
  for (let c = 0; c < output.length; c++) output[c].set(this.outViews[c]);
  return true;
}

Allocate one region per channel in init() and keep a view per channel. Two channels at 128 frames is 1 kB of linear memory — the cost is irrelevant, and the resulting code has no per-quantum branching on channel count.

Expected output

A working kernel is silent about its success, so instrument it deliberately. Time the render call and report the worst case back to the main thread every second or so — never per quantum, because postMessage from the audio thread is itself work.

const t0 = performance.now();
this.inst.exports.render(128);
const dt = performance.now() - t0;
if (dt > this.worst) this.worst = dt;
if (++this.n % 375 === 0) { this.port.postMessage({ worstMs: this.worst }); this.worst = 0; }
// console, once per second
{ worstMs: 0.41 }   // 15% of the 2.67 ms budget — comfortable
{ worstMs: 0.44 }
{ worstMs: 2.91 }   // over budget: something allocated, or the machine was loaded

A worst case above about half the quantum is a warning even if you hear nothing: it means a busy machine will glitch. Investigate before shipping rather than after a support ticket.

Sample flow through a worklet-hosted kernel The audio graph hands the processor planar channel arrays. Each is copied into a preallocated region of linear memory, the kernel renders in place, and the results are copied back into the output channels. No allocation happens on this path. inputs[0][c] 128 floats per channel set() preallocated arena in region out region render() — filter state kept between quanta no malloc, no grow, no await outputs[0][c] straight to the graph Filter state — delay lines, biquad history, reverb tails — lives inside the arena and persists across calls, which is why the instance must outlive the quantum.

Getting audio in and out of the graph

A kernel is only useful once it is wired into a real signal path, and the surrounding graph decides how much headroom your callback actually has. Three sources cover almost every case: a MediaElementSource for playback of a file, a MediaStreamSource for microphone input, and an OscillatorNode or buffer source for synthesis. Each connects to your node the same way, and each brings a different latency profile with it.

Microphone input is the demanding one. The browser adds its own capture buffer ahead of your processor, and requesting a low latencyHint shortens that buffer at the cost of a smaller safety margin — exactly the margin your kernel is competing for. If you are building a live effect, start with latencyHint: 'interactive' and only move to a numeric hint after you have measured the worst-case render time on real hardware. Playback of a file is far more forgiving, because nothing downstream is waiting on a human.

const stream = await navigator.mediaDevices.getUserMedia({
  audio: { echoCancellation: false, noiseSuppression: false, autoGainControl: false },
});
const src = ctx.createMediaStreamSource(stream);
src.connect(node);                // your worklet node
node.connect(ctx.destination);

Disabling the built-in processing matters when your kernel is the effect: echo cancellation and automatic gain will fight anything you do, and they are enabled by default. For an analysis-only kernel you usually leave them on and connect the node to a GainNode with zero gain so the graph still pulls quanta through your processor without routing them to the speakers.

One more practical detail: an AudioContext starts suspended until a user gesture resumes it, so your first render never happens until someone clicks. Wire the resume() call to the same button that starts the feature, and check ctx.state before assuming silence is a bug in your kernel.

Parameters and control changes

Audio parameters change while the graph runs, and the temptation is to postMessage a new value from the UI. That works but arrives whenever the thread gets around to it, which means a knob turn can land mid quantum and produce a zipper noise. Two better options exist.

For sample-accurate automation, declare parameterDescriptors and read the parameters argument in process(). The browser hands you either a single value or a 128-element array when the value is being automated, and you pass it into the kernel so the change is applied per sample. For less critical controls — a preset switch, a bypass toggle — a SharedArrayBuffer of control values written by the main thread and read by the processor gives lock-free, allocation-free updates, using the same Atomics discipline you would use anywhere else.

Whichever you pick, smooth the parameter inside the kernel. Jumping a filter cutoff from 200 Hz to 2 kHz between quanta is audible as a click regardless of how the value arrived; a one-pole smoother over a few milliseconds costs two multiplies per sample and removes the artefact entirely.

The audio thread's rules The render callback runs on a real-time thread with a hard deadline. Anything that can block, allocate or wait belongs somewhere else. arithmetic on a fixed arena the only thing that belongs here allocation may trigger a pause and miss the deadline await, fetch, postMessage waits never; the thread cannot block parameter updates via a shared buffer or the message port Allocate every buffer before the first render callback and never again while audio is running. Report the maximum callback duration, not the average — the maximum is what listeners hear.

Gotchas

  • AudioWorkletNode constructed before addModule resolves. Await the module registration first, or you get “Failed to construct ‘AudioWorkletNode’: unregistered name”.
  • The first second is silent. The instance was created from a message that arrived after the first callbacks. Pass the bytes in processorOptions instead.
  • Views come back zero-length. Something inside the kernel called the allocator and grew memory. Preallocate in init(), and make the kernel’s allocator a fixed arena so growth is impossible.
  • Works on desktop, glitches on mobile. The quantum is the same but the CPU is not. Measure the worst case on the slowest device you support, not the fastest.
  • process() returns false and the node dies. Returning false tells the browser this processor will never produce output again. Return true unless you mean it.

Performance note

A biquad cascade over 128 frames costs a few microseconds — the interesting cost is everything around it. The two set() calls copy 512 bytes each and are negligible. What actually shows up in the worst case is the browser’s own graph overhead plus scheduling jitter, which on a loaded machine can be a millisecond on its own. That is why the target is under half the quantum: your kernel’s time is only part of what has to fit inside it.

Frequently Asked Questions

Can I use threads inside an AudioWorklet? No. The worklet scope has no Worker constructor, and spawning work elsewhere then waiting for it would violate the deadline anyway. Parallelism in audio comes from splitting the graph across nodes, not from threading a single kernel.

Is ScriptProcessorNode easier? It is deprecated, it runs on the main thread, and it introduces latency measured in tens of milliseconds. Anything new should use AudioWorklet; the extra setup is an hour once.

How big can the Wasm module be? Synchronous new WebAssembly.Module() is allowed on this thread for reasonably sized buffers, and DSP kernels are typically 10–80 kB. If your kernel is megabytes, compile it down or split the analysis work into a regular worker and keep only the real-time path in the worklet.

← Back to Media Processing & Codecs in Wasm