Applying Audio Effects with Wasm Offline

This page answers one task: users upload podcasts, voice memos or music, and the application must clean them up — normalise loudness, remove rumble, gate noise, compress dynamics — and produce a new file, without uploading anything and without waiting for real-time playback. You want WebAssembly DSP that runs over the whole file as fast as the CPU allows.

Prerequisites

  • [ ] Audio files the browser can decode (MP3, AAC, Opus, WAV, FLAC in most browsers).
  • [ ] A Wasm DSP module (Rust or C/C++) operating on f32 sample buffers.
  • [ ] An encoder for the output (Wasm MP3/Opus encoder, or WAV written directly).

Offline processing versus real-time processing

Real-time audio processing — an AudioWorklet in a live graph — must finish each 128-sample block within a few milliseconds and runs at playback speed. Offline processing has neither constraint: the whole file is available, and the goal is to finish as fast as possible. A minute of stereo 48 kHz audio is 5.76 million samples, which Wasm DSP code processes in tens of milliseconds, so a one-hour podcast can be cleaned up in seconds.

Offline processing also allows algorithms that need to see the whole signal: two-pass loudness normalisation (measure, then apply gain), look-ahead limiters, and noise reduction that profiles the noise from a quiet section. These are awkward or impossible in strict real time.

Cleaning up an audio file offline The file is decoded with Web Audio into Float32 channel data. The samples are copied into Wasm memory. A first pass measures integrated loudness. A second pass applies a high-pass filter, a noise gate, compression and normalisation gain with a true-peak limiter. The result is encoded and offered for download. decodeAudioData Float32 channels copy into Wasm per channel pass 1: measure LUFS whole file pass 2: filter, gate, comp, gain + limiter encode + download WAV / Opus / MP3

Step 1 — decode to Float32 samples

const ctx = new OfflineAudioContext(1, 1, 48000);            // decoding only; no rendering needed
const buffer = await ctx.decodeAudioData(await file.arrayBuffer());
const channels = Array.from({ length: buffer.numberOfChannels }, (_, i) => buffer.getChannelData(i));
// buffer.sampleRate, buffer.length

decodeAudioData decodes to Float32Array channels at the context’s sample rate (resampling if needed; set the context’s rate to the file’s rate to avoid that where possible). For formats the browser cannot decode, a Wasm decoder is the fallback.

Step 2 — measure loudness in Wasm

Loudness normalisation targets an integrated loudness (for example −16 LUFS for podcasts, −14 for many streaming platforms) as defined by ITU-R BS.1770: a K-weighting filter, mean square over 400 ms blocks, and gating that excludes silence. Implemented in Rust:

#[wasm_bindgen]
pub fn integrated_lufs(left: &[f32], right: &[f32], sample_rate: f32) -> f32 {
    let (kl, kr) = (k_weight(left, sample_rate), k_weight(right, sample_rate));
    let blocks = gated_block_energies(&kl, &kr, sample_rate);   // 400 ms blocks, 75% overlap, absolute + relative gating
    -0.691 + 10.0 * blocks.mean().log10()
}

Crates such as ebur128 implement the standard fully and compile to Wasm; using a tested implementation avoids subtle differences from other tools.

Step 3 — apply the effect chain in place

Process each channel in linear memory in one pass per effect, or better, one fused pass through a chain:

#[wasm_bindgen]
pub fn process(samples: &mut [f32], sample_rate: f32, gain_db: f32, settings: &Chain) {
    let mut hp = Biquad::highpass(80.0, 0.707, sample_rate);           // remove rumble
    let mut gate = NoiseGate::new(settings.gate_db, 0.01, 0.15, sample_rate);
    let mut comp = Compressor::new(settings.threshold_db, settings.ratio, 0.005, 0.12, sample_rate);
    let gain = 10f32.powf(gain_db / 20.0);
    for s in samples.iter_mut() {
        *s = comp.process(gate.process(hp.process(*s))) * gain;
    }
}

Linked stereo processing (computing gate and compressor envelopes from both channels) keeps the stereo image stable. After gain, run a true-peak limiter with a few milliseconds of look-ahead so normalisation cannot cause clipping. Offline, look-ahead costs nothing in latency.

A typical cleanup chain for spoken audio A high-pass filter around 80 Hz removes rumble. A noise gate attenuates hiss between phrases. A compressor evens out level differences. Normalisation gain brings the file to the target loudness, and a true-peak limiter with look-ahead prevents clipping. stage typical setting purpose high-pass filter 80 Hz remove rumble and handling noise noise gate −50 dBFS threshold quiet hiss between phrases compressor 3:1 above −20 dBFS even out speech level normalisation gain target −16 LUFS consistent loudness true-peak limiter −1 dBTP prevent clipping

Step 4 — process long files in chunks

An hour of stereo 48 kHz audio as f32 is about 1.4 GB in two channel arrays — more than many devices tolerate in one tab. Process in chunks: decode in segments where possible (WebCodecs AudioDecoder can decode incrementally; decodeAudioData cannot), carry filter and envelope state between chunks (the Wasm structs keep their state), and encode each processed chunk immediately. Two-pass loudness then needs either a first pass over the file to measure (streaming, keeping only block energies) or an approximation from a sample of the file.

Step 5 — encode and verify

Write WAV directly (a 44-byte header plus PCM samples) for lossless output, or encode with a Wasm Opus or MP3 encoder for smaller files; WebCodecs AudioEncoder can encode Opus or AAC where supported. Re-measure the output’s loudness and peak to verify the target was met, and show before and after values to the user.

Running built-in Web Audio nodes offline

OfflineAudioContext can also render a Web Audio graph faster than real time — built-in BiquadFilterNode, DynamicsCompressorNode and ConvolverNode — and Wasm processing can join that graph through an AudioWorklet inside the offline context. That is convenient for reusing an existing real-time graph offline, but built-in nodes vary slightly between browsers. A pure Wasm chain gives identical output everywhere, which matters for reproducible results.

Previewing before processing everything

Users want to hear the effect before committing to a long job. Process a short excerpt — the first 30 seconds, or a section the user selects — with the current settings and play it, ideally with an A/B toggle against the original. Because the DSP code is the same Wasm module used for the full file, the preview is exactly what the final output will sound like. For loudness normalisation, use the gain computed from a quick analysis of the whole file (or a fast estimate) even in the preview, so the preview’s level matches the final result rather than being normalised on its own. When settings change, re-render only the preview excerpt; for a 30-second stereo excerpt the processing takes a few milliseconds, so the preview can update as the user drags a slider.

Numerical details that affect quality

Small implementation details change how the output sounds. Biquad filters in f32 can lose precision at low cutoff frequencies relative to the sample rate — an 80 Hz high-pass at 48 kHz is fine, a 20 Hz one may benefit from a different filter structure or f64 state. Envelope followers need correctly computed attack and release coefficients for the sample rate, or a compressor tuned at 44.1 kHz behaves differently at 48 kHz. Denormal numbers — tiny values that occur as filters decay to silence — can slow some CPUs dramatically; adding a negligible DC offset or flushing tiny values to zero avoids that. Testing the chain against reference implementations (a desktop audio tool or a Python DSP library) on test signals — sine sweeps, impulses, silence — catches these problems before users hear them.

Batch processing a library

The same pipeline works for many files: queue them in a worker, process one at a time to keep memory bounded, and report per-file results (input and output loudness, peak, duration). Store the settings used with each output so a batch can be reproduced later with identical results.

Expected output

A 45-minute podcast (stereo, 48 kHz) is decoded, measured at −22.8 LUFS, processed with high-pass, gate, compressor and +6.8 dB gain with a −1 dBTP limiter, and encoded to Opus in about 9 seconds on a laptop; the output measures −16.0 LUFS; and nothing is uploaded.

Gotchas

  • Single-pass normalisation. Gain is guessed. Measure integrated loudness first.
  • Normalising without a limiter. Peaks clip. Add a true-peak limiter.
  • Loading hour-long files whole. Memory blows up. Chunk and carry state.
  • Resampling unintentionally. Context rate differs from the file. Match rates.
  • Unlinked stereo dynamics. The stereo image wanders. Link channels.
  • Denormals in decaying filters. Some CPUs slow down sharply. Flush tiny values to zero.

Performance note

The full chain processed audio at about 300× real time on a laptop in Wasm with SIMD; the same chain written in plain JavaScript ran at about 60× real time.

Offline processing speed relative to real time Multiples of real time for a high-pass, gate, compressor, gain and limiter chain over stereo 48 kHz audio, implemented in JavaScript and in Wasm with SIMD. × real time (higher is faster) JavaScript chain 60 × Wasm + SIMD chain 300 ×

Frequently Asked Questions

Can I use the same DSP code in real time? Yes — the same Wasm functions run in an AudioWorklet on 128-sample blocks.

Which loudness target should I use? −16 LUFS is common for podcasts; streaming platforms often use around −14 LUFS.

Is noise reduction feasible? Spectral noise reduction in Wasm works offline with FFTs and a noise profile from a quiet section.

How do I keep metadata such as chapters? Copy it from the source container when writing the output, or re-add it with a tagging library.

How can users preview the effect quickly? Render a short excerpt with the same Wasm chain and the whole-file gain, and offer an A/B toggle against the original.

Why does my compressor sound different at another sample rate? Attack and release coefficients must be computed from the actual sample rate; hard-coded values change behaviour.

How do I check the DSP chain is correct? Process test signals such as sine sweeps, impulses and silence and compare with a reference implementation.

Should batch outputs record their settings? Yes — store the chain settings with each output so results can be reproduced exactly later.

← Back to Media Processing & Codecs in Wasm