Media Processing & Codecs in Wasm

Media is the workload WebAssembly was effectively designed for: large buffers of bytes, arithmetic-heavy inner loops, and mature C implementations that nobody wants to rewrite in JavaScript. A single 1080p frame is roughly 8.3 MB of RGBA; a one-minute stereo track at 48 kHz is 23 MB of float samples. Moving that much data across the boundary carelessly costs more than the processing, so the engineering problem is rarely the codec itself — it is owning the buffers, keeping them in one place, and letting the module work in the memory the data already lives in.

Prerequisites

  • [ ] A toolchain that can build a C or Rust media library — emcc 3.1+ or wasm-pack 0.13+.
  • [ ] A dev server that sets Cross-Origin-Opener-Policy and Cross-Origin-Embedder-Policy if you plan to use threads.
  • [ ] Chrome 94+ or Firefox 130+ for WebCodecs interop examples; everything else works anywhere.
  • [ ] A test asset set: one short H.264 clip, one large JPEG, one WAV file. Synthetic inputs hide real bottlenecks.

What a media pipeline looks like in memory

The shape that works is always the same. The module owns a region of linear memory sized for the largest frame you expect, JavaScript writes into that region through a typed-array view, the module processes in place or into a second region, and JavaScript reads the output through a second view. The frame never exists twice unless you ask for it twice.

Frame buffers inside one linear memory Linear memory is divided into an input region the page writes into, scratch space the codec uses, and an output region the page reads from. Pointers to each region are handed to JavaScript once, and reused for every frame. one instance, one memory, three regions input frame region page writes here codec scratch never crosses the boundary output frame region page reads a view of this per frame, the page does exactly two things: view.set(frameBytes) — one copy in new Uint8ClampedArray(buf, out, len) — zero copies out Allocating per frame instead of reusing these regions is the single most common cause of a pipeline that starts fast and degrades. It also churns the allocator, which triggers memory.grow, which detaches every view the page is holding.

The region pointers are obtained once, at startup, and cached:

const cap = 1920 * 1080 * 4;
const inPtr  = mod.exports.alloc_input(cap);
const outPtr = mod.exports.alloc_output(cap);
// reused for every frame for the life of the instance

If your module can guarantee it never grows memory after this point — which a fixed arena can — the views can be cached too, and the per-frame cost drops to one set() and one function call.

Building a codec library for the browser

Most real pipelines start from an existing C library. The build is unremarkable; the flags that matter are the ones that decide what the JavaScript side can reach and how big the result is.

emcc decoder.c libjpeg/*.c \
  -O3 -msimd128 \
  -sMODULARIZE=1 -sEXPORT_ES6=1 \
  -sEXPORTED_FUNCTIONS='["_alloc_input","_alloc_output","_decode_frame","_free"]' \
  -sEXPORTED_RUNTIME_METHODS='["HEAPU8"]' \
  -sALLOW_MEMORY_GROWTH=1 -sINITIAL_MEMORY=64MB \
  -o decoder.mjs

-msimd128 is usually worth 1.5–3× on pixel loops and costs a few kilobytes; ship it behind a capability check with a baseline build as covered in shipping SIMD and baseline builds together. -sINITIAL_MEMORY is set high deliberately: a decoder that grows from 16 MB to 100 MB during its first frame pays for several reallocations and invalidates every view in flight. For a Rust pipeline the equivalent is wasm-pack build --release with RUSTFLAGS="-C target-feature=+simd128" and an arena crate instead of the default allocator.

Keeping the main thread out of it

Decoding a frame takes milliseconds; decoding thirty takes a frame budget’s worth of blocked UI. Media pipelines belong in a worker, with the instance and its memory owned entirely by that worker. The page sends compressed input and receives either a Transferable result or, better, an ImageBitmap the worker has already produced.

// worker.js — owns the instance, never posts the whole heap back
import init, { decode } from './decoder.mjs';
const mod = await init();
self.onmessage = async ({ data }) => {
  const { bytes, id } = data;
  const rgba = decode(mod, bytes);                     // view into linear memory
  const bmp = await createImageBitmap(new ImageData(rgba, w, h));
  self.postMessage({ id, bmp }, [bmp]);                // transferred, not copied
};

Posting the ImageBitmap rather than the pixel array matters: a transfer moves ownership without a copy, while posting a typed array backed by linear memory forces a structured clone of every byte. The same reasoning drives compiling Wasm in a worker to free the main thread — the compile itself is expensive enough to be worth moving.

Two ways back from the worker, one of which copies Posting a typed array backed by linear memory structurally clones every byte on every frame. Producing an ImageBitmap in the worker and transferring it moves ownership with no copy, and the page can draw it directly. postMessage(pixelArray) 8.3 MB in worker 8.3 MB copy structured clone of every byte allocation pressure on both sides ≈ 4–8 ms per 1080p frame, wasted postMessage(bitmap, [bitmap]) ImageBitmap same object ownership moves, nothing is copied drawImage takes it directly ≈ 0.1 ms, independent of frame size The same rule applies to audio: build the AudioBuffer in the worker where possible, and transfer rather than clone.

Payload budgets for media modules

Media libraries are large. Knowing the rough numbers before you start saves a redesign:

Module Typical size (Brotli) Notes
JPEG decoder (libjpeg-turbo subset) 90–160 kB Comfortable for first load
PNG + zlib 60–110 kB Often already covered by the browser
AVIF / JPEG XL decoder 400 kB–1.2 MB Load lazily, cache aggressively
Opus encode + decode 250–400 kB Worth it for real-time audio
ffmpeg.wasm full build 10–25 MB Only for explicit user-initiated work

Nothing on that list should be fetched during first paint except the small ones. The pattern that works is: render the page, let the user choose a file, and only then fetch the codec while showing progress — the fetch and the user’s file picker overlap, so the perceived cost is close to zero. Pair that with caching the compiled module so the second use is instant.

Which browser API does the heavy lifting A Wasm codec rarely works alone. Each workload pairs the module with a platform API that supplies frames, samples or a place to put the result. video transcoding WebCodecs supplies decoded frames image decoding createImageBitmap and canvas for display audio processing AudioWorklet supplies 128-sample quanta filter pipelines OffscreenCanvas keeps the work off the main thread The module does the arithmetic; the platform API does the I/O and the scheduling around it. Where a native API already does the job, use it — a hardware decoder beats any module.

Gotchas and failure modes

  • RangeError: WebAssembly.Memory(): could not allocate memory — a 4K frame pipeline asking for 512 MB on a mobile device. Cap your working set, process in tiles, and treat allocation failure as a path you handle rather than an exception you log.
  • Colour space and stride surprises. A decoder that emits YUV planes with row padding will produce a sheared image if you assume tightly packed RGBA. Read the stride the library reports, never assume width * 4.
  • Detached views mid-frame. Any call that can allocate may grow memory. If you cache views, you must either guarantee no growth after startup or rebuild them after every call — see why memory.grow invalidates pointers.
  • Audio glitches from garbage collection, not from Wasm. If DSP runs in an AudioWorklet, the processing must not allocate at all. A single allocation in the audio callback is audible.
  • SharedArrayBuffer is not defined in a threaded build served without cross-origin isolation. Fix the headers as in configuring COOP/COEP headers, or ship the single-threaded build.

Verifying a media pipeline

Correctness for media is not “it looks right” — it is a checksum against a reference implementation. Decode the same asset with the native library and with the compiled module and compare bytes:

# reference, on the host
djpeg -outfile ref.raw test.jpg && sha256sum ref.raw

# under a standalone runtime, the same code path the browser will run
wasmtime run --dir=. decoder.wasm -- test.jpg out.raw && sha256sum out.raw

Identical digests mean the compile did not change behaviour. If they differ, the usual causes are floating-point differences from -ffast-math, an uninitialised scratch buffer that happened to be zero natively, or a SIMD path that diverges on edge pixels. For lossy pipelines where exact equality is not expected, compare PSNR against a stored baseline and fail the build when it regresses — a check that belongs in CI alongside the size regression gate.

When the browser already has a decoder

Before compiling anything, check what the platform gives you. Browsers decode JPEG, PNG, WebP, GIF and — increasingly — AVIF natively, on a background thread, with hardware acceleration on some platforms. A compiled decoder for those formats will lose to createImageBitmap almost every time, and it costs the user a download. The cases where your own decoder wins are narrower than they look:

  • A format the browser does not support, or supports only in some of the browsers you target. JPEG XL and certain HDR profiles are the current examples, and the gap moves every year.
  • Access to intermediate data. createImageBitmap gives you an opaque handle; if you need raw coefficients, per-plane YUV, or metadata the browser drops, you need the library.
  • Deterministic output across browsers. Native decoders differ in chroma upsampling and colour management. If two users must see byte-identical pixels — a medical or measurement context — you have to own the decode.
  • Encoding, not decoding. The platform’s encoding options are thin: quality and format, with no control over the parameters that matter for a specific corpus. A compiled encoder gives you all of them.

The pragmatic architecture uses both. Try the native path, fall back to the module, and record which path ran so you can see the real distribution in the field rather than guessing at it:

async function decode(blob, mime) {
  if (await supportsNatively(mime)) return createImageBitmap(blob);  // free, fast
  return decodeWithWasm(new Uint8Array(await blob.arrayBuffer()));   // our module
}

Streaming instead of whole-file processing

The naive pipeline reads an entire file into memory, hands it to the module, and waits. For a 40 MB video or a 200 MB WAV, that means the page holds the whole file, the module holds a copy, and the peak footprint is comfortably past what a phone will tolerate. Streaming fixes it, and the structure is not much harder.

Read the source through a ReadableStream reader, push each chunk into a fixed input region, and let the module consume chunks and emit whatever it can. The module keeps its own parser state between calls, so it needs an explicit “feed” and “drain” API rather than a single process_everything export:

const reader = file.stream().getReader();
const CHUNK = mod.exports.input_capacity();
const inView = new Uint8Array(mod.exports.memory.buffer, mod.exports.input_ptr(), CHUNK);
for (;;) {
  const { value, done } = await reader.read();
  if (done) break;
  for (let off = 0; off < value.length; off += CHUNK) {
    const slice = value.subarray(off, off + CHUNK);
    inView.set(slice);
    mod.exports.feed(slice.length);          // parse what we can, buffer the rest
    drainOutput(mod);                        // emit finished frames as they appear
  }
}
mod.exports.finish();

The peak memory becomes the chunk size plus whatever internal state the codec keeps, which for most formats is small and bounded. The user-visible effect is larger than the memory saving: the first frame appears while the rest of the file is still downloading, which turns a ten-second wait into an immediate result. This is the same argument that makes streaming instantiation the default for the module itself.

Tiling and threads for large surfaces

Very large images — scanned documents, satellite tiles, print-resolution exports — defeat a single-buffer design. A 12000 × 9000 RGBA surface is 432 MB, which no browser tab should be holding. The answer is tiling: process a band at a time, with an overlap large enough to cover the filter kernel’s radius so seams do not appear at the boundaries.

Tiling also makes threading trivial, because tiles are independent. With a threaded build and a worker pool, each worker takes a tile index, processes into its own slice of the shared output buffer, and signals completion. There is no locking because no two workers touch the same output bytes — the only synchronisation is the barrier at the end, which Atomics handles cleanly.

The speedup is close to linear up to the physical core count and then flattens, and it is worth measuring rather than assuming: on a four-core laptop a tiled blur typically lands at 3.2–3.6× the single-threaded time, with the shortfall going to memory bandwidth rather than to coordination overhead. If your kernel is bandwidth-bound rather than compute-bound — a simple copy or a channel swap — threads will buy almost nothing, and the honest move is to keep the single-threaded build and its smaller payload.

The data layout contract

Every bug that survives the first day of a media integration is a layout misunderstanding. The module and the page have to agree on five things, and none of them is inferable from the byte count alone: pixel format, channel order, stride, origin, and whether values are premultiplied.

Pixel format and channel order are the obvious pair. A C library that says “RGB” may mean packed 24-bit with no alpha, while ImageData always wants 32-bit RGBA. Swapping red and blue produces an image that looks plausible in a thumbnail and obviously wrong once someone notices skin tones, which is why this bug reliably reaches production. Write the conversion explicitly at the boundary, in the module where it is cheap, rather than in JavaScript where it is a per-pixel loop in the wrong language.

Stride is the one that bites hardest. Many decoders align each row to a 4-, 8- or 16-byte boundary, so a row of 1021 pixels occupies more bytes than 1021 * 4. Reading such a buffer as if it were tightly packed shears the image progressively down the frame — the classic diagonal-smear screenshot. Always ask the library for the stride it used and copy row by row when it differs from the packed width:

// stride-aware copy out of linear memory into a packed RGBA buffer
const packed = new Uint8ClampedArray(w * h * 4);
for (let y = 0; y < h; y++) {
  const src = new Uint8ClampedArray(memory.buffer, outPtr + y * stride, w * 4);
  packed.set(src, y * w * 4);
}

Origin — whether row zero is the top or the bottom of the image — differs between graphics APIs and image libraries, and a vertically flipped result is usually a one-line fix in the copy loop rather than a reprocessing step. Premultiplied alpha is the subtlest of the five: compositing premultiplied data as if it were straight alpha darkens edges, and the difference only shows on semi-transparent pixels, so a test image with a hard-edged alpha mask will pass while real content looks wrong.

Write these five properties down in a comment next to the export signature, and assert what you can at runtime in development builds. A module that exports frame_stride() and frame_format() alongside its pointer costs nothing and removes an entire class of afternoon-long debugging sessions. The broader version of this discipline — documenting the memory layout as part of the interface — is covered in returning structs from Wasm to JavaScript.

Frequently Asked Questions

Why is my Wasm decoder slower than the browser’s built-in one? Because the built-in one is native code with SIMD and often hardware acceleration, running off the main thread, with no boundary crossing and no download. A compiled decoder competes on capability, not on raw speed for formats the platform already handles.

How do I avoid a copy when the source is a File? You cannot avoid the first one — bytes have to get from the file into linear memory — but you can make it the only one. Read into the module’s input region directly with view.set(), process in place, and return a pointer rather than a new array. One copy per item is the floor; anything more is a bug.

Can I use ffmpeg.wasm for real-time processing? Not for anything with a latency budget. The full build is tens of megabytes and its process model assumes files, not streams. For real-time work compile the specific codec you need with a streaming API, or use WebCodecs for the decode and Wasm only for the parts the platform does not do.

Does SIMD help audio as much as it helps images? Usually more, because audio kernels are pure float loops over contiguous samples with no branching. FIR filters, mixing and resampling commonly see 2–4×. The constraint in audio is not throughput but jitter: the callback must finish within the quantum every single time.

Guides in this topic

← Back to Production Wasm: Workloads & Deployment