Decoding Modern Image Formats with Wasm Codecs
This guide answers one task: display an image in a format the user’s browser cannot decode natively, by falling back to a compiled decoder — and do it without downloading the decoder for the majority of users whose browser handles the format perfectly well on its own.
Prerequisites
- [ ] A decoder build for the format you need:
libavif,libjxlor the Squoosh codec bundles compiled withemcc. - [ ]
emcc3.1+ if you are compiling the codec yourself. - [ ] A canvas or
ImageBitmaprendering path — the decoded pixels have to go somewhere. - [ ] Sample images in the target format, including one with alpha and one with an odd width, which is where stride bugs surface.
Detect first, download second
The decoder is between 400 kB and 1.2 MB compressed. Most of your users do not need it, and the ones who do should not pay for it until you know. Feature detection for image formats is asynchronous and cheap: decode a tiny embedded sample and see whether the browser produces a bitmap.
const PROBE = {
avif: 'data:image/avif;base64,AAAAIGZ0eXBhdmlmAAAAAGF2aWZtaWYxbWlhZk1BMUI...',
jxl: 'data:image/jxl;base64,/woIELASCAgQAFwASxLFgkWAHL0xqnCBCV0qDp901Te...',
};
const support = new Map();
async function supportsFormat(fmt) {
if (support.has(fmt)) return support.get(fmt);
const ok = await new Promise((res) => {
const img = new Image();
img.onload = () => res(img.width > 0);
img.onerror = () => res(false);
img.src = PROBE[fmt];
});
support.set(fmt, ok);
return ok;
}
Cache the answer for the session — the probe is a decode, and doing it per image is wasteful. Persisting
it to localStorage is tempting but wrong: browser updates change the answer, and a stale false costs
every future visit a megabyte of decoder it no longer needs.
Building the decoder
If you compile the codec yourself rather than using a prebuilt bundle, the export surface should be tiny: allocate, decode, describe, free. Keeping it small is what makes the JavaScript side simple and what keeps the binary from carrying an entire runtime.
emcc libavif/src/*.c aom/*.c decode_shim.c \
-O3 -msimd128 -flto \
-sMODULARIZE=1 -sEXPORT_ES6=1 -sENVIRONMENT=worker \
-sEXPORTED_FUNCTIONS='["_avif_alloc","_avif_decode","_avif_width","_avif_height","_avif_stride","_avif_free"]' \
-sALLOW_MEMORY_GROWTH=1 \
-o avif-decoder.mjs
-sENVIRONMENT=worker is worth setting deliberately: it strips the Node and shell shims from the glue,
saving several kilobytes, and it documents that this module is never meant to run on the main thread.
-flto typically removes another 10–15% for codec libraries, which are full of small functions that
inline well across translation units. The full set of levers is in
reducing Wasm bundle size with wasm-opt.
Decoding into pixels the canvas accepts
The decode itself is four calls. The subtlety is entirely in what comes back: a pointer, a width, a height and — critically — a stride, which is the number of bytes per row including any padding the decoder added for alignment.
export function decodeAvif(mod, bytes) {
const { memory, avif_alloc, avif_decode, avif_width, avif_height, avif_stride, avif_free } = mod.exports;
const inPtr = avif_alloc(bytes.length);
new Uint8Array(memory.buffer, inPtr, bytes.length).set(bytes);
const outPtr = avif_decode(inPtr, bytes.length); // may grow memory — do not cache views
if (outPtr === 0) { avif_free(inPtr); throw new Error('avif: decode failed'); }
const w = avif_width(), h = avif_height(), stride = avif_stride();
const packed = new Uint8ClampedArray(w * h * 4);
for (let y = 0; y < h; y++) {
packed.set(new Uint8ClampedArray(memory.buffer, outPtr + y * stride, w * 4), y * w * 4);
}
avif_free(inPtr);
return new ImageData(packed, w, h);
}
Note the order: every view is constructed after avif_decode returns, because the decode allocates and
may therefore have grown linear memory, detaching anything built earlier. This is the failure described
in why memory.grow invalidates pointers,
and image decoding is where it bites most often, because the output allocation is large and its size
depends on the input.
Expected output
A correct decode produces an ImageData whose dimensions match the file’s header, and whose first pixel
matches the reference decoder. Verify against the command-line tool rather than against your eyes:
avifdec sample.avif reference.png
# then in the page, after decoding:
# console.log(w, h, [...packed.slice(0, 8)])
# 1200 800 [ 34, 78, 121, 255, 35, 79, 122, 255 ]
If the first pixel matches but the image shears diagonally, the stride is wrong. If the colours look inverted in the blue channel, the decoder emitted BGRA. If the image is upside down, row zero is at the bottom in this library’s convention. All three are one-line fixes in the copy loop, and all three look like “the decoder is broken” until you check the layout contract.
Where the decoder should live
A decoder is a long-lived, expensive object, and treating it as a module-level singleton inside a worker is almost always right. Instantiating per image throws away the compiled module and reallocates the heap; instantiating on the main thread guarantees a jank spike at exactly the moment the user is looking at the page.
The structure that scales is a small pool. One worker is enough for a gallery that decodes images as they
scroll into view; two to four make sense when a user drops a folder of files and expects thumbnails
quickly. Beyond the physical core count you gain nothing, because decoding is compute-bound, and each
extra worker carries its own multi-megabyte linear memory.
// main thread: a minimal pool that keeps one instance per worker alive
const pool = Array.from({ length: Math.min(4, navigator.hardwareConcurrency || 2) },
() => new Worker('/decode-worker.js', { type: 'module' }));
let next = 0;
export function decodeInPool(bytes, fmt) {
const w = pool[next++ % pool.length];
return new Promise((resolve, reject) => {
const id = crypto.randomUUID();
const onMsg = (e) => {
if (e.data.id !== id) return;
w.removeEventListener('message', onMsg);
e.data.error ? reject(new Error(e.data.error)) : resolve(e.data.bitmap);
};
w.addEventListener('message', onMsg);
w.postMessage({ id, bytes, fmt }, [bytes.buffer]); // transfer the input, do not clone it
});
}
Transferring the input buffer rather than cloning it matters for large files: a 4 MB AVIF cloned into four workers is 16 MB of pointless copying. Transfer moves ownership, which is safe here because the main thread has no further use for the bytes once the worker owns them.
Caching decoded results
Decoding is expensive enough that repeating it is a visible cost, and the usual page behaviour — scroll away, scroll back — repeats it constantly. Two caches are worth having, and they solve different problems.
The first is a bitmap cache keyed by URL, holding ImageBitmap objects for the images currently near the
viewport. ImageBitmap lives in GPU-adjacent memory and draws essentially for free, so a small cache of
twenty or thirty covers typical scrolling. Call close() on entries you evict; without it, the memory
is reclaimed only when the garbage collector gets around to it, which on a fast scroll is too late.
The second is a compiled-module cache so that the decoder itself is instantiated once per session and compiled once per device, as described in caching compiled Wasm modules in IndexedDB. Between them, the second visit to a page that uses a compiled decoder costs a few milliseconds of instantiation and no download at all — which is what makes the fallback path acceptable in production rather than merely possible.
Do not cache the raw decoded pixel arrays. A 12-megapixel surface is 48 MB, and a cache of five of them
is larger than most tabs should ever be; keep the compact ImageBitmap and let the browser manage where
the pixels actually sit.
Gotchas
RuntimeError: memory access out of boundsimmediately after decode. The output pointer was read with a stale view, or the length was computed withwidth * 4on a padded buffer.- A 10-second freeze on a large AVIF. AV1 intra decoding is genuinely expensive. Run it in a worker; on the main thread it will blow every frame budget you have.
- Alpha looks wrong on composited images. Some decoders emit premultiplied alpha.
ImageDataexpects straight alpha, so un-premultiply in the module before returning, where it is a cheap loop. - Progressive or animated files return only the first frame. Animation is a separate API in most of these libraries; the shim you compile has to expose it explicitly or you will silently drop frames.
- The module works locally and 404s in production. The
.wasmnext to the glue is fetched relative to the glue’s URL. Bundlers frequently move one and not the other — see bundling Wasm ESM with Vite.
Performance note
For a 12-megapixel AVIF, a compiled decoder on a modern laptop takes roughly 350–600 ms single-threaded, against 90–150 ms for the browser’s native decoder, which uses hardware paths you cannot reach. SIMD closes perhaps a third of that gap and threads close more, but the compiled path will not win on speed — it wins on availability. Budget accordingly: show a placeholder, decode in a worker, and swap the image in when it arrives rather than blocking the render on it.
Frequently Asked Questions
Is it worth shipping a JPEG XL decoder today? Only if your content genuinely benefits and you control the corpus — photography archives and print workflows are the usual cases. For general web imagery, AVIF and WebP have broad native support and a compiled fallback costs more than the format saves.
Can I decode straight into a WebGL texture?
Yes, and it avoids one copy. Decode into linear memory, then pass the typed-array view directly to
texSubImage2D. The driver copies from there, so the intermediate ImageData is unnecessary — see
rendering with WebGL from a Wasm module.
How do I keep the decoder from being downloaded by crawlers and bots? By never loading it eagerly. Because the fetch happens inside a capability-gated code path that runs only after a failed native decode, anything that does not attempt to display the image never requests it.
Related
- Building a Wasm image filter pipeline — what to do with the pixels once you have them.
- Resizing images off the main thread — the worker pattern in full.
- Avoiding copies when passing image buffers — the zero-copy rules this page applies.
← Back to Media Processing & Codecs in Wasm