Porting a JavaScript Image Algorithm to Wasm

This page answers one task: a per-pixel JavaScript image operation — here a box blur over an RGBA canvas — is too slow for interactive use, and you want to port it to WebAssembly in measured steps, so you know which change produced which part of the speed-up.

Prerequisites

  • [ ] The JavaScript implementation and a test image of realistic size (12 megapixels used here).
  • [ ] Rust with wasm-bindgen (or C with Emscripten), and SIMD enabled in the build for the last step.
  • [ ] A benchmark harness that warms up and repeats, as in building a reproducible Wasm benchmark harness.

Why image code is a good porting target

Image algorithms are the classic WebAssembly win: they loop over millions of bytes, do integer arithmetic, have no object allocation in the inner loop, and map naturally onto SIMD. They are also a good teaching case because the speed-up comes from several distinct changes, and it is easy to attribute gains incorrectly. A port that “made it 8× faster” may owe 1.3× to the language, 2× to a better algorithm that would also help JavaScript, and 3× to SIMD. Measuring each step separately tells you what to keep, what to backport to JavaScript, and what the browser fallback will lose.

The example is a box blur with radius r: each output pixel is the average of a (2r+1)² neighbourhood. The JavaScript version is the naive four-nested-loop implementation, run on a Uint8ClampedArray from getImageData.

The porting steps and the time after each Milliseconds for a radius-8 box blur on a 12-megapixel image after each step: naive JavaScript, a straight Wasm port, separable passes, integer sliding windows, and SIMD. naive JavaScript 3,900 ms straight Wasm port 2,950 ms separable passes 420 ms sliding window 150 ms SIMD 48 ms

Step 1 — a straight port

Translate the JavaScript loops literally into Rust, taking and returning byte slices:

#[wasm_bindgen]
pub fn box_blur_naive(src: &[u8], w: usize, h: usize, r: usize) -> Vec<u8> {
    let mut dst = vec![0u8; src.len()];
    for y in 0..h { for x in 0..w { for c in 0..4 {
        let (mut sum, mut n) = (0u32, 0u32);
        for dy in y.saturating_sub(r)..=(y + r).min(h - 1) {
            for dx in x.saturating_sub(r)..=(x + r).min(w - 1) {
                sum += src[(dy * w + dx) * 4 + c] as u32; n += 1;
            }
        }
        dst[(y * w + x) * 4 + c] = (sum / n) as u8;
    }}}
    dst
}

On the 12-megapixel test image with r = 8, the JavaScript took 3,900 ms and this port 2,950 ms — a 1.3× gain. That is typical for a straight port of well-typed JavaScript: the engine’s JIT already did a good job, and the language change alone is not where big wins come from.

Step 2 — keep pixels in linear memory

&[u8] parameters and Vec<u8> results make wasm-bindgen copy 48 MB in and 48 MB out on every call. For repeated operations — a blur slider — allocate input and output buffers in the module once and let JavaScript write pixels into them directly:

const inPtr = wasm.alloc(w * h * 4), outPtr = wasm.alloc(w * h * 4);
new Uint8ClampedArray(wasm.memory.buffer, inPtr, w * h * 4).set(ctx.getImageData(0, 0, w, h).data);
wasm.box_blur_into(inPtr, outPtr, w, h, radius);
ctx.putImageData(new ImageData(new Uint8ClampedArray(wasm.memory.buffer, outPtr, w * h * 4), w, h), 0, 0);

This saved about 140 ms per call — small next to the algorithm here, but it becomes significant once the algorithm is fast. The technique is detailed in avoiding copies when passing image buffers.

Step 3 — fix the algorithm

The naive blur does (2r+1)² reads per pixel. A box blur is separable: blurring horizontally and then vertically with a 2r+1 window gives the same result with 2 × (2r+1) reads per pixel. With r = 8 that is 34 reads instead of 289. Implementing the two passes dropped the time to 420 ms — a 7× gain from the algorithm alone. Note that this change would help the JavaScript version equally; doing it in both is the honest comparison, and backporting it is a cheap improvement for browsers that use the JavaScript fallback.

Step 4 — sliding windows and integer arithmetic

Each pass can keep a running sum: add the pixel entering the window, subtract the one leaving it, so the cost per pixel no longer depends on the radius. Use integer division only once per output, or multiply by a precomputed reciprocal in fixed point:

let inv = (1u32 << 16) / (2 * r as u32 + 1);              // fixed-point reciprocal
// inside the horizontal pass, per channel:
sum += row[(x + r + 1).min(w - 1) * 4 + c] as u32;
sum -= row[x.saturating_sub(r) * 4 + c] as u32;
out[x * 4 + c] = ((sum * inv) >> 16) as u8;

This dropped the time to 150 ms and made it independent of the radius. Again, a JavaScript implementation benefits from the same idea, though less, because JavaScript’s numbers are doubles and integer tricks are not always preserved by the JIT.

What each change contributed The language change alone gave a 1.3 times gain. Avoiding copies saved a fixed cost per call. Separable passes and sliding windows were algorithmic changes worth about 20 times and apply to JavaScript too. SIMD was a Wasm-specific gain of about 3 times on top. language and copies straight port 1.3× no copies saves ~140 ms/call Wasm-specific modest on their own algorithm separable passes 7× sliding window 2.8× also helps JavaScript the biggest win SIMD 16 bytes per instruction ~3× on top Wasm-only advantage final multiplier

Step 5 — add SIMD and move to a worker

With the algorithm in good shape, SIMD processes all four channels of four pixels per instruction. Using core::arch::wasm32 intrinsics — u16x8 adds and subtracts on widened bytes for the running sums, and a multiply-high for the reciprocal — the vertical pass in particular vectorises well because it operates on whole rows. Building with -C target-feature=+simd128 brought the time to 48 ms. Writing those intrinsics is covered in writing v128 SIMD intrinsics in Rust. Finally, run the blur in a worker so the slider stays responsive even on slow devices, and ship a non-SIMD build for browsers without SIMD support.

Keeping the result correct

Every optimisation step is a chance to change the output. The separable and sliding-window versions round differently from the naive one unless care is taken, and edge handling — clamping versus mirroring at the image border — must match exactly. Keep the naive JavaScript as the reference and compare every optimised version against it on a set of test images, including tiny ones (1×1, 2×3) and ones whose sizes are not multiples of the SIMD width. Accept a maximum difference of one intensity level only if you have decided to, and test it explicitly. The general approach is described in keeping JavaScript and Wasm results identical.

What this means for planning a port

The lesson generalises beyond blurs. When a port is planned, list the candidate improvements and classify them: language and runtime gains (usually modest for well-typed code), data-movement gains (significant for repeated calls on large buffers), algorithmic gains (often the largest, and available in JavaScript too), and Wasm-specific capabilities such as SIMD, 64-bit integers and manual memory layout. Do the algorithmic work first — sometimes in JavaScript, where it may already be enough — and use WebAssembly for the parts that only it can accelerate. That ordering avoids crediting the port with gains that came from rethinking the algorithm, and it keeps the JavaScript fallback reasonably fast for browsers that need it.

Expected output

The blur slider updates a 12-megapixel image in about 50 ms on a laptop and under 200 ms on a mid-range phone, from 3.9 s originally; every optimised version matches the reference within one intensity level on the test images; and the improved separable algorithm also runs in the JavaScript fallback.

Gotchas

  • Crediting the language for algorithmic gains. Measure each change separately.
  • Copying the image on every call. Keep buffers in linear memory for repeated operations.
  • Edge handling drifting between versions. Test tiny images and odd sizes.
  • SIMD tails. Widths that are not multiples of the vector size need scalar remainders.
  • Forgetting the non-SIMD build. Older browsers need it.
  • Benchmarking a single run. Image timings vary with caches and thermal state. Repeat and report the median.

Performance note

From 3,900 ms to 48 ms overall — an 81× improvement — of which roughly 20× came from algorithmic changes that also apply to JavaScript, 3× from SIMD, and 1.3× from the language change itself. The backported JavaScript with the same algorithm ran in 210 ms.

Final comparison, JavaScript and Wasm with the same algorithm Milliseconds for the radius-8 box blur on a 12-megapixel image with the original naive JavaScript, JavaScript using the improved sliding-window algorithm, Wasm with the same algorithm, and Wasm with SIMD. ms per blur naive JavaScript 3,900 ms JS, sliding window 210 ms Wasm, sliding window 150 ms Wasm + SIMD 48 ms

Frequently Asked Questions

Should I port the naive algorithm first? Yes, briefly, as a correctness baseline — then optimise and measure each step.

Would WebGL or WebGPU be faster? For blurs at interactive rates on large images, often yes. Wasm wins when you need exact results, CPU-only environments or easier integration.

Does Uint8ClampedArray clamping matter? It clamps on assignment in JavaScript; in Rust use saturating arithmetic or explicit clamps to match.

How do I handle very large images? Process in tiles or strips to bound memory, especially on phones.

Can Emscripten do the same? Yes — the steps are identical in C, with wasm_simd128.h for the SIMD stage.

Why did copying matter so little here? The naive algorithm was slow enough to hide it. Once the algorithm took 48 ms, a 140 ms copy would have dominated, so removing it was essential.

← Back to Porting JavaScript Hot Paths to Wasm