Porting a JavaScript Image Algorithm to Wasm
This page answers one task: a per-pixel JavaScript image operation — here a box blur over an RGBA canvas — is too slow for interactive use, and you want to port it to WebAssembly in measured steps, so you know which change produced which part of the speed-up.
Prerequisites
- [ ] The JavaScript implementation and a test image of realistic size (12 megapixels used here).
- [ ] Rust with wasm-bindgen (or C with Emscripten), and SIMD enabled in the build for the last step.
- [ ] A benchmark harness that warms up and repeats, as in building a reproducible Wasm benchmark harness.
Why image code is a good porting target
Image algorithms are the classic WebAssembly win: they loop over millions of bytes, do integer arithmetic, have no object allocation in the inner loop, and map naturally onto SIMD. They are also a good teaching case because the speed-up comes from several distinct changes, and it is easy to attribute gains incorrectly. A port that “made it 8× faster” may owe 1.3× to the language, 2× to a better algorithm that would also help JavaScript, and 3× to SIMD. Measuring each step separately tells you what to keep, what to backport to JavaScript, and what the browser fallback will lose.
The example is a box blur with radius r: each output pixel is the average of a (2r+1)² neighbourhood. The JavaScript version is the naive
four-nested-loop implementation, run on a Uint8ClampedArray from getImageData.
Step 1 — a straight port
Translate the JavaScript loops literally into Rust, taking and returning byte slices:
#[wasm_bindgen]
pub fn box_blur_naive(src: &[u8], w: usize, h: usize, r: usize) -> Vec<u8> {
let mut dst = vec![0u8; src.len()];
for y in 0..h { for x in 0..w { for c in 0..4 {
let (mut sum, mut n) = (0u32, 0u32);
for dy in y.saturating_sub(r)..=(y + r).min(h - 1) {
for dx in x.saturating_sub(r)..=(x + r).min(w - 1) {
sum += src[(dy * w + dx) * 4 + c] as u32; n += 1;
}
}
dst[(y * w + x) * 4 + c] = (sum / n) as u8;
}}}
dst
}
On the 12-megapixel test image with r = 8, the JavaScript took 3,900 ms and this port 2,950 ms — a 1.3× gain. That is typical for a straight port of
well-typed JavaScript: the engine’s JIT already did a good job, and the language change alone is not where big wins come from.
Step 2 — keep pixels in linear memory
&[u8] parameters and Vec<u8> results make wasm-bindgen copy 48 MB in and 48 MB out on every call. For repeated operations — a blur slider — allocate
input and output buffers in the module once and let JavaScript write pixels into them directly:
const inPtr = wasm.alloc(w * h * 4), outPtr = wasm.alloc(w * h * 4);
new Uint8ClampedArray(wasm.memory.buffer, inPtr, w * h * 4).set(ctx.getImageData(0, 0, w, h).data);
wasm.box_blur_into(inPtr, outPtr, w, h, radius);
ctx.putImageData(new ImageData(new Uint8ClampedArray(wasm.memory.buffer, outPtr, w * h * 4), w, h), 0, 0);
This saved about 140 ms per call — small next to the algorithm here, but it becomes significant once the algorithm is fast. The technique is detailed in avoiding copies when passing image buffers.
Step 3 — fix the algorithm
The naive blur does (2r+1)² reads per pixel. A box blur is separable: blurring horizontally and then vertically with a 2r+1 window gives the same
result with 2 × (2r+1) reads per pixel. With r = 8 that is 34 reads instead of 289. Implementing the two passes dropped the time to 420 ms — a 7×
gain from the algorithm alone. Note that this change would help the JavaScript version equally; doing it in both is the honest comparison, and
backporting it is a cheap improvement for browsers that use the JavaScript fallback.
Step 4 — sliding windows and integer arithmetic
Each pass can keep a running sum: add the pixel entering the window, subtract the one leaving it, so the cost per pixel no longer depends on the radius. Use integer division only once per output, or multiply by a precomputed reciprocal in fixed point:
let inv = (1u32 << 16) / (2 * r as u32 + 1); // fixed-point reciprocal
// inside the horizontal pass, per channel:
sum += row[(x + r + 1).min(w - 1) * 4 + c] as u32;
sum -= row[x.saturating_sub(r) * 4 + c] as u32;
out[x * 4 + c] = ((sum * inv) >> 16) as u8;
This dropped the time to 150 ms and made it independent of the radius. Again, a JavaScript implementation benefits from the same idea, though less, because JavaScript’s numbers are doubles and integer tricks are not always preserved by the JIT.
Step 5 — add SIMD and move to a worker
With the algorithm in good shape, SIMD processes all four channels of four pixels per instruction. Using core::arch::wasm32 intrinsics — u16x8 adds
and subtracts on widened bytes for the running sums, and a multiply-high for the reciprocal — the vertical pass in particular vectorises well because it
operates on whole rows. Building with -C target-feature=+simd128 brought the time to 48 ms. Writing those intrinsics is covered in
writing v128 SIMD intrinsics in Rust.
Finally, run the blur in a worker so the slider stays responsive even on slow devices, and ship a non-SIMD build for browsers without SIMD support.
Keeping the result correct
Every optimisation step is a chance to change the output. The separable and sliding-window versions round differently from the naive one unless care is taken, and edge handling — clamping versus mirroring at the image border — must match exactly. Keep the naive JavaScript as the reference and compare every optimised version against it on a set of test images, including tiny ones (1×1, 2×3) and ones whose sizes are not multiples of the SIMD width. Accept a maximum difference of one intensity level only if you have decided to, and test it explicitly. The general approach is described in keeping JavaScript and Wasm results identical.
What this means for planning a port
The lesson generalises beyond blurs. When a port is planned, list the candidate improvements and classify them: language and runtime gains (usually modest for well-typed code), data-movement gains (significant for repeated calls on large buffers), algorithmic gains (often the largest, and available in JavaScript too), and Wasm-specific capabilities such as SIMD, 64-bit integers and manual memory layout. Do the algorithmic work first — sometimes in JavaScript, where it may already be enough — and use WebAssembly for the parts that only it can accelerate. That ordering avoids crediting the port with gains that came from rethinking the algorithm, and it keeps the JavaScript fallback reasonably fast for browsers that need it.
Expected output
The blur slider updates a 12-megapixel image in about 50 ms on a laptop and under 200 ms on a mid-range phone, from 3.9 s originally; every optimised version matches the reference within one intensity level on the test images; and the improved separable algorithm also runs in the JavaScript fallback.
Gotchas
- Crediting the language for algorithmic gains. Measure each change separately.
- Copying the image on every call. Keep buffers in linear memory for repeated operations.
- Edge handling drifting between versions. Test tiny images and odd sizes.
- SIMD tails. Widths that are not multiples of the vector size need scalar remainders.
- Forgetting the non-SIMD build. Older browsers need it.
- Benchmarking a single run. Image timings vary with caches and thermal state. Repeat and report the median.
Performance note
From 3,900 ms to 48 ms overall — an 81× improvement — of which roughly 20× came from algorithmic changes that also apply to JavaScript, 3× from SIMD, and 1.3× from the language change itself. The backported JavaScript with the same algorithm ran in 210 ms.
Frequently Asked Questions
Should I port the naive algorithm first? Yes, briefly, as a correctness baseline — then optimise and measure each step.
Would WebGL or WebGPU be faster? For blurs at interactive rates on large images, often yes. Wasm wins when you need exact results, CPU-only environments or easier integration.
Does Uint8ClampedArray clamping matter?
It clamps on assignment in JavaScript; in Rust use saturating arithmetic or explicit clamps to match.
How do I handle very large images? Process in tiles or strips to bound memory, especially on phones.
Can Emscripten do the same?
Yes — the steps are identical in C, with wasm_simd128.h for the SIMD stage.
Why did copying matter so little here? The naive algorithm was slow enough to hide it. Once the algorithm took 48 ms, a 140 ms copy would have dominated, so removing it was essential.
Related
- Vectorizing image convolution with SIMD — SIMD for general kernels.
- Building a Wasm image filter pipeline — chaining filters.
- Estimating the speed-up before porting — predicting outcomes.
- Passing canvas pixels to Wasm without extra copies — the canvas side.
← Back to Porting JavaScript Hot Paths to Wasm