Passing Canvas Pixels to Wasm Without Extra Copies

This page answers one task: a WebAssembly module filters or analyses pixels from a <canvas> — image editing, video effects, computer vision — and the time spent copying pixel data between the canvas, JavaScript arrays and Wasm memory rivals the filter itself. You want the shortest path from canvas to Wasm and back.

Prerequisites

  • [ ] A canvas 2D context (or OffscreenCanvas) and a Wasm module that processes RGBA pixels.
  • [ ] Access to the module’s memory and an allocation function (wasm-bindgen, Emscripten or raw exports).
  • [ ] Familiarity with typed arrays and ImageData.

Where the copies come from

A naive pipeline copies pixels four times per frame: getImageData copies from the canvas into a new ImageData buffer; passing that Uint8ClampedArray to a wasm-bindgen function as &[u8] copies it into linear memory; returning a Vec<u8> copies the result out into a new JavaScript array; and putImageData copies it back to the canvas. For a 4K frame (33 MB of RGBA), each copy costs several milliseconds.

Two of those copies are unavoidable with the 2D canvas API — reading from and writing to the canvas always copy, because the canvas’s backing store is owned by the browser and may live on the GPU. The other two can be eliminated by having the canvas API read from and write to linear memory directly: copy getImageData’s result into a buffer the module already owns, and build the output ImageData as a view over linear memory so putImageData reads straight from Wasm memory.

Naive versus direct pixel path The naive path copies pixels from canvas to ImageData, into Wasm memory, out to a new array and back to the canvas, four copies per frame. The direct path copies from canvas into a Wasm-owned buffer, processes in place, and calls putImageData on an ImageData that views Wasm memory, two copies per frame. naive (4 copies) getImageData → new array array → linear memory Vec out → new array putImageData → canvas slow for large frames direct (2 copies) getImageData → set into Wasm buffer filter in place ImageData view over memory putImageData from memory minimum for 2D canvas

Step 1 — allocate a pixel buffer inside the module

Have the module own a buffer sized for the frame and expose its address:

use wasm_bindgen::prelude::*;

#[wasm_bindgen]
pub struct Frame { pixels: Vec<u8>, width: u32, height: u32 }

#[wasm_bindgen]
impl Frame {
    #[wasm_bindgen(constructor)]
    pub fn new(width: u32, height: u32) -> Frame {
        Frame { pixels: vec![0; (width * height * 4) as usize], width, height }
    }
    pub fn ptr(&self) -> *const u8 { self.pixels.as_ptr() }
    pub fn len(&self) -> usize { self.pixels.len() }
    pub fn grayscale(&mut self) {
        for px in self.pixels.chunks_exact_mut(4) {
            let y = ((px[0] as u32 * 77 + px[1] as u32 * 150 + px[2] as u32 * 29) >> 8) as u8;
            px[0] = y; px[1] = y; px[2] = y;
        }
    }
}

Allocate once per frame size and reuse it for every frame; reallocating per frame churns the allocator and may grow memory, which detaches views.

Step 2 — copy canvas pixels into the buffer

import init, { Frame } from "./pkg/filters.js";
const { memory } = await init();

const frame = new Frame(canvas.width, canvas.height);
function wasmPixels() {
  return new Uint8ClampedArray(memory.buffer, frame.ptr(), frame.len());   // fresh view each time
}

const src = ctx.getImageData(0, 0, canvas.width, canvas.height);   // copy 1 (unavoidable)
wasmPixels().set(src.data);                                          // into Wasm memory

set on a typed array is a fast memory copy. This replaces the copy that wasm-bindgen would make when passing &[u8], and leaves the data where the module can process it in place.

Step 3 — process in place and write back from a view

frame.grayscale();                                                   // no copies
const out = new ImageData(wasmPixels(), canvas.width, canvas.height); // a view, not a copy
ctx.putImageData(out, 0, 0);                                         // copy 2 (unavoidable)

new ImageData(uint8ClampedArray, width, height) wraps the existing array without copying when the array’s length matches; because the array views linear memory, putImageData reads directly from it. Create the ImageData right before use: if memory grows in between, the view detaches and the ImageData becomes unusable.

A frame's pixels from canvas through Wasm and back getImageData copies pixels out of the canvas. A typed-array set copies them into the module's frame buffer in linear memory. The Wasm filter processes them in place. An ImageData is created as a view over that buffer, and putImageData copies it back to the canvas. getImageData copy out of canvas view.set(data) into Wasm buffer frame. grayscale() in place new ImageData(view) no copy putImageData copy into canvas

Step 4 — move the work to a worker with OffscreenCanvas

For video or continuous effects, run the whole loop in a worker so the main thread stays free. Transfer the canvas with transferControlToOffscreen(), and use the same direct path in the worker with the OffscreenCanvas 2D context. Frames from a <video> can reach the worker as VideoFrame objects (WebCodecs) or ImageBitmaps and be drawn into the offscreen canvas, then read back with getImageData. For video specifically, VideoFrame.copyTo can copy pixel data directly into a buffer — including a view over Wasm memory — skipping the canvas read entirely; see Media Processing & Codecs.

Step 5 — measure the copies

Time each stage with performance.now() for your frame size. On a typical laptop, copying 33 MB takes a few milliseconds per copy; the filter itself may take less. If copies dominate, consider smaller working resolutions for previews, processing only changed regions (getImageData with a sub-rectangle), or moving to WebGL/WebGPU, where pixels stay on the GPU and only results come back.

Detached views and memory growth

Every view over memory.buffer — the Uint8ClampedArray, the ImageData built on it — becomes detached if linear memory grows. A detached view has length 0, and putImageData with a detached ImageData throws or draws nothing. Avoid growth during the frame loop by allocating frame buffers up front, and always create views fresh after any call that might allocate. In threaded builds with shared memory, views stay valid through growth but may not see the new size until recreated; and ImageData cannot be constructed over a SharedArrayBuffer-backed array in some browsers, so copy into a non-shared buffer for putImageData in that case.

Alpha and colour formats

Canvas ImageData is RGBA with unpremultiplied alpha, 8 bits per channel, in sRGB unless the context was created with another colorSpace. Wasm filters must use the same layout, or convert. Some image libraries expect BGRA or premultiplied alpha; converting in Wasm is cheap compared with an extra copy, but it must be done consistently, or colours shift subtly at semi-transparent edges.

Processing regions instead of whole frames

Many interactions touch a small part of the image: a brush stroke, a selection, a cursor-following magnifier. Copying the whole frame for a 64×64 change wastes nearly all the work. getImageData(x, y, w, h) and putImageData(data, x, y) operate on sub-rectangles, so read only the dirty region into a small Wasm buffer, process it, and write it back. If the filter needs neighbouring pixels — a blur radius, an edge detector — read a slightly larger rectangle (the region plus the filter radius on each side) and write back only the inner part. Track dirty rectangles per frame and merge overlapping ones to keep the number of calls small. For editors, keeping the full-resolution document inside the module and drawing only the visible, changed area to the canvas avoids most reads from the canvas altogether: the canvas becomes an output surface, and the module holds the source of truth.

Keeping the document in Wasm memory

That last point generalises. If the module holds the image — layers, history, the full-resolution pixels — then the expensive canvas read happens once, when an image is loaded, and every later operation reads and writes Wasm memory directly. Only display needs putImageData, and only for the visible area at screen resolution. Undo history can be stored compactly inside the module (for example as compressed tiles), rather than as JavaScript copies of whole frames, which also keeps the JavaScript heap small and garbage collection quiet during editing.

Expected output

For a 3840×2160 frame, the naive path takes 34 ms per frame (four copies plus filter) and the direct path 19 ms (two copies plus filter); the module allocates its frame buffer once; no detached-view errors occur during a 10-minute run; and the worker version keeps the main thread’s frame time under 2 ms.

Gotchas

  • Passing ImageData.data as &[u8] to wasm-bindgen. It copies. Set into a Wasm-owned buffer instead.
  • Returning Vec<u8> results. Another copy. Process in place.
  • Views kept across allocations. They detach on growth. Recreate per use.
  • Allocating per frame. Churn and growth. Allocate once per size.
  • Assuming BGRA or premultiplied alpha. Canvas uses straight RGBA.
  • Whole-frame copies for small edits. Read and write only the dirty region.

Performance note

Removing the two avoidable copies saved 15 ms per 4K frame — more than the grayscale filter itself, which took 4 ms.

Time per 4K frame by pipeline Milliseconds per 3840 by 2160 frame for a grayscale filter using the naive four-copy pipeline and the direct two-copy pipeline with an ImageData view over Wasm memory. ms per frame naive (4 copies) 34 ms direct (2 copies) 19 ms filter alone 4 ms

Frequently Asked Questions

Can Wasm read the canvas without any copy? Not with the 2D API; the browser owns the backing store. WebGL/WebGPU keep pixels on the GPU instead.

Does getImageData with willReadFrequently help? It tells the browser to keep a CPU-side copy, making repeated reads faster.

Is ImageBitmap faster? For drawing, yes; for reading pixels you still need getImageData or VideoFrame.copyTo.

What about premultiplied alpha? ImageData is unpremultiplied; convert if your filter needs premultiplied values.

How do I avoid copying the whole frame for a small edit? Read and write only the dirty sub-rectangle, padded by the filter radius, with getImageData and putImageData offsets.

Should an editor keep its image in Wasm or in the canvas? In Wasm — read the canvas once at load, then draw only visible changed areas back, so most operations never copy from the canvas.

← Back to Zero-Copy Data Transfer Patterns