Sharing a Canvas Framebuffer with Wasm

This guide answers one task: let a WebAssembly module write pixels directly and get them onto a canvas every frame, with the smallest number of copies the platform allows.

Prerequisites

  • [ ] A module that renders into an RGBA byte buffer in linear memory.
  • [ ] A 2D canvas context, or a WebGL context if you take the texture route.
  • [ ] A fixed framebuffer size decided at startup, or a clean way to reallocate on resize.
  • [ ] A frame-time measurement, because the whole point is the cost of the transfer.

When software rendering is the right choice

The GPU is faster at almost everything, so software rendering needs a reason. There are several good ones. Emulators reproduce a machine’s exact pixel output and cannot express it as shaders. Raster graphics tools — pixel editors, plotters, map renderers — need precise control over every pixel and deterministic results. Ports of software renderers already exist and work. And some algorithms, such as certain scanline or flood-fill techniques, express naturally as sequential pixel writes and awkwardly as GPU passes.

There is also the portability argument: a software renderer has no context to lose, no driver quirks, no shader compilation, and works identically on every device. For a small canvas that is a real advantage.

Two routes from linear memory to the screen An ImageData backed by a copy of the module's pixels is handed to putImageData, or the pixels are uploaded as a texture and drawn by the GPU. The texture route avoids a CPU-side blit at the cost of needing a GL context. linear memory RGBA pixels from the module ImageData → putImageData one CPU copy per frame texSubImage2D → draw quad driver copy, GPU scales for free canvas what the user sees

The ImageData route

ImageData requires a Uint8ClampedArray, and its constructor will accept one that views linear memory — but putImageData then reads from that view, so the pixels the module wrote go straight to the canvas without an intermediate array of your own.

const w = 640, h = 480;
const ptr = mod.exports.framebuffer_ptr();               // allocated once at startup
const pixels = new Uint8ClampedArray(mod.exports.memory.buffer, ptr, w * h * 4);
const image = new ImageData(pixels, w, h);               // wraps the view, no copy

function frame() {
  mod.exports.render();                                  // module writes RGBA in place
  ctx.putImageData(image, 0, 0);
  requestAnimationFrame(frame);
}

This works only if the module never grows its memory — a growth detaches pixels, and the ImageData built over it becomes useless. Preallocate the framebuffer at startup and treat growth as forbidden in this design.

Note the pixel format: ImageData is RGBA with straight, non-premultiplied alpha, in that byte order, regardless of the platform’s endianness. A module that writes BGRA — as many software renderers ported from Windows do — produces an image with red and blue swapped, which is a one-line fix in the module and a long confusion if you look for it elsewhere.

The texture route

If a WebGL context is already present, uploading the framebuffer as a texture and drawing a full-screen quad is often faster, and it gets scaling and filtering for free.

gl.bindTexture(gl.TEXTURE_2D, tex);
gl.texSubImage2D(gl.TEXTURE_2D, 0, 0, 0, w, h, gl.RGBA, gl.UNSIGNED_BYTE,
                 new Uint8Array(mod.exports.memory.buffer, ptr, w * h * 4));
gl.drawArrays(gl.TRIANGLE_STRIP, 0, 4);

The advantage is that the GPU handles the scale to the canvas’s display size, so a 320 × 240 emulator framebuffer can fill a 4K display with nearest-neighbour filtering and no CPU cost. With putImageData you would have to draw at native size and scale with CSS, which gives you the browser’s filtering rather than yours.

Dirty rectangles: only move what changed

Most software renderers do not repaint the whole surface every frame. A pixel editor changes a brush stroke; a map renderer changes a panned strip; an emulator’s status bar is static. Uploading only the changed region cuts the transfer proportionally.

const { x, y, w: dw, h: dh } = mod.exports.dirty_rect();  // module reports what it touched
if (dw > 0) {
  ctx.putImageData(image, 0, 0, x, y, dw, dh);            // source-rect form
}

putImageData’s seven-argument form takes a source rectangle, so the same ImageData object serves both whole-frame and partial updates. With WebGL, texSubImage2D takes an offset and size for the same purpose, though it needs the source rows to be contiguous — which for a sub-rectangle they are not, so either upload full rows spanning the dirty region or set UNPACK_ROW_LENGTH in WebGL 2.

Double buffering

A module that renders progressively — a ray tracer refining an image, a renderer drawing scanline by scanline — will show tearing if the canvas reads the same buffer it is writing. Two framebuffers, swapped after each complete frame, fix it.

const front = new Uint8ClampedArray(mod.exports.memory.buffer, mod.exports.fb_ptr(0), w * h * 4);
const back  = new Uint8ClampedArray(mod.exports.memory.buffer, mod.exports.fb_ptr(1), w * h * 4);
const imgs = [new ImageData(front, w, h), new ImageData(back, w, h)];

function frame() {
  const done = mod.exports.render_step();          // returns which buffer is complete, or -1
  if (done >= 0) ctx.putImageData(imgs[done], 0, 0);
  requestAnimationFrame(frame);
}

The memory cost is one extra full framebuffer — 1.2 MB at 640 × 480, 8.3 MB at 1080p — which is usually worth it. The alternative, synchronising with the render thread, is more complex and offers no benefit for a single-threaded renderer.

Why a second buffer is worth 1.2 MB With one buffer the canvas can display a partially drawn frame, showing a seam where new and old content meet. With two, the display always reads a complete frame while the other is being written. single buffer new frame old frame visible seam double buffered front: displayed back: being drawn swap only when complete The swap is a pointer change inside the module — no pixels move, and the ImageData objects for both buffers are built once at startup.

Handling resize cleanly

A framebuffer has a fixed size, and the element showing it does not. Deciding how those relate is part of the design rather than an afterthought, and there are three defensible answers.

Fix the framebuffer and scale it. An emulator or a pixel-art tool has a native resolution that should not change; render at that size and let the canvas scale, with image-rendering: pixelated when you want hard pixels rather than smoothing. This is the simplest option and needs no reallocation at all.

Match the element and reallocate. A map renderer or a plotting tool should draw at the display’s real resolution, which means reallocating the framebuffer, rebuilding the views and the ImageData, and telling the module its new dimensions whenever the size settles:

const ro = new ResizeObserver(([entry]) => {
  const dpr = window.devicePixelRatio || 1;
  const w = Math.round(entry.contentRect.width * dpr);
  const h = Math.round(entry.contentRect.height * dpr);
  if (w === fbW && h === fbH) return;               // observers fire for other reasons
  clearTimeout(pending);
  pending = setTimeout(() => reallocate(w, h), 120); // debounce the drag
});

Or render at a fixed resolution below the display and upscale deliberately, which is what a renderer does when it cannot afford native resolution. That is a quality decision made once rather than a resize problem, and it pairs naturally with the texture route where the GPU does the upscale for nothing.

Whichever you choose, remember devicePixelRatio: a canvas sized in CSS pixels on a high-density display renders at a quarter of the resolution the screen can show, and the result looks soft in a way that is hard to diagnose from a screenshot taken on a standard display.

Expected output

The numbers to watch are the render time and the transfer time, separately:

resolution 640 × 480 (1.23 MB per frame)
render     4.8 ms   (module)
transfer   0.42 ms  (putImageData, full frame)
transfer   0.06 ms  (putImageData, dirty rect 180 × 120)
frame      5.4 ms

If transfer time is a large fraction of the frame, either the resolution is high enough to justify the texture route, or you are copying the pixels into a second array before constructing the ImageData — which is the most common accidental cost here.

How the pixels reach the screen The module writes into linear memory. The difference between these rows is how many copies happen between that write and the pixels appearing. getImageData + putImageData two copies per frame, plus a readback stall ImageData over memory one copy at putImageData the view itself is free OffscreenCanvas in a worker one copy, off-thread the main thread never touches the pixels at all A view over linear memory costs nothing to create; the copy is in the upload, not the view. Rebuild the view after any memory.grow, or it silently addresses a detached buffer.

Gotchas

  • Detached array after growth. The ImageData silently stops working. Preallocate and never grow.
  • Colours swapped. The module writes BGRA; ImageData wants RGBA. Fix it in the module.
  • Alpha darkening the image. ImageData uses straight alpha; a renderer producing premultiplied pixels must un-premultiply, or write 255 into the alpha channel if opacity is not used.
  • Canvas CSS size not equal to its backing size. Produces blurry scaling. Set canvas.width and canvas.height explicitly, and scale with CSS deliberately if you want it.
  • getImageData called to read back. Slow and usually unnecessary — the module already has the pixels.
  • Resizing without reallocating. The framebuffer, the views and the ImageData all need rebuilding when the dimensions change.

Performance note

At 640 × 480, putImageData of a full frame costs about 0.42 ms and the equivalent WebGL texture upload plus quad draw costs 0.28 ms — close enough that either is fine. At 1920 × 1080 the gap widens to 3.1 ms against 0.9 ms, and the texture route also scales for free rather than forcing CSS filtering. A dirty rectangle covering 10% of the surface costs roughly 10% of the full-frame time in both routes, which is usually the largest available saving.

Frequently Asked Questions

Can I use OffscreenCanvas for this? Yes, and it is the better arrangement when the renderer is in a worker: transfer the canvas control to the worker, and both the module and the drawing live off the main thread.

Does createImageBitmap help? It does if you need to scale or if you are producing frames faster than they are consumed — it hands the browser an object it can draw efficiently. For a direct per-frame blit, putImageData is more direct.

Is a 2D context slower than WebGL for this? Only at larger sizes. Below roughly 800 × 600 the difference is under a millisecond, and the 2D context avoids context loss handling entirely — a genuine simplification for a tool that must not lose the user’s work.

← Back to Graphics, Games & Simulation