Rendering with WebGL from a Wasm Module

This guide answers one task: render geometry whose vertex data lives inside a WebAssembly module’s linear memory, at frame rate, without copying that data into a JavaScript array on the way to the GPU.

Prerequisites

  • [ ] A module that exposes pointers to its vertex data and never grows memory after startup.
  • [ ] A WebGL 2 context. WebGL 1 works with more extensions and more caveats.
  • [ ] Familiarity with buffers, shaders and attribute layout; this page covers the interop, not GL.
  • [ ] A frame-time measurement, so claims about improvement can be checked.

The core trick: views are valid GL sources

WebGL’s upload functions accept an ArrayBufferView. A typed array constructed over memory.buffer is such a view, so the upload reads the module’s memory directly — the driver copies from there, and no intermediate JavaScript array exists.

const { memory, positions_ptr, vertex_count } = mod.exports;

const positions = new Float32Array(memory.buffer, positions_ptr(), vertex_count() * 3);
gl.bindBuffer(gl.ARRAY_BUFFER, vbo);
gl.bufferSubData(gl.ARRAY_BUFFER, 0, positions);

That is the whole integration. Everything else in this page is about doing it without introducing a stall or an invalidated view.

One copy, made by the driver The typed array is a window onto linear memory rather than a container. The only copy is the one the graphics driver performs when filling the GPU buffer, and it reads the module's bytes directly. linear memory positions written by the module window Float32Array view offset + length, no bytes owned driver copy GPU buffer persistent, reused each frame Building a plain JavaScript array in the middle adds a full copy plus allocation per frame — the thing this arrangement exists to avoid. Constructing the view itself is free; it stores an offset and a length and nothing else.

Allocate GPU buffers once

Call bufferData at startup to size the buffer, and bufferSubData every frame to fill it. bufferData reallocates GPU storage; doing it per frame forces the driver to orphan the old allocation and can stall the pipeline while in-flight draws finish with it.

// startup: reserve for the maximum you will ever draw
gl.bindBuffer(gl.ARRAY_BUFFER, vbo);
gl.bufferData(gl.ARRAY_BUFFER, MAX_VERTS * 3 * 4, gl.DYNAMIC_DRAW);

// each frame: fill only the part in use
gl.bufferSubData(gl.ARRAY_BUFFER, 0, positions.subarray(0, liveCount * 3));

DYNAMIC_DRAW tells the driver the contents change often, which affects where it places the allocation. Using STATIC_DRAW for data you rewrite every frame produces correct output and measurably worse performance on some drivers.

Keep views valid

A view over memory.buffer becomes detached the moment the module grows its memory. In a frame loop that means garbage output or a thrown error, appearing at whatever moment the module happened to allocate.

Two approaches work, and they are not equally good. The robust one is a module that preallocates everything at startup and never grows, in which case views built once remain valid forever and the loop does no work at all to maintain them. The fallback, for modules you do not control, is rebuilding views each frame:

let cached = null, cachedBuffer = null;
function positionsView(mod) {
  const buf = mod.exports.memory.buffer;
  if (buf !== cachedBuffer) {                       // cheap identity check
    cachedBuffer = buf;
    cached = new Float32Array(buf, mod.exports.positions_ptr(), MAX_VERTS * 3);
  }
  return cached;
}

The identity comparison costs nothing and rebuilds only when the buffer actually changed, which is the right compromise when you cannot guarantee a fixed memory.

Instancing: thousands of objects, one draw call

Drawing a thousand sprites with a thousand draw calls spends the frame in call overhead. Instancing issues one call and supplies per-instance attributes from a buffer — which the module writes, in exactly the layout the GPU wants.

// per-instance data: x, y, scale, rotation — written by the module into one contiguous block
const inst = new Float32Array(memory.buffer, mod.exports.instances_ptr(), count * 4);
gl.bindBuffer(gl.ARRAY_BUFFER, instanceVbo);
gl.bufferSubData(gl.ARRAY_BUFFER, 0, inst);

gl.vertexAttribPointer(aOffset, 4, gl.FLOAT, false, 16, 0);
gl.vertexAttribDivisor(aOffset, 1);                 // advance once per instance
gl.drawArraysInstanced(gl.TRIANGLE_STRIP, 0, 4, count);

The module’s job becomes filling a flat array of per-instance records, which is the same struct-of-arrays discipline that makes the simulation fast in the first place. Ten thousand instances at four floats each is 160 kB uploaded per frame — trivial for the bus, and a single draw call.

Expected output

A working integration draws the right geometry and shows a flat frame time. Instrument both halves separately so you know which one moved:

sim   p50 4.1 ms  p95 5.0 ms
gl    p50 1.8 ms  p95 2.4 ms
frame p50 6.3 ms  p95 7.9 ms   (60 Hz, 16.7 ms budget)
draws 1 instanced call, 9,800 instances, 156 kB uploaded

If the upload time scales with vertex count more steeply than the byte count suggests, you are probably copying through an intermediate array somewhere — a Array.from, a spread, or a slice() that was added to “be safe”.

Ten thousand objects, two ways Issuing one draw call per object spends the frame in call overhead regardless of how little each object draws. One instanced call uploads a flat per-instance array and lets the GPU iterate. one call per object 10,000 draw calls uniform updates between each state validation dominates ≈ 38 ms — misses the frame one instanced call 1 draw call 156 kB uploaded from linear memory GPU iterates the instance buffer ≈ 1.8 ms — comfortable The module's work is identical in both; only the shape of what it writes and how it is submitted changed.

Structuring the per-frame call sequence

The order of operations inside a frame affects both correctness and cost, and a stable structure is worth fixing early.

Step the simulation first, so the data the frame draws is the data the frame computed — drawing before stepping introduces a frame of latency that is easy to add and hard to notice. Then update only the buffers whose contents changed: static geometry uploaded once at startup does not need touching, and a frame that re-uploads everything every time wastes bandwidth proportional to how much of the scene is static, which in most applications is most of it.

Bind state in an order that minimises changes. Group draws by shader program first, then by texture, then by buffer, because each of those bindings has a different cost and the expensive ones should change least often. With instancing this usually collapses to a handful of binds per frame regardless of object count.

function frame(now) {
  stepSimulation(now);                       // module updates its arrays
  gl.bindVertexArray(vao);                   // one bind, set up at startup
  gl.bufferSubData(gl.ARRAY_BUFFER, 0, instancesView.subarray(0, live * 4));
  gl.useProgram(prog);
  gl.drawArraysInstanced(gl.TRIANGLE_STRIP, 0, 4, live);
  requestAnimationFrame(frame);
}

Vertex array objects are worth using for exactly this reason: they capture the attribute configuration once so the per-frame sequence does not re-specify pointers and divisors. In WebGL 2 they are core, and they remove a dozen calls per frame from the loop above.

Textures from linear memory

Texture uploads take a view the same way buffers do, which makes a module that generates or decodes pixels able to feed the GPU directly.

const pixels = new Uint8Array(memory.buffer, mod.exports.framebuffer_ptr(), w * h * 4);
gl.bindTexture(gl.TEXTURE_2D, tex);
gl.texSubImage2D(gl.TEXTURE_2D, 0, 0, 0, w, h, gl.RGBA, gl.UNSIGNED_BYTE, pixels);

Use texSubImage2D into a texture allocated once with texStorage2D, for the same reason bufferSubData beats bufferData. And be aware of UNPACK_ALIGNMENT: WebGL defaults to 4-byte row alignment, which is fine for RGBA and wrong for a three-channel texture whose width is not a multiple of four — the symptom is a diagonal skew identical to a stride bug.

Stalls: the thing that ruins a frame

WebGL is asynchronous by design. Commands queue and the GPU executes them later, which is why a frame can issue thousands of calls in two milliseconds. Any operation that needs a result from the GPU breaks that and waits for the queue to drain.

readPixels, getError in a loop, getParameter on state the driver has to query, and finish() are all synchronisation points that can cost 5–15 ms. If you need pixels back — for picking, for a screenshot, for feeding the result to the module — use a pixel buffer object with an asynchronous fence in WebGL 2, and read the result a frame or two later rather than immediately.

Checking gl.getError() every frame in production is a common accidental version of this. Wrap error checking in a development-only flag.

Where the frame budget goes Each GL call from the module is a crossing. Batching geometry into fewer, larger draw calls removes most of them, and the arithmetic is unchanged. one draw per object 1,200 crossings per frame — 14 ms of pure overhead batched by material 36 crossings 0.4 ms; the same triangles reach the GPU either way The geometry is identical in both rows — only the number of times the boundary is crossed differs. Upload vertex data once and update with a sub-range write rather than recreating the buffer per frame.

Gotchas

  • Detached view after growth. Compare memory.buffer identity, or use a non-growing module.
  • bufferData per frame. Reallocates; use bufferSubData into a preallocated buffer.
  • Unaligned texture rows. Set gl.pixelStorei(gl.UNPACK_ALIGNMENT, 1) for tightly packed non-RGBA data.
  • Context loss ignored. Browsers can drop the WebGL context at any time. Handle webglcontextlost and rebuild resources, or the canvas goes black permanently.
  • Interleaved attributes assumed. If the module writes struct-of-arrays, use separate buffers or correct strides; a wrong stride draws a recognisable but wrong mesh.
  • Uploading more than changed. subarray to the live range instead of uploading the whole capacity.

Performance note

Uploading 160 kB of instance data per frame from linear memory costs about 0.15 ms on a laptop — effectively free. Ten thousand individual draw calls cost roughly 38 ms; one instanced call with the same geometry costs 1.8 ms. Introducing a single readPixels into the frame added 9 ms in the same test. The ordering of those numbers is consistent across machines: draw call count and synchronisation dominate, data transfer does not.

Frequently Asked Questions

Should the module call GL itself through Emscripten? For a port, yes — it is what makes an existing engine work unchanged. For new code, calling GL from JavaScript over data the module owns is simpler to debug and integrates better with the rest of a page.

Does WebGL 2 matter, or is WebGL 1 enough? WebGL 2 gives instancing, vertex array objects, pixel buffer objects and integer textures without extensions. All of those are relevant here, and support is now broad enough that WebGL 1 is a fallback rather than a target.

Can I render in a worker? Yes, with OffscreenCanvas. Move the module and the context together, and keep the main thread for input and interface.

← Back to Graphics, Games & Simulation