Graphics, Games & Simulation

Real-time work has a property no other WebAssembly workload shares: a deadline that arrives sixty times a second and never moves. Everything in this topic follows from that. The compiled module is usually doing the simulation — physics, particles, pathfinding, animation — while the GPU does the drawing, and the interesting engineering is in the seam: how state gets from linear memory into draw calls without a copy, and how the loop stays inside its budget on a machine you do not own.

Prerequisites

  • [ ] A toolchain that can produce a web build: emcc for C and C++ engines, wasm-pack for Rust.
  • [ ] A canvas with a WebGL 2 or WebGPU context.
  • [ ] Familiarity with the browser frame loop — requestAnimationFrame, and why setInterval is wrong.
  • [ ] A test machine that is not your development machine. Frame budgets are hardware opinions.

The frame budget is the whole design

At 60 Hz you have 16.67 ms per frame for everything: input handling, simulation, the JavaScript that issues draw calls, the browser’s own compositing, and whatever else the page is doing. A comfortable target for your own work is under 10 ms, because the rest is not yours.

That budget decides the architecture. Simulation state lives in linear memory and stays there. Draw calls read directly from typed-array views over that memory. Nothing is serialised, nothing is converted to JavaScript objects, and nothing allocates during a frame — an allocation is a garbage collection waiting to land in the middle of one.

Where a frame actually goes Simulation in the module and draw call submission from JavaScript are the parts you control. GPU execution and the browser's compositing are not, and together they routinely consume a third of the frame. 16.67 ms at 60 Hz simulation in Wasm · 5 ms draw calls · 2 ms GPU · 4 ms composite · 3 ms yours to optimise not yours — budget around it A single allocation that triggers a collection costs 3–20 ms, which is why a frame loop that allocates drops frames unpredictably rather than consistently. Preallocate every buffer at startup and reuse them for the life of the session. At 120 Hz the budget halves to 8.3 ms, and the parts you do not control do not halve with it.

Two architectures, and the one you should prefer

There are two ways to structure a WebAssembly-driven renderer, and the choice shapes everything else.

In the module-drives-everything model, used by Emscripten ports of existing engines, the compiled code calls OpenGL ES functions that Emscripten translates into WebGL calls. The engine’s original render loop runs unchanged, driven by emscripten_set_main_loop. This is the fastest route for an existing codebase and the reason so many native games run in browsers at all.

In the module-simulates, JavaScript-draws model, the compiled code owns state and physics, and a thin JavaScript layer reads that state through typed-array views and issues draw calls itself. It requires writing the renderer, but it composes with the rest of a web application, gives direct access to WebGPU, and keeps the module small.

For a port, take the first. For a new feature inside an existing web application, take the second — it is much easier to integrate, and the performance difference is small because the expensive part was never the draw call submission.

Getting vertex data to the GPU without copying

The critical path is the same in both models: positions, colours and indices live in linear memory, and the GPU needs them. WebGL’s buffer upload functions accept a typed-array view, and a view over linear memory is a legitimate argument — so the upload reads straight from the module’s memory with no intermediate array.

const positions = new Float32Array(memory.buffer, mod.exports.positions_ptr(), count * 3);
gl.bindBuffer(gl.ARRAY_BUFFER, vbo);
gl.bufferSubData(gl.ARRAY_BUFFER, 0, positions);      // reads directly from linear memory

Two rules keep this correct. Rebuild the view after anything that might have grown memory, or use a module that never grows after startup — the second is strongly preferable in a frame loop. And use bufferSubData into a preallocated buffer rather than bufferData, which reallocates GPU storage and stalls the pipeline.

Fixed timestep, variable rendering

Simulation and rendering should not share a clock. A physics step that uses whatever time elapsed since the last frame is non-deterministic, behaves differently on a 144 Hz monitor, and explodes when a frame takes 300 ms because the user switched tabs.

The standard structure accumulates real time and consumes it in fixed steps:

const STEP = 1 / 120;                       // simulate at 120 Hz regardless of display
let acc = 0, last = performance.now();

function frame(now) {
  acc += Math.min((now - last) / 1000, 0.25);   // clamp: never simulate more than 0.25 s
  last = now;
  while (acc >= STEP) { mod.exports.step(STEP); acc -= STEP; }
  render(acc / STEP);                            // interpolation factor for smooth motion
  requestAnimationFrame(frame);
}

The clamp matters more than it looks. Without it, a tab restored after two minutes in the background tries to simulate twelve thousand steps in one frame, and the page hangs. With it, the simulation simply loses that time, which is the correct behaviour for everything except a deterministic replay.

Fixed steps, interpolated frames Real elapsed time accumulates and is consumed in equal simulation steps. Rendering interpolates between the last two states, so motion stays smooth even when the display refresh and the simulation rate do not divide evenly. frames (variable) 16.7 ms 22.1 ms — a slow frame 15.9 ms simulation steps (fixed 8.3 ms) always equal, always deterministic The slow frame consumes three steps instead of two; the leftover fraction becomes the interpolation factor so nothing visibly stutters. Without interpolation this design looks worse than a variable timestep, which is why the two halves must ship together.

Memory layout decides throughput

For simulation, the arrangement of data in linear memory matters more than the arithmetic performed on it. A particle system storing an array of structs — position, velocity, colour and lifetime interleaved per particle — touches every field on every pass even when a pass only needs positions, and it defeats vectorisation because the values a SIMD lane wants are not adjacent.

The alternative, a struct of arrays, stores all x coordinates together, then all y, then all velocities. An integration pass reads two contiguous streams and writes one, which is the ideal shape for both the cache and v128 instructions:

pub struct Particles {
    x: Vec<f32>, y: Vec<f32>,
    vx: Vec<f32>, vy: Vec<f32>,
    life: Vec<f32>,
}

pub fn integrate(p: &mut Particles, dt: f32) {
    for i in 0..p.x.len() {
        p.x[i] += p.vx[i] * dt;      // two streams in, one out — vectorises cleanly
        p.y[i] += p.vy[i] * dt;
    }
}

The same layout also makes the GPU upload trivial: a contiguous array of positions is exactly what bufferSubData wants, with no gather step in between. Converting an existing array-of-structs engine is invasive, which is an argument for choosing the layout before there is much code rather than after.

Capacity planning follows the same logic. Decide the maximum entity count at startup, allocate for it, and treat exceeding it as a design error rather than a reason to reallocate mid-frame. A pool with a free list gives you the dynamism without the allocation, and it keeps indices stable — which matters, because in a struct-of-arrays layout an index is the identity.

Input, and why it is harder than it looks

Input arrives as DOM events on the main thread, and the simulation wants it as state at a specific simulation step. Bridging that means buffering events and applying them at step boundaries rather than whenever they fire.

The simplest correct approach is a small input state block inside linear memory that JavaScript writes and the module reads at the start of each step. Key states, pointer position and button flags fit in a few dozen bytes, and updating them from event handlers costs nothing. Avoid calling into the module from an event handler directly — the event may arrive mid-frame, and a simulation that mutates during a step produces inconsistencies that are extremely hard to reproduce.

Pointer lock, gamepad polling and touch all layer on top of the same structure. Gamepads are polled rather than evented, so read them once per frame and write into the same block.

Running the simulation in a worker

Moving the module to a worker is attractive — the main thread then only issues draw calls — but WebGL contexts are not transferable, so the renderer must either stay on the main thread or use OffscreenCanvas to move the context into the worker with the module.

The second option is the better one when available: simulation and rendering both live in the worker, and the main thread handles only input and interface. With shared memory the two can also be split, with the worker simulating into a SharedArrayBuffer that the main thread reads for drawing — at the cost of requiring cross-origin isolation and careful synchronisation, since a half-updated state read mid-frame produces visible tearing.

Porting an existing engine

Engine ports are a well-trodden path, and the surprises are consistent. The main loop must be restructured — a while (running) loop cannot exist in a browser, because it never yields, so Emscripten’s main-loop machinery takes over the driving. File access becomes a virtual file system that needs preloading or fetching. Threads need cross-origin isolation. And the asset bundle, not the code, is usually what makes the download unacceptable.

The pages under this topic cover each: the loop restructuring in porting a C game loop to Emscripten, and the GPU paths in the WebGL and WebGPU guides.

Assets are usually the real problem

Teams porting a game to the browser tend to spend their effort on the compiled code and then discover that the code was never the issue. A modest native game ships hundreds of megabytes of textures, meshes and audio, and none of that gets smaller by being compiled to WebAssembly.

Three things help, in order of impact. Compress textures in a GPU-native format — Basis Universal transcodes to whatever the device supports, so one asset serves every platform at a fraction of the size of PNG. Stream assets by need rather than bundling them: the first level, the first area, the first few seconds of audio, with the rest fetched while the player is occupied. And separate code from content entirely so that a code update does not invalidate a two-hundred-megabyte asset cache.

The loading experience is part of this. A browser game that shows a progress bar reaching 100% and then sits for eight seconds instantiating is worse than one that reports each phase honestly. Report the fetch, the decode and the instantiation separately, and start the audio context on the first user gesture rather than at load, since browsers will refuse it otherwise.

Measure the total: transfer size, time to first interactive frame, and memory resident once the first scene is running. Those three numbers decide whether anyone stays long enough to see the frame rate you worked on.

What the module is actually doing In every graphics workload the browser owns the device. The module's job is the arithmetic that feeds it — vertices, physics state, pixels — never the drawing itself. WebGL rendering builds vertex and index data, issues draw calls WebGPU rendering builds buffers and encodes command passes physics simulation integrates state; no drawing at all software rasterising writes pixels into a shared framebuffer The first three keep the pixels on the GPU; only the last moves them through linear memory. Whichever it is, the per-frame boundary crossings decide whether the frame budget holds.

Gotchas and failure modes

  • Allocating in the frame loop. The single most common cause of irregular stutter. Preallocate.
  • Rebuilding typed-array views every frame. Cheap but not free, and unnecessary if memory never grows.
  • bufferData instead of bufferSubData. Reallocates GPU storage and stalls.
  • Simulation tied to frame rate. Different behaviour on different displays, and an explosion after a long pause.
  • Reading GPU state back per frame. readPixels and getParameter force a synchronisation point and can cost more than everything else combined.
  • Assuming 60 Hz. High-refresh displays exist, and so do power-saving modes that drop to 30.

Determinism, replays and multiplayer

Any feature that replays a session, synchronises two clients, or validates a score depends on the simulation producing identical results from identical inputs. WebAssembly helps here more than most platforms: integer arithmetic is exactly specified, and floating-point follows IEEE 754 with defined rounding, so the same module fed the same inputs produces the same bits on every engine.

The caveats are the ones you introduce. Threads change reduction order and therefore floating-point results, so a deterministic simulation must be single-threaded or must fix its reduction order explicitly. Anything seeded from Math.random, the wall clock or performance.now is non-deterministic by construction — seed from a value that is part of the recorded input instead. And iteration over a hash map whose order depends on pointer values will differ between runs; use a deterministic container in the simulation path.

Getting this right buys more than multiplayer. A recorded input stream plus a deterministic simulation is the best debugging tool available for a real-time system: a bug report becomes a file that reproduces the failure exactly, on your machine, as many times as you need. Teams that build this early tend to keep it forever, and teams that skip it spend their evenings trying to reproduce something a player saw once.

Verifying frame time honestly

Frame rate is a lagging indicator; frame time distribution is what users perceive. A steady 58 fps with occasional 40 ms frames feels worse than a steady 50.

const times = [];
function frame(now) {
  times.push(now - last); last = now;
  if (times.length === 600) {
    times.sort((a, b) => a - b);
    console.log({ p50: times[300], p95: times[570], p99: times[594], worst: times[599] });
    times.length = 0;
  }
  requestAnimationFrame(frame);
}

Report p95 and p99 alongside the median, and instrument the simulation separately from the frame so you know which half is responsible. The browser’s own performance panel is the right tool for finding where the time goes, but these numbers are what you track over time.

Guides in this topic

Frequently Asked Questions

Is WebAssembly fast enough for a real game? Yes, for the simulation. Compiled physics, animation and AI run at a large multiple of equivalent JavaScript, and the GPU does the drawing either way. What limits browser games is usually asset size and input latency rather than compute.

Should the renderer be in Wasm or JavaScript? Draw call submission is cheap in both. Keep it wherever integration is easiest — JavaScript for a feature inside a web application, the module for a port where the engine already owns rendering.

Does SIMD help? Substantially for particles, skinning, and any vector maths over arrays — 2–4× is typical. It helps physics broad-phase less than people expect, because that work is branch-heavy rather than arithmetic.

How do I keep the page responsive while the game runs? Use OffscreenCanvas and put simulation and rendering in a worker. Failing that, keep the frame budget honest: a game that consistently uses 14 ms of a 16.7 ms frame leaves nothing for the rest of the page.

How much does the boundary crossing cost per frame? Less than people fear. A call into the module is on the order of tens of nanoseconds, so even a few hundred calls per frame are irrelevant. What costs is crossing with data — converting arrays, building objects, or copying buffers — which is why the design keeps state on one side and passes indices and pointers rather than values.

Should audio go through the module too? Mixing and synthesis, yes, in an AudioWorklet with the same no-allocation discipline as any other real-time path. Triggering sounds is better done from JavaScript, because the sound system needs to know about page state — muting when the tab is hidden, respecting the user’s volume — that the simulation should not care about.

What frame rate should I target? Design for a variable rate and a fixed simulation step. Targeting a specific frame rate bakes in an assumption that high-refresh displays, battery-saver modes and lower-powered devices all violate, and the interpolated fixed-step loop handles every one of them without special cases.

Can I share the simulation between a browser client and a server? Yes, and it is one of the better reasons to put it in a compiled module: the same binary validates moves server-side under a standalone runtime and predicts them client-side in the tab, with identical results because the instruction semantics are identical.

← Back to Production Wasm: Workloads & Deployment