Choosing Between the Wasm and WebGPU Backends

This guide answers one task: pick an execution backend for browser inference, and build a fallback chain that gets the best available option on each device without ever leaving a user with nothing.

Prerequisites

  • [ ] A runtime that supports both providers — onnxruntime-web 1.19+ or transformers.js 3+.
  • [ ] Chrome 113+, Edge 113+ or Safari 18+ for WebGPU; Firefox support varies by platform.
  • [ ] A model you can run under both, with a fixture input and known-good output.
  • [ ] A way to measure honestly — warm-up discarded, percentiles rather than means.

The two cost curves

Every backend has a fixed cost and a marginal cost, and the whole decision follows from where those two curves cross.

The WebAssembly backend has a small fixed cost — instantiate the module, allocate an arena — and a marginal cost proportional to the arithmetic, executed four float lanes at a time. The WebGPU backend has a large fixed cost — request an adapter, compile shader pipelines, allocate GPU buffers, upload weights — and a marginal cost that is much lower because thousands of lanes run at once.

For a model that takes 8 ms on the CPU, the GPU cannot win: the upload alone costs more than the whole inference. For a model that takes 400 ms on the CPU, the GPU usually wins by a factor of three to eight. The interesting region is in between, and it is exactly where most production models sit, which is why this decision needs measurement rather than a rule of thumb.

Fixed cost versus marginal cost The WebAssembly backend starts cheaply and scales linearly with arithmetic. The WebGPU backend pays a large setup cost and then scales far more slowly, so it overtakes once the model is big enough to amortise the setup. arithmetic per inference → total latency crossover Wasm: low setup, steep slope WebGPU: high setup, shallow slope Setup is paid once per session, so a page that runs one inference and a page that runs a thousand sit on opposite sides of the same crossover.

Availability is not a detail

A backend that is faster on your machine and missing on a third of your users’ machines is not faster. WebGPU availability depends on the browser, the operating system, the GPU and the driver — and browsers disable it on known-bad driver versions, so support can disappear between two visits from the same user.

async function pickProviders() {
  const list = [];
  if ('gpu' in navigator) {
    try {
      const adapter = await navigator.gpu.requestAdapter();
      if (adapter) list.push('webgpu');
    } catch { /* no adapter, or blocked */ }
  }
  list.push('wasm');                 // always last, always present
  return list;
}

const session = await ort.InferenceSession.create(url, { executionProviders: await pickProviders() });

Note that 'gpu' in navigator is not enough on its own: the property can exist while requestAdapter() resolves to null on a machine whose GPU is blocklisted. Always request the adapter before committing to the provider.

Operator coverage and silent fallback

Runtimes do not implement every operator on every backend. When a graph contains an operator the GPU backend lacks, the runtime typically falls back to the CPU for that node — which means a round trip between GPU and CPU memory for every inference, at every such node. One unsupported operator in the middle of a network can make the GPU path slower than pure WebAssembly.

The way to see this is the runtime’s own logging, turned up during development:

ort.env.logLevel = 'verbose';
// look for lines about node assignment / placement per execution provider

If a substantial run of nodes lands on the CPU provider, either simplify the graph — many such operators come from an exporter emitting an exotic node for something ordinary — or accept the WebAssembly path for that model. Re-exporting with a lower opset, or with constant folding applied, frequently removes the offending node entirely.

Memory behaves differently on each path

On the WebAssembly backend, weights and activations live in linear memory and count against the tab’s budget in the obvious way. On the WebGPU backend, weights live in GPU buffers, activations are transient GPU allocations, and the tab holds much less — but the GPU has its own limits, and exceeding maxBufferSize or maxStorageBufferBindingSize produces a validation error rather than a graceful degradation.

That difference has a practical consequence for large models: a model that cannot fit in a mobile tab may fit on the GPU, and a model that fits comfortably on the CPU may exceed a single GPU buffer limit and need splitting. Neither is a reason to prefer one backend in general; both are reasons to test the actual model on the actual devices before deciding.

The upload itself is a cost people forget. Moving 200 MB of weights into GPU buffers takes real time at session creation, and it happens again if the GPU device is lost — which browsers do when a tab is backgrounded for long enough on some platforms. Handle device.lost and rebuild rather than assuming the session lives forever.

Measuring the two honestly

A fair comparison controls for four things: warm-up, input, thread count and what is included in the measurement.

async function bench(session, feeds, n = 50) {
  await session.run(feeds);                       // discard the first
  const times = [];
  for (let i = 0; i < n; i++) {
    const t = performance.now();
    await session.run(feeds);
    times.push(performance.now() - t);
  }
  times.sort((a, b) => a - b);
  return { p50: times[n >> 1], p95: times[Math.floor(n * 0.95)] };
}

Report the median and the 95th percentile, not the mean — GPU timings in particular have a long tail caused by scheduling and other tabs competing for the device. Measure session creation separately and state it separately, because for a page that runs one inference, creation is the latency the user experiences.

Two numbers, not one Session creation and per-inference time must be reported separately. A backend with fast inference and slow creation loses on a page that runs one inference and wins on a page that runs many. Wasm, 4 threads, SIMD session create: 180 ms per inference p50: 42 ms per inference p95: 51 ms one-shot total: 222 ms WebGPU session create: 610 ms per inference p50: 9 ms per inference p95: 23 ms one-shot total: 619 ms · 100 runs: 1.5 s Same model, same machine, opposite conclusions depending on how many inferences the page runs. Numbers are illustrative — measure your own.

Building the fallback chain

The shipping architecture is a chain, not a choice. Try the best option, verify it produced a sane result, and degrade on failure — including failures that happen after a successful session creation, because a GPU device can be lost mid-session.

const CHAIN = [
  { name: 'webgpu', opts: { executionProviders: ['webgpu'] } },
  { name: 'wasm-threaded', opts: { executionProviders: ['wasm'] }, requires: () => crossOriginIsolated },
  { name: 'wasm', opts: { executionProviders: ['wasm'] } },
];

for (const step of CHAIN) {
  if (step.requires && !step.requires()) continue;
  try {
    const s = await ort.InferenceSession.create(url, step.opts);
    await s.run(fixtureFeeds);                    // proves it actually runs, not just loads
    report({ backend: step.name });
    return s;
  } catch (e) { report({ backendFailed: step.name, error: String(e) }); }
}
throw new Error('no usable backend');

Running a fixture inference before accepting a session is the part people skip, and it is the part that catches driver bugs — sessions that create successfully and then produce zeros or throw on the first real input. Log which step succeeded; the field distribution is almost never what the team predicted, and it is the data you need to decide whether the GPU path is worth maintaining at all.

Where the crossover sits The GPU backend pays a fixed setup and transfer cost per inference. Below a certain model size the Wasm backend wins outright, and above it the GPU wins by a wide margin. small model, 4 MB Wasm 18 ms WebGPU 31 ms — setup dominates the work medium model, 40 MB Wasm 140 ms WebGPU 44 ms large model, 400 MB Wasm 1,980 ms Ship both and pick at runtime: availability varies by device, and a GPU request can simply be refused. Measure on the hardware your users have — a desktop discrete GPU tells you nothing about a mid-range phone.

Gotchas

  • WebGPU works in development and not in production. WebGPU requires a secure context. It works on localhost over HTTP and silently disappears on an internal HTTP staging host.
  • First inference on the GPU is enormously slow. Shader compilation happens lazily per pipeline. Warm up with a real-shaped input, not a zero-size one, or the compile lands on the user’s first click.
  • Output differs between backends. Different accumulation order, different precision. Compare with a tolerance; a strict equality check between backends will always fail.
  • Unsupported data type on the GPU path. Some quantized formats are CPU-only. An int8 model that flies on WebAssembly may not run on WebGPU at all.
  • The GPU path starves the rest of the page. Heavy inference competes with rendering for the same device. If the page animates while inferring, budget for both or move the work to idle time.

Performance note

Across typical vision and embedding models, WebGPU delivers 3–8× the throughput of the threaded WebAssembly backend once past the crossover, and roughly 10–20× the single-threaded backend. Below about 20 ms of CPU work per inference it is usually a loss. The most reliable predictor is not file size but the largest matrix multiplication in the graph: if the model’s dominant GEMM is smaller than roughly 512 × 512, the GPU has nothing to get its teeth into.

Frequently Asked Questions

Should I ship both and pick at runtime? Yes. The WebAssembly backend is required as a fallback regardless, and the WebGPU path is additive. The cost is the extra runtime files, which are cached and only fetched when that path is selected.

What about WebNN? Promising and worth a probe in the chain ahead of WebGPU when available, since it can reach platform accelerators directly. Operator coverage is still the limiting factor, so treat it as an optional first step rather than a replacement.

Does the choice change on mobile? Substantially. Mobile GPUs have much smaller buffer limits and more aggressive power management, and thermal throttling shows up as p95 latency drifting upward over a long session. Measure over minutes, not seconds, before committing to the GPU path on phones.

← Back to Machine Learning Inference in the Browser