Multi-Threaded Inference with Wasm Threads
This guide answers one task: make browser inference use more than one core, verify that it actually did, and choose a thread count that helps rather than hurts — because the threaded backend is usually the largest single speedup available and it fails silently when misconfigured.
Prerequisites
- [ ] Control over the HTTP response headers for your document and its assets.
- [ ] A runtime with a threaded build —
onnxruntime-web,transformers.js, or an Emscripten build with-pthread. - [ ] A model whose runtime is dominated by matrix multiplication, which is nearly all of them.
- [ ] A test device with at least four cores; scaling on two is not informative.
The two headers that gate everything
Threaded WebAssembly requires SharedArrayBuffer, and SharedArrayBuffer requires the document to be
cross-origin isolated. That state is granted only when both of these headers are present on the document
response:
Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp
The consequences reach further than the document. With require-corp in force, every cross-origin
subresource — images, fonts, scripts, iframes, analytics beacons — must opt in with
Cross-Origin-Resource-Policy: cross-origin or it is blocked outright. That is why enabling threads is a
site-wide decision rather than a page-level one, and why it frequently breaks embedded third-party
content the first time it is turned on.
Cross-Origin-Embedder-Policy: credentialless is a softer alternative that loads cross-origin resources
without credentials rather than requiring opt-in. It is widely supported now and breaks far less, at the
cost of cookies not being sent to those resources. Try credentialless first; fall back to
require-corp only if something genuinely needs it.
Turning it on
With isolation in place, enabling threads is one setting. The runtime detects SharedArrayBuffer, loads
the threaded variant of its own module, and spawns a pool of workers.
import * as ort from 'onnxruntime-web';
const cores = navigator.hardwareConcurrency || 2;
ort.env.wasm.numThreads = crossOriginIsolated ? Math.min(4, Math.max(1, cores >> 1)) : 1;
ort.env.wasm.simd = true;
console.log({ crossOriginIsolated, cores, threads: ort.env.wasm.numThreads });
Halving hardwareConcurrency is deliberate. The property reports logical processors, which on a machine
with simultaneous multithreading is twice the physical core count. Matrix kernels are already saturating
the execution units, so two threads on one physical core contend for the same cache and registers and
deliver almost nothing — sometimes less than nothing.
Capping at four is also deliberate for typical models. Beyond four threads the synchronisation between GEMM tiles starts to cost more than the extra parallelism returns, and the browser has other work to do. Measure on your own model before raising the cap.
Why scaling stops short of linear
A perfectly parallel workload on four cores would be four times faster. Inference is not perfectly parallel, and understanding the gap prevents chasing it.
Part of every model is sequential: activation functions, normalisations, reshapes, the graph executor’s own bookkeeping. Amdahl’s law applies directly — if 15% of the runtime is sequential, four threads cap out at about 2.8× no matter how good the parallel part is. Part of the work is memory-bandwidth-bound rather than compute-bound, and adding threads does not add bandwidth. And the thread pool itself costs something: a barrier per operator, cache lines bouncing between cores, and the occasional scheduling delay when the browser has other work.
The observed result across typical vision and transformer models is 1.7–2.0× on two threads, 2.5–3.5× on four, and 3–4.5× on eight — with the eight-thread figure depending heavily on whether those are physical cores. If you are seeing less than 1.5× on four threads, the model is probably too small for the pool to pay for itself, and single-threaded is the right answer.
Detecting silently single-threaded traffic
The dangerous failure is not an error; it is production quietly running on one thread while your local
machine runs on four. A deploy that drops a header, a CDN that strips one, a new third-party script that
forces you to relax require-corp — all of these downgrade performance without breaking anything.
report({
metric: 'inference',
crossOriginIsolated,
threads: ort.env.wasm.numThreads,
hardwareConcurrency: navigator.hardwareConcurrency,
p50: timings.p50,
});
Send those four fields with every timing sample. The moment isolation regresses, the ratio of isolated to non-isolated sessions moves and you can see it the same day rather than three sprints later when someone notices the feature feels slow. It is also the only way to know the real distribution across your users, which is rarely what a team assumes: embedded contexts, some in-app browsers and certain enterprise proxies never achieve isolation at all.
Verifying the pool is real
Configuration that looks right is not evidence. Three checks, run once after any change to headers, hosting or the runtime version, tell you whether the pool actually exists.
The first is the isolation flag itself: crossOriginIsolated must be true in the context that creates
the session, which for a worker means checking inside the worker rather than on the page. The second is
the network panel: the runtime fetches a different .wasm file for the threaded build, and seeing the
non-threaded one load is immediate proof that something upstream decided threads were unavailable.
The third and most convincing is a timing comparison you run yourself:
async function threadScaling(url, feeds, counts = [1, 2, 4]) {
const rows = [];
for (const n of counts) {
ort.env.wasm.numThreads = n;
const s = await ort.InferenceSession.create(url, { executionProviders: ['wasm'] });
await s.run(feeds); // warm up
const t = performance.now();
for (let i = 0; i < 20; i++) await s.run(feeds);
rows.push({ threads: n, msPerRun: (performance.now() - t) / 20 });
await s.release?.();
}
console.table(rows);
}
If the three rows are within a few percent of each other, the pool is not doing anything regardless of
what the configuration says. Note that numThreads must be set before the session is created — changing
it afterwards has no effect, which is a common source of benchmarks that appear to show threads making no
difference at all.
Threads in a worker
Running inference in a worker and running it with threads are independent choices, and you usually want both. The worker keeps the main thread free; the threads make the inference itself faster. A worker can spawn nested workers, which is what the threaded runtime does internally, and that works in every current browser as long as the isolation state is inherited — which it is, since workers inherit their creator’s agent cluster.
What you should not do is run several inference workers each with their own thread pool. Four workers with four threads each means sixteen threads competing for four cores, plus four copies of the model in memory. One worker holding one session with a four-thread pool is faster and uses a quarter of the memory. If you genuinely need concurrent inferences, queue them into the single session — the runtime serialises them anyway.
Gotchas
SharedArrayBuffer is not defined. Isolation is off. Check the document’s response headers, not the asset’s.- Everything works locally, threads are off in production. A CDN or reverse proxy is stripping or
overriding the headers. Verify with
curl -Iagainst the production URL. - Enabling COEP breaks third-party embeds. Expected. Move to
credentialless, or host the resource yourself, or ask the provider forCross-Origin-Resource-Policy: cross-origin. - More threads made it slower. Oversubscription, usually from trusting
hardwareConcurrencyon a machine with simultaneous multithreading. Halve it. - Thread count changes nothing. The runtime loaded the non-threaded variant because that file was
missing from your vendor directory. Check the network panel for which
.wasmwas actually fetched.
Performance note
On a four-physical-core laptop with a quantized transformer encoder: single-threaded p50 was 118 ms, two threads 64 ms (1.84×), four threads 41 ms (2.88×), eight threads 38 ms (3.11×). Memory grew by about 9 MB per additional thread for per-thread arenas. The four-thread configuration is the obvious choice — the eighth thread bought 3 ms and cost 36 MB.
Frequently Asked Questions
Is credentialless safe to use?
It is a standard, supported mode designed for exactly this. Cross-origin resources load without
credentials, so anything requiring cookies — an authenticated image CDN, for instance — will fail and
needs hosting differently. For most sites it is strictly better than require-corp.
Do threads help on mobile? Less than on desktop. Mobile chips have fewer performance cores and aggressive thermal management, so a four-thread pool can throttle within seconds. Two threads is often the sweet spot on phones, and worth measuring separately rather than sharing the desktop configuration.
Can I enable isolation for one route only? Yes — the headers are per-response, so a single route can be isolated while the rest of the site is not. This is the usual way to ship threads without auditing every third-party script on every page.
Related
- Configuring COOP/COEP headers for SharedArrayBuffer — the server-side half in detail.
- Using Atomics for Wasm thread synchronization — what the pool does underneath.
- Measuring inference latency in the browser — producing the numbers quoted here.