Using Wasm Threads in Node.js worker_threads
This page answers one task: a threaded WebAssembly module — or one you want to parallelise — must run in Node.js, for a CLI, a build tool or a server, and
you want shared memory across worker_threads working without the browser’s cross-origin isolation requirements getting in the way.
Prerequisites
- [ ] Node.js 18 or later (current LTS recommended).
- [ ] A module built with shared-memory support (atomics and bulk-memory), or a plan to build one.
- [ ] Familiarity with
worker_threadsbasics.
What is different from the browser
In browsers, SharedArrayBuffer and shared WebAssembly.Memory are available only on cross-origin isolated pages, which requires COOP and COEP headers.
Node has no such requirement: SharedArrayBuffer is always available, and worker_threads can share memories and buffers freely. Two other differences
matter. Atomics.wait is allowed on Node’s main thread, because blocking it freezes a process rather than a page — though in a server it still blocks the
event loop and all request handling. And Node’s workers are heavier to start than many expect (tens of milliseconds each, each with its own V8 isolate),
so pools should be created once and reused.
Otherwise the model is the same: one compiled WebAssembly.Module and one shared WebAssembly.Memory are posted to each worker; each worker instantiates
the module with the shared memory; threads coordinate through atomics in that memory.
Step 1 — share a module and memory with workers
// main.mjs
import { Worker } from "node:worker_threads";
import { readFile } from "node:fs/promises";
const module = await WebAssembly.compile(await readFile(new URL("./kernel.wasm", import.meta.url)));
const memory = new WebAssembly.Memory({ initial: 256, maximum: 16384, shared: true });
const control = new Int32Array(new SharedArrayBuffer(64));
const workers = Array.from({ length: 4 }, (_, id) =>
new Worker(new URL("./worker.mjs", import.meta.url), { workerData: { module, memory, control, id } }));
// worker.mjs
import { workerData } from "node:worker_threads";
const { module, memory, control, id } = workerData;
const instance = await WebAssembly.instantiate(module, { env: { memory } });
instance.exports.thread_init(id);
// ... wait for jobs on `control` with Atomics.wait, run exports, signal completion ...
WebAssembly.Module and shared memories can be passed via workerData or postMessage; they are shared, not copied. The same chunking and signalling
patterns as in the browser apply.
Step 2 — run Emscripten pthreads builds in Node
Emscripten’s pthreads support Node directly: the generated JavaScript detects Node and uses worker_threads for its pool.
emcc src/*.c -O3 -pthread -sPTHREAD_POOL_SIZE=8 -sENVIRONMENT=node -o kernel.mjs
node kernel.mjs
-sPTHREAD_POOL_SIZE pre-creates workers at startup so pthread_create does not wait; -sENVIRONMENT=node drops browser code paths. Because Node allows
blocking on the main thread, the main-thread restrictions that complicate browser pthreads mostly disappear, though -sPROXY_TO_PTHREAD is still useful
for servers that must keep the event loop responsive.
Step 3 — run Rust rayon code in Node
wasm-bindgen-rayon is designed for browsers but works in Node with a little setup: the generated glue spawns its workers with Worker from the web API,
so in Node you either provide a Worker polyfill backed by worker_threads (packages such as web-worker exist for this) or build for Node with a pool
created by your own code. Many projects choose a simpler path in Node: compile the Rust code natively as a Node addon (with napi-rs) when native
threads are acceptable, and keep the Wasm build for browsers. If a single portable artefact matters more, use the Wasm build with a Node worker polyfill
and test it in CI.
Step 4 — keep the event loop free in servers
Atomics.wait on the main thread works in Node but stops the event loop: no requests are accepted, no timers fire. In a server, wait with
Atomics.waitAsync (available in current Node) or use message passing (worker.postMessage and once("message")) to learn when work is done:
async function runJob(job) {
writeJob(job);
Atomics.store(control, 1, 0);
Atomics.add(control, 2, 1);
Atomics.notify(control, 2);
while (Atomics.load(control, 1) < workers.length) {
const r = Atomics.waitAsync(control, 1, Atomics.load(control, 1));
if (r.async) await r.value;
}
}
For CLIs and build tools where the process does nothing else, blocking waits are simpler and fine.
Step 5 — size the pool for the machine
os.availableParallelism() (Node 18.14+) reports the number of CPUs the process may use, which respects container CPU limits better than
os.cpus().length. Size pools to it, minus one for the main thread in servers. In containers with CPU quotas, more workers than the quota allows only adds
contention. Libraries such as piscina manage pools of workers for task-based workloads; for shared-memory Wasm jobs, a fixed pool coordinated by atomics
is usually simpler and faster.
Memory limits and growth
Node’s default heap limits apply to the JavaScript heap, not to Wasm memory, but the process’s total memory is still bounded by the machine or container. Shared memories must declare a maximum; reserve generously but realistically, because some engines reserve virtual address space up front for the maximum of shared memories. As in the browser, grow memory before starting parallel work, never during it, and re-create views in every thread after growth.
Debugging and profiling
node --inspect exposes workers to Chrome DevTools, which lists them as separate targets; each can be paused and profiled. --cpu-prof writes CPU
profiles for the main thread and, with --cpu-prof-dir, for workers as well. Wasm frames appear with function names if the module keeps its name
section. Deadlock diagnosis works as in the browser — pause each worker and read its stack — described in
debugging deadlocks in threaded Wasm.
Shutting down cleanly
Workers keep a Node process alive. A CLI that finishes its work but leaves pool workers sleeping in Atomics.wait never exits, which looks like a hang.
Give the pool an explicit shutdown: set a “stop” flag in the control buffer, notify all workers so they wake, let each worker see the flag and return from
its loop, then call worker.terminate() on any that have not exited after a short grace period. Alternatively, call worker.unref() on pool workers so
they do not keep the process alive on their own; the process then exits when the main thread has nothing left to do. For servers, hook shutdown into the
process’s signal handling (SIGTERM) so in-flight jobs finish or are cancelled before the pool is torn down, and so container orchestrators see a clean
exit rather than a forced kill.
Error handling across the pool
A trap in one worker’s instance — out-of-bounds access, a Rust panic — ends that worker’s current job with an exception in the worker. Catch it in the worker, record it in the control buffer or post it as a message, and signal completion anyway, or the coordinator waits forever for a job that will never finish. Because all workers share memory, a trap may have left shared data half-updated; treat any trap in a shared-memory job as invalidating that job’s output, and if the module’s internal state might be corrupted, restart the whole pool from the compiled module rather than continuing with a possibly inconsistent shared heap.
Expected output
A CLI processes a 2 GB dataset with an 8-worker pool sharing one memory, 5.6× faster than single-threaded; a server runs the same kernel per request
with Atomics.waitAsync, keeping p99 latency for unrelated endpoints unchanged during heavy jobs; the Emscripten pthreads build runs in Node with a
pre-created pool; and pool size follows os.availableParallelism() inside a 4-CPU container.
Gotchas
- Blocking waits in servers. They stop the event loop. Use
waitAsyncor messages. - Creating workers per request. Startup costs tens of milliseconds. Reuse a pool.
os.cpus().lengthin containers. Ignores quotas. UseavailableParallelism().- Assuming rayon glue works in Node unchanged. It expects web
Worker. Polyfill or use another path. - Growing shared memory during jobs. Views go stale. Grow first.
- Workers that never exit. Sleeping workers keep the process alive. Add an explicit shutdown.
Performance note
Starting a worker_threads worker and instantiating a 2 MB module from a shared compiled Module took about 35 ms per worker; reusing the pool made
per-job overhead under 0.1 ms.
Frequently Asked Questions
Do I need COOP/COEP in Node? No — they are browser headers; Node always allows shared memory.
Can Deno and Bun do the same? Deno supports shared memory across its workers; Bun’s worker support is improving — test your workload.
Is a native addon faster than threaded Wasm? Often somewhat, but Wasm runs on every platform without compiling per OS and architecture.
Can workers share one instance? No — each worker needs its own instance; they share memory and the compiled module.
Why does my CLI not exit after the work is done?
Pool workers sleeping in Atomics.wait keep the process alive; signal a stop flag and terminate them, or unref() the workers.
Related
- Running Wasm off the event loop in Node.js — single-threaded offloading.
- Parallelizing a loop across workers with shared memory — the chunking pattern.
- Replacing a native Node addon with Wasm — the trade-offs.
- Building a Wasm thread pool — pool design.
← Back to SharedArrayBuffer, Atomics & Threading