Using Wasm Threads in Node.js worker_threads

This page answers one task: a threaded WebAssembly module — or one you want to parallelise — must run in Node.js, for a CLI, a build tool or a server, and you want shared memory across worker_threads working without the browser’s cross-origin isolation requirements getting in the way.

Prerequisites

  • [ ] Node.js 18 or later (current LTS recommended).
  • [ ] A module built with shared-memory support (atomics and bulk-memory), or a plan to build one.
  • [ ] Familiarity with worker_threads basics.

What is different from the browser

In browsers, SharedArrayBuffer and shared WebAssembly.Memory are available only on cross-origin isolated pages, which requires COOP and COEP headers. Node has no such requirement: SharedArrayBuffer is always available, and worker_threads can share memories and buffers freely. Two other differences matter. Atomics.wait is allowed on Node’s main thread, because blocking it freezes a process rather than a page — though in a server it still blocks the event loop and all request handling. And Node’s workers are heavier to start than many expect (tens of milliseconds each, each with its own V8 isolate), so pools should be created once and reused.

Otherwise the model is the same: one compiled WebAssembly.Module and one shared WebAssembly.Memory are posted to each worker; each worker instantiates the module with the shared memory; threads coordinate through atomics in that memory.

Wasm threads in the browser versus Node.js In the browser, shared memory needs cross-origin isolation headers and Atomics.wait is forbidden on the main thread. In Node.js, shared memory is always available, Atomics.wait works on the main thread but blocks the event loop, and workers are heavier to start, so pools should be reused. browser needs COOP + COEP headers main thread cannot block Web Workers isolation required Node.js worker_threads SharedArrayBuffer always on main thread can wait (blocks loop) workers costly to start no headers needed

Step 1 — share a module and memory with workers

// main.mjs
import { Worker } from "node:worker_threads";
import { readFile } from "node:fs/promises";

const module = await WebAssembly.compile(await readFile(new URL("./kernel.wasm", import.meta.url)));
const memory = new WebAssembly.Memory({ initial: 256, maximum: 16384, shared: true });
const control = new Int32Array(new SharedArrayBuffer(64));

const workers = Array.from({ length: 4 }, (_, id) =>
  new Worker(new URL("./worker.mjs", import.meta.url), { workerData: { module, memory, control, id } }));
// worker.mjs
import { workerData } from "node:worker_threads";
const { module, memory, control, id } = workerData;
const instance = await WebAssembly.instantiate(module, { env: { memory } });
instance.exports.thread_init(id);
// ... wait for jobs on `control` with Atomics.wait, run exports, signal completion ...

WebAssembly.Module and shared memories can be passed via workerData or postMessage; they are shared, not copied. The same chunking and signalling patterns as in the browser apply.

Step 2 — run Emscripten pthreads builds in Node

Emscripten’s pthreads support Node directly: the generated JavaScript detects Node and uses worker_threads for its pool.

emcc src/*.c -O3 -pthread -sPTHREAD_POOL_SIZE=8 -sENVIRONMENT=node -o kernel.mjs
node kernel.mjs

-sPTHREAD_POOL_SIZE pre-creates workers at startup so pthread_create does not wait; -sENVIRONMENT=node drops browser code paths. Because Node allows blocking on the main thread, the main-thread restrictions that complicate browser pthreads mostly disappear, though -sPROXY_TO_PTHREAD is still useful for servers that must keep the event loop responsive.

Step 3 — run Rust rayon code in Node

wasm-bindgen-rayon is designed for browsers but works in Node with a little setup: the generated glue spawns its workers with Worker from the web API, so in Node you either provide a Worker polyfill backed by worker_threads (packages such as web-worker exist for this) or build for Node with a pool created by your own code. Many projects choose a simpler path in Node: compile the Rust code natively as a Node addon (with napi-rs) when native threads are acceptable, and keep the Wasm build for browsers. If a single portable artefact matters more, use the Wasm build with a Node worker polyfill and test it in CI.

A Node.js process running a threaded Wasm module The main thread compiles the module once and creates a shared memory. It starts a pool of worker_threads, passing both. Each worker instantiates the module with the shared memory. Jobs are signalled through atomics, workers process them in parallel, and the main thread awaits completion without blocking request handling. compile once WebAssembly.compile shared Memory no isolation headers worker pool created at startup instantiate per worker same module + memory jobs via Atomics main awaits, no block

Step 4 — keep the event loop free in servers

Atomics.wait on the main thread works in Node but stops the event loop: no requests are accepted, no timers fire. In a server, wait with Atomics.waitAsync (available in current Node) or use message passing (worker.postMessage and once("message")) to learn when work is done:

async function runJob(job) {
  writeJob(job);
  Atomics.store(control, 1, 0);
  Atomics.add(control, 2, 1);
  Atomics.notify(control, 2);
  while (Atomics.load(control, 1) < workers.length) {
    const r = Atomics.waitAsync(control, 1, Atomics.load(control, 1));
    if (r.async) await r.value;
  }
}

For CLIs and build tools where the process does nothing else, blocking waits are simpler and fine.

Step 5 — size the pool for the machine

os.availableParallelism() (Node 18.14+) reports the number of CPUs the process may use, which respects container CPU limits better than os.cpus().length. Size pools to it, minus one for the main thread in servers. In containers with CPU quotas, more workers than the quota allows only adds contention. Libraries such as piscina manage pools of workers for task-based workloads; for shared-memory Wasm jobs, a fixed pool coordinated by atomics is usually simpler and faster.

Memory limits and growth

Node’s default heap limits apply to the JavaScript heap, not to Wasm memory, but the process’s total memory is still bounded by the machine or container. Shared memories must declare a maximum; reserve generously but realistically, because some engines reserve virtual address space up front for the maximum of shared memories. As in the browser, grow memory before starting parallel work, never during it, and re-create views in every thread after growth.

Debugging and profiling

node --inspect exposes workers to Chrome DevTools, which lists them as separate targets; each can be paused and profiled. --cpu-prof writes CPU profiles for the main thread and, with --cpu-prof-dir, for workers as well. Wasm frames appear with function names if the module keeps its name section. Deadlock diagnosis works as in the browser — pause each worker and read its stack — described in debugging deadlocks in threaded Wasm.

Shutting down cleanly

Workers keep a Node process alive. A CLI that finishes its work but leaves pool workers sleeping in Atomics.wait never exits, which looks like a hang. Give the pool an explicit shutdown: set a “stop” flag in the control buffer, notify all workers so they wake, let each worker see the flag and return from its loop, then call worker.terminate() on any that have not exited after a short grace period. Alternatively, call worker.unref() on pool workers so they do not keep the process alive on their own; the process then exits when the main thread has nothing left to do. For servers, hook shutdown into the process’s signal handling (SIGTERM) so in-flight jobs finish or are cancelled before the pool is torn down, and so container orchestrators see a clean exit rather than a forced kill.

Error handling across the pool

A trap in one worker’s instance — out-of-bounds access, a Rust panic — ends that worker’s current job with an exception in the worker. Catch it in the worker, record it in the control buffer or post it as a message, and signal completion anyway, or the coordinator waits forever for a job that will never finish. Because all workers share memory, a trap may have left shared data half-updated; treat any trap in a shared-memory job as invalidating that job’s output, and if the module’s internal state might be corrupted, restart the whole pool from the compiled module rather than continuing with a possibly inconsistent shared heap.

Expected output

A CLI processes a 2 GB dataset with an 8-worker pool sharing one memory, 5.6× faster than single-threaded; a server runs the same kernel per request with Atomics.waitAsync, keeping p99 latency for unrelated endpoints unchanged during heavy jobs; the Emscripten pthreads build runs in Node with a pre-created pool; and pool size follows os.availableParallelism() inside a 4-CPU container.

Gotchas

  • Blocking waits in servers. They stop the event loop. Use waitAsync or messages.
  • Creating workers per request. Startup costs tens of milliseconds. Reuse a pool.
  • os.cpus().length in containers. Ignores quotas. Use availableParallelism().
  • Assuming rayon glue works in Node unchanged. It expects web Worker. Polyfill or use another path.
  • Growing shared memory during jobs. Views go stale. Grow first.
  • Workers that never exit. Sleeping workers keep the process alive. Add an explicit shutdown.

Performance note

Starting a worker_threads worker and instantiating a 2 MB module from a shared compiled Module took about 35 ms per worker; reusing the pool made per-job overhead under 0.1 ms.

Per-job overhead with and without a reused worker pool Milliseconds of overhead per job when creating four workers per job compared with dispatching to a pool of four workers created at startup. ms overhead per job new workers per job 140 ms reused pool 0.1 ms

Frequently Asked Questions

Do I need COOP/COEP in Node? No — they are browser headers; Node always allows shared memory.

Can Deno and Bun do the same? Deno supports shared memory across its workers; Bun’s worker support is improving — test your workload.

Is a native addon faster than threaded Wasm? Often somewhat, but Wasm runs on every platform without compiling per OS and architecture.

Can workers share one instance? No — each worker needs its own instance; they share memory and the compiled module.

Why does my CLI not exit after the work is done? Pool workers sleeping in Atomics.wait keep the process alive; signal a stop flag and terminate them, or unref() the workers.

← Back to SharedArrayBuffer, Atomics & Threading