Building a Wasm Thread Pool

This page answers one task: you want several Web Workers to run functions of the same WebAssembly module in parallel, sharing its memory, without Rayon or Emscripten’s pthreads doing it for you — to understand the mechanism or to build a pool tailored to your workload.

Prerequisites

  • [ ] A module built with shared memory and atomics (+atomics,+bulk-memory, a declared maximum memory, imported shared memory).
  • [ ] A cross-origin isolated page.
  • [ ] Familiarity with sharing memory between Wasm and Web Workers.

What a thread is in WebAssembly

There is no thread-creation instruction in WebAssembly. A “thread” is a Web Worker running its own instance of the module, created from the same compiled WebAssembly.Module and given the same shared WebAssembly.Memory as an import. Instances share memory but not globals: each has its own stack pointer global and its own thread-local-storage base, so each needs its own region of memory for its stack and TLS block. Code, heap and static data are shared.

A pool therefore has four jobs. Start N workers and instantiate the module in each with the shared memory. Give each worker its own stack and TLS area before it runs any code. Provide a way to hand work to idle workers and wake them. And shut everything down cleanly. Libraries do all four; doing them by hand shows what those libraries are taking care of.

Anatomy of a Wasm thread pool The main thread compiles the module once and creates a shared memory. Each worker instantiates the module with that memory and gets its own stack and TLS region. A job queue and a signal counter in shared memory let the main thread hand out work and wake idle workers. main thread compile once, create shared memory, submit jobs workers 1..N own instance, own stack + TLS job queue in shared memory, Atomics-protected shared linear memory code data, heap, stacks, TLS blocks isolation COOP + COEP required

Step 1 — compile once, create the shared memory, start workers

const module = await WebAssembly.compileStreaming(fetch("compute.wasm"));
const memory = new WebAssembly.Memory({ initial: 256, maximum: 16384, shared: true });

const workers = [];
for (let i = 0; i < N; i++) {
  const w = new Worker(new URL("./pool-worker.js", import.meta.url), { type: "module" });
  w.postMessage({ type: "init", module, memory, id: i });
  workers.push(w);
}
await Promise.all(workers.map((w) => new Promise((r) => w.addEventListener("message", r, { once: true }))));

Posting the compiled WebAssembly.Module avoids each worker compiling it again; posting the shared Memory gives all of them the same buffer. The main thread usually instantiates the module too, to allocate memory and submit jobs, but should not run long computations itself.

Step 2 — give each worker its own stack and TLS

Each instance starts with the stack pointer the linker chose — the same for all, which would make them overwrite each other’s stacks. Before running anything, a worker must allocate its own stack region and point its stack pointer there. Toolchains export helpers for this: wasm-bindgen-rayon and Emscripten do it internally; with a hand-written runtime, export functions that allocate a stack and TLS block from the shared heap and set the globals:

// pool-worker.js
self.onmessage = async ({ data }) => {
  if (data.type !== "init") return;
  const instance = await WebAssembly.instantiate(data.module, { env: { memory: data.memory } });
  const ex = instance.exports;
  const stackTop = ex.thread_alloc_stack(64 * 1024);     // returns top of a new 64 KB stack
  ex.thread_set_stack_pointer(stackTop);
  ex.__wasm_init_tls(ex.thread_alloc_tls());              // TLS block for thread_local! / _Thread_local
  self.postMessage("ready");
  workLoop(ex, data.id);
};

__wasm_init_tls is generated by wasm-ld for modules with thread-local storage; __tls_size and __tls_align give the block’s requirements. Data segments must be initialised exactly once — by the first instance — which the linker arranges through a start function guarded by an atomic flag when memory is shared. The linking details are in linking C and Rust objects into one module.

Step 3 — a job queue in shared memory

The simplest effective queue is a fixed array of job slots plus two counters — submitted and claimed — in shared memory. The main thread writes a job’s parameters into the next slot and increments submitted; idle workers claim jobs by atomically incrementing claimed:

#[no_mangle]
pub extern "C" fn worker_loop() {
    loop {
        let job = loop {
            let c = CLAIMED.load(SeqCst);
            if c < SUBMITTED.load(SeqCst) {
                if CLAIMED.compare_exchange(c, c + 1, SeqCst, SeqCst).is_ok() { break c; }
            } else {
                if SHUTDOWN.load(SeqCst) { return; }
                unsafe { core::arch::wasm32::memory_atomic_wait32(SUBMITTED.as_ptr() as *mut i32, SUBMITTED.load(SeqCst) as i32, -1); }
            }
        };
        run_job(&JOBS[job as usize % JOBS.len()]);
        DONE.fetch_add(1, SeqCst);
        unsafe { core::arch::wasm32::memory_atomic_notify(DONE.as_ptr() as *mut i32, 1); }
    }
}

compare_exchange ensures each job is claimed by exactly one worker. Idle workers sleep in memory.atomic.wait32 on the submitted counter instead of spinning, and the submitter wakes them with memory.atomic.notify after publishing a job. Completed-job counting with a notify lets the submitter — on the main thread, via Atomics.waitAsync — learn when a batch is finished, as in using Atomics.waitAsync on the main thread.

Submitting a job and having a worker run it The submitter writes the job into a slot and increments the submitted counter, then notifies. A sleeping worker wakes, claims the job with compare-and-exchange on the claimed counter, runs it, increments the done counter and notifies anyone waiting for completion. write job slot parameters in shared memory submitted += 1 then notify worker wakes from memory.atomic.wait32 CAS claimed exactly one worker wins run the job done += 1, notify

Step 4 — size the pool

Use at most navigator.hardwareConcurrency − 1 workers so the main thread keeps a core, and fewer on phones. Memory is a second limit: every worker needs a stack and TLS block in the shared memory, plus its own JavaScript heap. Start the pool once and keep it — creating workers and instantiating takes tens of milliseconds each — and make it lazy if parallel work is rare on the page.

Step 5 — shut down cleanly

Set a shutdown flag in shared memory, notify all waiters so sleeping workers wake and see it, then wait for them to acknowledge before terminating:

Atomics.store(shutdownFlag, 0, 1);
Atomics.notify(submittedCounter, 0);                 // wake every sleeping worker
await Promise.all(workers.map((w) => once(w, "message")));   // each posts "stopped" as it exits its loop
workers.forEach((w) => w.terminate());

Terminating without the handshake works too, but a worker killed in the middle of a job may hold a lock in shared memory, leaving the module unusable for any remaining threads.

Load balancing and job granularity

A shared queue balances load automatically: whichever worker is free claims the next job, so uneven job sizes are absorbed as long as there are more jobs than workers. Granularity is the main tuning knob. Very small jobs make the queue’s atomic operations and wake-ups dominate; very large jobs leave workers idle at the end of a batch while one straggler finishes. A good starting point is four to eight jobs per worker per batch, each taking at least a millisecond. For image work, that means tiles of a few hundred thousand pixels rather than rows; for simulations, chunks of a few thousand elements. Work stealing — per-worker queues where idle workers take from busy ones, as Rayon does — handles recursive and nested parallelism better than a single shared queue, but adds complexity that flat batches of independent jobs do not need. Measure the batch’s wall time against the sum of job times divided by the worker count; the difference is the overhead and imbalance the design leaves on the table.

Expected output

A pool of 7 workers processes a batch of 64 image tiles in parallel; the Performance panel shows seven busy worker tracks; the main thread receives a completion signal via Atomics.waitAsync; and shutdown terminates all workers with no errors in the console.

Gotchas

  • All workers sharing one stack. Without per-thread stacks they corrupt each other’s frames. Allocate a stack per worker before running code.
  • Data segments initialised twice. Re-initialising wipes shared data. Let only the first instance initialise memory.
  • Spinning instead of waiting. Idle loops waste cores and battery. Sleep in memory.atomic.wait32.
  • Blocking on the main thread. memory.atomic.wait32 traps there. Submit from the main thread, wait with waitAsync.
  • Forgetting the maximum memory. Shared memories must declare one, and every stack comes out of it. Budget for N stacks.
  • Terminating mid-job. Locks may stay held. Use a shutdown handshake.

Performance note

On an 8-core laptop, processing 64 tiles of a 24-megapixel image took 402 ms on one thread and 71 ms with a 7-worker pool. Waking a sleeping worker via memory.atomic.notify took about 20–50 µs. Pool start-up — 7 workers instantiating a 1.2 MB module from a posted Module — took 38 ms once.

Processing 64 image tiles with a hand-built pool Milliseconds to process sixty-four image tiles with a single thread and with pools of two, four and seven workers sharing one module and memory. ms per batch 1 thread 402 ms 2 workers 207 ms 4 workers 109 ms 7 workers 71 ms

Frequently Asked Questions

Should I build my own pool or use a library? Use Rayon with wasm-bindgen-rayon, or Emscripten’s pthreads, unless you need control they do not offer. Building one is the best way to understand them.

Can each worker have its own non-shared memory? Yes — that is a pool of independent instances, which needs no isolation but cannot share data without copying.

How do I pass large inputs to a job? Write them into shared memory once and put only offsets and lengths in the job slot.

What happens if a job traps? That worker’s call aborts and shared state may be inconsistent; treat the pool as broken and restart it.

Can workers allocate from the shared heap at the same time? Only if the allocator is thread-safe. Threaded builds of dlmalloc and Rust’s allocator use a lock; a custom allocator needs one too.

Why not use postMessage for jobs? You can, and it needs no shared queue — but each job is then delivered through the worker’s event loop, which is slower and cannot be claimed by whichever worker is free.

← Back to SharedArrayBuffer, Atomics & Threading