Building a Wasm Thread Pool
This page answers one task: you want several Web Workers to run functions of the same WebAssembly module in parallel, sharing its memory, without Rayon or Emscripten’s pthreads doing it for you — to understand the mechanism or to build a pool tailored to your workload.
Prerequisites
- [ ] A module built with shared memory and atomics (
+atomics,+bulk-memory, a declared maximum memory, imported shared memory). - [ ] A cross-origin isolated page.
- [ ] Familiarity with sharing memory between Wasm and Web Workers.
What a thread is in WebAssembly
There is no thread-creation instruction in WebAssembly. A “thread” is a Web Worker running its own instance of the module, created from the same
compiled WebAssembly.Module and given the same shared WebAssembly.Memory as an import. Instances share memory but not globals: each has its own stack
pointer global and its own thread-local-storage base, so each needs its own region of memory for its stack and TLS block. Code, heap and static data are
shared.
A pool therefore has four jobs. Start N workers and instantiate the module in each with the shared memory. Give each worker its own stack and TLS area before it runs any code. Provide a way to hand work to idle workers and wake them. And shut everything down cleanly. Libraries do all four; doing them by hand shows what those libraries are taking care of.
Step 1 — compile once, create the shared memory, start workers
const module = await WebAssembly.compileStreaming(fetch("compute.wasm"));
const memory = new WebAssembly.Memory({ initial: 256, maximum: 16384, shared: true });
const workers = [];
for (let i = 0; i < N; i++) {
const w = new Worker(new URL("./pool-worker.js", import.meta.url), { type: "module" });
w.postMessage({ type: "init", module, memory, id: i });
workers.push(w);
}
await Promise.all(workers.map((w) => new Promise((r) => w.addEventListener("message", r, { once: true }))));
Posting the compiled WebAssembly.Module avoids each worker compiling it again; posting the shared Memory gives all of them the same buffer. The main
thread usually instantiates the module too, to allocate memory and submit jobs, but should not run long computations itself.
Step 2 — give each worker its own stack and TLS
Each instance starts with the stack pointer the linker chose — the same for all, which would make them overwrite each other’s stacks. Before running anything, a worker must allocate its own stack region and point its stack pointer there. Toolchains export helpers for this: wasm-bindgen-rayon and Emscripten do it internally; with a hand-written runtime, export functions that allocate a stack and TLS block from the shared heap and set the globals:
// pool-worker.js
self.onmessage = async ({ data }) => {
if (data.type !== "init") return;
const instance = await WebAssembly.instantiate(data.module, { env: { memory: data.memory } });
const ex = instance.exports;
const stackTop = ex.thread_alloc_stack(64 * 1024); // returns top of a new 64 KB stack
ex.thread_set_stack_pointer(stackTop);
ex.__wasm_init_tls(ex.thread_alloc_tls()); // TLS block for thread_local! / _Thread_local
self.postMessage("ready");
workLoop(ex, data.id);
};
__wasm_init_tls is generated by wasm-ld for modules with thread-local storage; __tls_size and __tls_align give the block’s requirements. Data
segments must be initialised exactly once — by the first instance — which the linker arranges through a start function guarded by an atomic flag when
memory is shared. The linking details are in
linking C and Rust objects into one module.
Step 3 — a job queue in shared memory
The simplest effective queue is a fixed array of job slots plus two counters — submitted and claimed — in shared memory. The main thread writes a job’s
parameters into the next slot and increments submitted; idle workers claim jobs by atomically incrementing claimed:
#[no_mangle]
pub extern "C" fn worker_loop() {
loop {
let job = loop {
let c = CLAIMED.load(SeqCst);
if c < SUBMITTED.load(SeqCst) {
if CLAIMED.compare_exchange(c, c + 1, SeqCst, SeqCst).is_ok() { break c; }
} else {
if SHUTDOWN.load(SeqCst) { return; }
unsafe { core::arch::wasm32::memory_atomic_wait32(SUBMITTED.as_ptr() as *mut i32, SUBMITTED.load(SeqCst) as i32, -1); }
}
};
run_job(&JOBS[job as usize % JOBS.len()]);
DONE.fetch_add(1, SeqCst);
unsafe { core::arch::wasm32::memory_atomic_notify(DONE.as_ptr() as *mut i32, 1); }
}
}
compare_exchange ensures each job is claimed by exactly one worker. Idle workers sleep in memory.atomic.wait32 on the submitted counter instead of
spinning, and the submitter wakes them with memory.atomic.notify after publishing a job. Completed-job counting with a notify lets the submitter — on
the main thread, via Atomics.waitAsync — learn when a batch is finished, as in
using Atomics.waitAsync on the main thread.
Step 4 — size the pool
Use at most navigator.hardwareConcurrency − 1 workers so the main thread keeps a core, and fewer on phones. Memory is a second limit: every worker
needs a stack and TLS block in the shared memory, plus its own JavaScript heap. Start the pool once and keep it — creating workers and instantiating takes
tens of milliseconds each — and make it lazy if parallel work is rare on the page.
Step 5 — shut down cleanly
Set a shutdown flag in shared memory, notify all waiters so sleeping workers wake and see it, then wait for them to acknowledge before terminating:
Atomics.store(shutdownFlag, 0, 1);
Atomics.notify(submittedCounter, 0); // wake every sleeping worker
await Promise.all(workers.map((w) => once(w, "message"))); // each posts "stopped" as it exits its loop
workers.forEach((w) => w.terminate());
Terminating without the handshake works too, but a worker killed in the middle of a job may hold a lock in shared memory, leaving the module unusable for any remaining threads.
Load balancing and job granularity
A shared queue balances load automatically: whichever worker is free claims the next job, so uneven job sizes are absorbed as long as there are more jobs than workers. Granularity is the main tuning knob. Very small jobs make the queue’s atomic operations and wake-ups dominate; very large jobs leave workers idle at the end of a batch while one straggler finishes. A good starting point is four to eight jobs per worker per batch, each taking at least a millisecond. For image work, that means tiles of a few hundred thousand pixels rather than rows; for simulations, chunks of a few thousand elements. Work stealing — per-worker queues where idle workers take from busy ones, as Rayon does — handles recursive and nested parallelism better than a single shared queue, but adds complexity that flat batches of independent jobs do not need. Measure the batch’s wall time against the sum of job times divided by the worker count; the difference is the overhead and imbalance the design leaves on the table.
Expected output
A pool of 7 workers processes a batch of 64 image tiles in parallel; the Performance panel shows seven busy worker tracks; the main thread receives a
completion signal via Atomics.waitAsync; and shutdown terminates all workers with no errors in the console.
Gotchas
- All workers sharing one stack. Without per-thread stacks they corrupt each other’s frames. Allocate a stack per worker before running code.
- Data segments initialised twice. Re-initialising wipes shared data. Let only the first instance initialise memory.
- Spinning instead of waiting. Idle loops waste cores and battery. Sleep in
memory.atomic.wait32. - Blocking on the main thread.
memory.atomic.wait32traps there. Submit from the main thread, wait withwaitAsync. - Forgetting the maximum memory. Shared memories must declare one, and every stack comes out of it. Budget for N stacks.
- Terminating mid-job. Locks may stay held. Use a shutdown handshake.
Performance note
On an 8-core laptop, processing 64 tiles of a 24-megapixel image took 402 ms on one thread and 71 ms with a 7-worker pool. Waking a sleeping worker via
memory.atomic.notify took about 20–50 µs. Pool start-up — 7 workers instantiating a 1.2 MB module from a posted Module — took 38 ms once.
Frequently Asked Questions
Should I build my own pool or use a library? Use Rayon with wasm-bindgen-rayon, or Emscripten’s pthreads, unless you need control they do not offer. Building one is the best way to understand them.
Can each worker have its own non-shared memory? Yes — that is a pool of independent instances, which needs no isolation but cannot share data without copying.
How do I pass large inputs to a job? Write them into shared memory once and put only offsets and lengths in the job slot.
What happens if a job traps? That worker’s call aborts and shared state may be inconsistent; treat the pool as broken and restart it.
Can workers allocate from the shared heap at the same time? Only if the allocator is thread-safe. Threaded builds of dlmalloc and Rust’s allocator use a lock; a custom allocator needs one too.
Why not use postMessage for jobs?
You can, and it needs no shared queue — but each job is then delivered through the worker’s event loop, which is slower and cannot be claimed by
whichever worker is free.
Related
- Building Rust Wasm with threads using Rayon — the library route.
- Avoiding data races in shared Wasm memory — correctness inside jobs.
- Instantiating one module many times — compile once, instantiate per worker.
- Wrapping a Wasm worker with Comlink — pools without shared memory.
← Back to SharedArrayBuffer, Atomics & Threading