Running Wasm Off the Event Loop in Node.js
This page answers one task: a Node.js service calls a WebAssembly function — image resizing, PDF rendering, compression, parsing — and while it runs, every other request waits. You want the Wasm work on other threads so the event loop keeps serving requests, without paying for repeated compilation or large copies.
Prerequisites
- [ ] A Node.js service (Express, Fastify, or plain
http) calling a Wasm module. - [ ] A Wasm call that takes long enough to matter (more than a few milliseconds).
- [ ] Node 18+ and optionally
piscinafor pooling.
Why Wasm calls block Node
Node runs JavaScript — and WebAssembly called from it — on a single main thread with an event loop. Asynchronous I/O keeps the loop free because the work happens elsewhere (the OS, libuv’s thread pool) while JavaScript waits for callbacks. A call into WebAssembly is not I/O: it runs on the main thread from start to finish. A 150 ms image resize stops the loop for 150 ms: no new connections accepted, no responses written, no timers fired. Under load, those pauses stack up into tail latency far worse than the work itself, and health checks may time out.
Wrapping the call in a promise does not help — the call is still synchronous. The remedy is to run the call on another thread with worker_threads,
where it blocks only that worker. The main thread posts the input, continues serving requests, and receives the result as a message.
Step 1 — measure event-loop delay first
Confirm the problem and get a baseline with Node’s built-in histogram:
import { monitorEventLoopDelay } from "node:perf_hooks";
const h = monitorEventLoopDelay({ resolution: 10 });
h.enable();
setInterval(() => {
console.log(`loop delay p99: ${(h.percentile(99) / 1e6).toFixed(1)} ms`);
h.reset();
}, 5000);
If p99 delay tracks the duration of Wasm calls under load, the module is blocking the loop.
Step 2 — run the module in a worker pool
piscina manages a pool of workers and a task queue. Compile the module once on the main thread and pass the compiled WebAssembly.Module to every worker,
so each instantiates without recompiling:
// server.mjs
import Piscina from "piscina";
import { readFile } from "node:fs/promises";
const module = await WebAssembly.compile(await readFile(new URL("./resize.wasm", import.meta.url)));
const pool = new Piscina({
filename: new URL("./resize-worker.mjs", import.meta.url).href,
workerData: { module },
maxThreads: Math.max(1, (await import("node:os")).availableParallelism() - 1),
});
app.post("/resize", async (req, reply) => {
const input = await req.arrayBuffer();
const out = await pool.run({ input, width: 800 }, { transferList: [input] });
reply.type("image/jpeg").send(Buffer.from(out));
});
// resize-worker.mjs
import { workerData } from "node:worker_threads";
const instance = await WebAssembly.instantiate(workerData.module, imports);
export default function ({ input, width }) {
return resizeWithWasm(instance, new Uint8Array(input), width); // returns an ArrayBuffer
}
Each worker holds its own instance for its lifetime; tasks reuse it.
Step 3 — transfer buffers instead of copying them
Posting an ArrayBuffer to a worker copies it by default. For large inputs and outputs, list buffers in transferList — ownership moves to the other
thread with no copy, and the sender’s buffer becomes detached. Return results the same way: piscina supports Piscina.move(buffer) in the worker to
transfer the result back. For a 10 MB image, transferring instead of copying saves several milliseconds per request and avoids doubling peak memory.
Step 4 — size the pool
Use os.availableParallelism(), which respects container CPU limits, and leave one core for the main thread. More workers than cores adds context
switching without throughput. Each worker holds its own Wasm memory, so memory use scales with pool size: a module that grows to 200 MB per instance needs
1.6 GB for eight workers. Set maxQueue so overload produces fast 503 responses rather than an unbounded queue, and set idleTimeout to release workers
after quiet periods if memory matters more than warm-up.
Step 5 — verify the effect
Load-test with a mix of heavy and light requests and compare event-loop delay and latency for the light ones before and after. Light endpoints — health checks, metadata lookups — should show latency independent of heavy work. Heavy requests themselves gain a little overhead (posting and transferring), but throughput rises because work runs in parallel.
Small calls do not belong in workers
Posting to a worker and back costs on the order of 0.1–0.3 ms. A Wasm function that runs in 50 µs is faster on the main thread, and moving it to a pool makes things worse. Batch small calls into one task, or keep them on the main thread. A reasonable rule: move calls that take more than a few milliseconds, or that can take much longer for some inputs.
Startup and warm-up
Workers start lazily in piscina by default; the first requests then pay worker startup and instantiation, tens of milliseconds each. Set minThreads so
workers start with the server, and run a warm-up task per worker at startup so engine tier-up happens before real traffic. For serverless environments
with frequent cold starts, the trade-offs differ — see
using Wasm in serverless Node functions.
Error handling and timeouts
A trap in a worker rejects that task’s promise with the error; the worker usually survives, but if the module’s state might be corrupted, discard the
instance — piscina can recycle a worker by throwing from the task with a flag that your worker code handles by exiting. Add a timeout per task
(AbortSignal.timeout passed as signal to pool.run) so a pathological input cannot hold a worker forever, and terminate workers that exceed it.
Streaming large inputs through a worker
Uploads of hundreds of megabytes should not be buffered whole on the main thread before being transferred. Node streams can be piped across threads: a
MessageChannel port or a transferable ReadableStream (Node supports transferring web streams to workers) lets the main thread forward chunks as they
arrive while the worker feeds them into the Wasm module incrementally — a streaming decoder, a hashing function, a compressor. Memory then stays bounded by
the chunk size and the module’s working set rather than by the upload size, and the first bytes of output can be sent back before the input has finished
arriving. Design the Wasm API for this from the start — push(chunk) and finish() functions rather than one process(all_bytes) call — because
retrofitting incremental processing into a module that assumes the whole input is available is much harder than designing for it.
Observability for pooled Wasm work
Once work moves off the main thread, the usual request traces lose detail: the time appears as a single await. Record per-task metrics in the worker — queue wait time (from submission to start), execution time, input size, and the instance’s memory size after the task — and report them alongside request metrics. Queue wait time is the early indicator of an undersized pool; memory size per instance reveals inputs that make the module grow; execution time by input size shows whether the algorithm scales as expected. piscina exposes queue and utilisation statistics directly, which can be exported to the service’s metrics system.
Expected output
Under a load of 50 resize requests per second mixed with 500 light requests per second, p99 latency for light requests drops from 310 ms to 6 ms; event-loop p99 delay stays under 5 ms; resize throughput rises 5× on an 8-core machine; inputs and outputs are transferred, not copied; and overload returns 503 once the queue limit is reached.
Gotchas
- Wrapping sync Wasm calls in promises. Still blocks. Use workers.
- Recompiling per worker. Slow startup. Pass a compiled
WebAssembly.Module. - Copying large buffers. Use
transferListandPiscina.move. - Unbounded queues. Overload turns into memory growth and timeouts. Set
maxQueue. - Moving tiny calls to workers. Messaging costs more than the work.
- No queue-wait metrics. An undersized pool shows up only as slow requests. Record wait and execution time per task.
Performance note
For a 150 ms resize, the worker round trip added 0.4 ms; light-request p99 latency under mixed load fell by 98%.
Frequently Asked Questions
Can libuv’s thread pool run Wasm?
No — it runs native work for Node’s built-in modules; use worker_threads for Wasm.
Is piscina required? No — a hand-written pool works; piscina handles queueing, timeouts and recycling.
Do workers share the compiled code?
Passing the WebAssembly.Module shares compilation; each worker still has its own instance and memory.
What about threads inside the module? Wasm threads with shared memory are a separate technique, covered in the threading guides.
How do I process very large uploads without buffering them? Stream chunks to the worker through a transferable stream or message port and feed them to an incremental Wasm API.
What does a growing queue wait time tell me? That the pool is undersized for the load or that some inputs take far longer than average; check execution times by input size.
Related
- Using Wasm threads in Node.js worker_threads — shared-memory parallelism.
- Caching compiled Wasm in Node.js — faster startup.
- Replacing a native Node addon with Wasm — the trade-offs.
- Timing out slow Wasm calls — deadlines.
← Back to Wasm in Node.js, Deno & Bun