Running Wasm Off the Event Loop in Node.js

This page answers one task: a Node.js service calls a WebAssembly function — image resizing, PDF rendering, compression, parsing — and while it runs, every other request waits. You want the Wasm work on other threads so the event loop keeps serving requests, without paying for repeated compilation or large copies.

Prerequisites

  • [ ] A Node.js service (Express, Fastify, or plain http) calling a Wasm module.
  • [ ] A Wasm call that takes long enough to matter (more than a few milliseconds).
  • [ ] Node 18+ and optionally piscina for pooling.

Why Wasm calls block Node

Node runs JavaScript — and WebAssembly called from it — on a single main thread with an event loop. Asynchronous I/O keeps the loop free because the work happens elsewhere (the OS, libuv’s thread pool) while JavaScript waits for callbacks. A call into WebAssembly is not I/O: it runs on the main thread from start to finish. A 150 ms image resize stops the loop for 150 ms: no new connections accepted, no responses written, no timers fired. Under load, those pauses stack up into tail latency far worse than the work itself, and health checks may time out.

Wrapping the call in a promise does not help — the call is still synchronous. The remedy is to run the call on another thread with worker_threads, where it blocks only that worker. The main thread posts the input, continues serving requests, and receives the result as a message.

Wasm on the main thread versus in a worker pool On the main thread, each Wasm call blocks the event loop for its full duration, so unrelated requests wait and tail latency grows. In a worker pool, calls run on other threads while the main thread keeps handling requests, at the cost of posting inputs and results between threads. main thread event loop blocked per call all requests wait p99 latency explodes only for tiny calls worker_threads pool loop stays free calls run in parallel message passing overhead heavy work

Step 1 — measure event-loop delay first

Confirm the problem and get a baseline with Node’s built-in histogram:

import { monitorEventLoopDelay } from "node:perf_hooks";
const h = monitorEventLoopDelay({ resolution: 10 });
h.enable();
setInterval(() => {
  console.log(`loop delay p99: ${(h.percentile(99) / 1e6).toFixed(1)} ms`);
  h.reset();
}, 5000);

If p99 delay tracks the duration of Wasm calls under load, the module is blocking the loop.

Step 2 — run the module in a worker pool

piscina manages a pool of workers and a task queue. Compile the module once on the main thread and pass the compiled WebAssembly.Module to every worker, so each instantiates without recompiling:

// server.mjs
import Piscina from "piscina";
import { readFile } from "node:fs/promises";

const module = await WebAssembly.compile(await readFile(new URL("./resize.wasm", import.meta.url)));
const pool = new Piscina({
  filename: new URL("./resize-worker.mjs", import.meta.url).href,
  workerData: { module },
  maxThreads: Math.max(1, (await import("node:os")).availableParallelism() - 1),
});

app.post("/resize", async (req, reply) => {
  const input = await req.arrayBuffer();
  const out = await pool.run({ input, width: 800 }, { transferList: [input] });
  reply.type("image/jpeg").send(Buffer.from(out));
});
// resize-worker.mjs
import { workerData } from "node:worker_threads";
const instance = await WebAssembly.instantiate(workerData.module, imports);
export default function ({ input, width }) {
  return resizeWithWasm(instance, new Uint8Array(input), width);   // returns an ArrayBuffer
}

Each worker holds its own instance for its lifetime; tasks reuse it.

Step 3 — transfer buffers instead of copying them

Posting an ArrayBuffer to a worker copies it by default. For large inputs and outputs, list buffers in transferList — ownership moves to the other thread with no copy, and the sender’s buffer becomes detached. Return results the same way: piscina supports Piscina.move(buffer) in the worker to transfer the result back. For a 10 MB image, transferring instead of copying saves several milliseconds per request and avoids doubling peak memory.

A request handled with a Wasm worker pool The main thread receives the upload and transfers the buffer to a pool worker without copying. The worker runs the Wasm resize on its own instance while the main thread serves other requests. The worker transfers the result back, and the main thread sends the response. request arrives upload buffer transfer to worker no copy worker runs Wasm own instance main loop keeps serving other requests result transferred back send response

Step 4 — size the pool

Use os.availableParallelism(), which respects container CPU limits, and leave one core for the main thread. More workers than cores adds context switching without throughput. Each worker holds its own Wasm memory, so memory use scales with pool size: a module that grows to 200 MB per instance needs 1.6 GB for eight workers. Set maxQueue so overload produces fast 503 responses rather than an unbounded queue, and set idleTimeout to release workers after quiet periods if memory matters more than warm-up.

Step 5 — verify the effect

Load-test with a mix of heavy and light requests and compare event-loop delay and latency for the light ones before and after. Light endpoints — health checks, metadata lookups — should show latency independent of heavy work. Heavy requests themselves gain a little overhead (posting and transferring), but throughput rises because work runs in parallel.

Small calls do not belong in workers

Posting to a worker and back costs on the order of 0.1–0.3 ms. A Wasm function that runs in 50 µs is faster on the main thread, and moving it to a pool makes things worse. Batch small calls into one task, or keep them on the main thread. A reasonable rule: move calls that take more than a few milliseconds, or that can take much longer for some inputs.

Startup and warm-up

Workers start lazily in piscina by default; the first requests then pay worker startup and instantiation, tens of milliseconds each. Set minThreads so workers start with the server, and run a warm-up task per worker at startup so engine tier-up happens before real traffic. For serverless environments with frequent cold starts, the trade-offs differ — see using Wasm in serverless Node functions.

Error handling and timeouts

A trap in a worker rejects that task’s promise with the error; the worker usually survives, but if the module’s state might be corrupted, discard the instance — piscina can recycle a worker by throwing from the task with a flag that your worker code handles by exiting. Add a timeout per task (AbortSignal.timeout passed as signal to pool.run) so a pathological input cannot hold a worker forever, and terminate workers that exceed it.

Streaming large inputs through a worker

Uploads of hundreds of megabytes should not be buffered whole on the main thread before being transferred. Node streams can be piped across threads: a MessageChannel port or a transferable ReadableStream (Node supports transferring web streams to workers) lets the main thread forward chunks as they arrive while the worker feeds them into the Wasm module incrementally — a streaming decoder, a hashing function, a compressor. Memory then stays bounded by the chunk size and the module’s working set rather than by the upload size, and the first bytes of output can be sent back before the input has finished arriving. Design the Wasm API for this from the start — push(chunk) and finish() functions rather than one process(all_bytes) call — because retrofitting incremental processing into a module that assumes the whole input is available is much harder than designing for it.

Observability for pooled Wasm work

Once work moves off the main thread, the usual request traces lose detail: the time appears as a single await. Record per-task metrics in the worker — queue wait time (from submission to start), execution time, input size, and the instance’s memory size after the task — and report them alongside request metrics. Queue wait time is the early indicator of an undersized pool; memory size per instance reveals inputs that make the module grow; execution time by input size shows whether the algorithm scales as expected. piscina exposes queue and utilisation statistics directly, which can be exported to the service’s metrics system.

Expected output

Under a load of 50 resize requests per second mixed with 500 light requests per second, p99 latency for light requests drops from 310 ms to 6 ms; event-loop p99 delay stays under 5 ms; resize throughput rises 5× on an 8-core machine; inputs and outputs are transferred, not copied; and overload returns 503 once the queue limit is reached.

Gotchas

  • Wrapping sync Wasm calls in promises. Still blocks. Use workers.
  • Recompiling per worker. Slow startup. Pass a compiled WebAssembly.Module.
  • Copying large buffers. Use transferList and Piscina.move.
  • Unbounded queues. Overload turns into memory growth and timeouts. Set maxQueue.
  • Moving tiny calls to workers. Messaging costs more than the work.
  • No queue-wait metrics. An undersized pool shows up only as slow requests. Record wait and execution time per task.

Performance note

For a 150 ms resize, the worker round trip added 0.4 ms; light-request p99 latency under mixed load fell by 98%.

Light-request p99 latency under mixed load Milliseconds of p99 latency for lightweight requests while the server also handles 50 Wasm resize requests per second, with Wasm on the main thread and in a worker pool. p99 latency (ms) Wasm on main thread 310 ms Wasm in worker pool 6 ms

Frequently Asked Questions

Can libuv’s thread pool run Wasm? No — it runs native work for Node’s built-in modules; use worker_threads for Wasm.

Is piscina required? No — a hand-written pool works; piscina handles queueing, timeouts and recycling.

Do workers share the compiled code? Passing the WebAssembly.Module shares compilation; each worker still has its own instance and memory.

What about threads inside the module? Wasm threads with shared memory are a separate technique, covered in the threading guides.

How do I process very large uploads without buffering them? Stream chunks to the worker through a transferable stream or message port and feed them to an incremental Wasm API.

What does a growing queue wait time tell me? That the pool is undersized for the load or that some inputs take far longer than average; check execution times by input size.

← Back to Wasm in Node.js, Deno & Bun