Using Wasm in Serverless Node Functions

This page answers one task: a Node.js serverless function — AWS Lambda, Google Cloud Functions, Azure Functions, Vercel or Netlify functions — needs a WebAssembly module, and you hit the usual problems: the .wasm file is missing after bundling, cold starts are slow, or the function runs out of memory. You want a setup that packages the module correctly and keeps both cold and warm invocations fast.

Prerequisites

  • [ ] A Node.js function (handler) and a deployment tool (SAM, Serverless Framework, CDK, or the platform’s CLI).
  • [ ] A bundler such as esbuild, if the project bundles code.
  • [ ] Access to the platform’s metrics for duration and memory.

The serverless execution model

A serverless function runs in an execution environment that the platform creates on demand. The first request to a new environment is a cold start: the platform starts a runtime, loads your code and runs module-level initialisation before calling the handler. Later requests to the same environment are warm: module-level state is still there, and only the handler runs. Environments are reused for a while and then discarded; under load, the platform runs many in parallel, each with its own cold start.

For WebAssembly, this means: compile and instantiate at module scope, so the cost is paid once per environment rather than once per request; keep the module small, because cold starts include reading and compiling it; and size memory for the module’s linear memory plus Node’s own heap, because both count against the function’s memory limit.

Cold and warm invocations of a function using Wasm On a cold start the runtime starts, the code bundle loads, the Wasm module is read and compiled and instantiated at module scope, and the first handler call runs. Warm invocations skip all initialisation and run only the handler with the already instantiated module. 0 ms environment starts 180 ms runtime + bundle loaded 310 ms Wasm compiled + instantiated 390 ms first handler done 560 ms warm request: 25 ms

Step 1 — make sure the .wasm file ships

Bundlers follow import statements; they do not know that a glue file reads app_bg.wasm from disk at runtime. If the file is not copied into the deployment package, the function fails with ENOENT only after deployment. Copy it explicitly and load it relative to the bundled code:

// esbuild.config.mjs
import { build } from "esbuild";
import { cp } from "node:fs/promises";
await build({ entryPoints: ["src/handler.mjs"], bundle: true, platform: "node", format: "esm", outfile: "dist/handler.mjs" });
await cp("node_modules/@acme/codec/pkg/codec_bg.wasm", "dist/codec_bg.wasm");
// src/handler.mjs
import { readFile } from "node:fs/promises";
const bytes = await readFile(new URL("./codec_bg.wasm", import.meta.url));

Alternatively, use esbuild’s binary loader to inline the module into the bundle as bytes — simplest for small modules, but it increases the JavaScript bundle that must be parsed on every cold start. For packages using --target nodejs glue from wasm-pack, check that the path the glue uses survives bundling, or bypass the glue’s loader and pass bytes yourself.

Step 2 — compile at module scope

// src/handler.mjs
import { readFile } from "node:fs/promises";
import { initSync, resize } from "./codec.js";

const bytes = await readFile(new URL("./codec_bg.wasm", import.meta.url));
initSync({ module: new WebAssembly.Module(bytes) });     // once per environment

export const handler = async (event) => {
  const input = Buffer.from(event.body, "base64");
  const out = resize(input, 800);
  return { statusCode: 200, headers: { "Content-Type": "image/jpeg" }, isBase64Encoded: true, body: Buffer.from(out).toString("base64") };
};

Top-level await in ES module handlers lets initialisation finish before the first invocation. Never compile inside the handler.

Step 3 — size memory for linear memory and the JS heap

The function’s memory setting caps the whole process: Node’s heap, buffers, and every WebAssembly memory. A module that grows its linear memory to 600 MB for large inputs fails in a 512 MB function with an out-of-memory error or a killed process. Measure the module’s peak memory.buffer.byteLength for the largest inputs you accept, add Node’s baseline (often 60–100 MB) and buffers for request and response, and choose the memory setting with headroom. On Lambda, memory also determines CPU share, so a larger setting often makes CPU-bound Wasm work faster and can cost less per request overall.

Where function memory goes The function's memory limit covers Node's runtime baseline, the JavaScript heap, request and response buffers, and the module's linear memory, which only grows within an environment's lifetime. Size the limit for the largest accepted input. consumer typical size notes Node runtime baseline 60–100 MB fixed JS heap + buffers ≈ 2–3× payload base64 doubles size Wasm linear memory peak for largest input never shrinks while warm headroom 20–30% avoids OOM kills

Step 4 — watch payload limits

Platforms limit request and response sizes (for example, 6 MB for synchronous Lambda invocations through API Gateway), and binary payloads are often base64-encoded, adding a third. A Wasm image or document processor can easily exceed these. For large inputs and outputs, pass object-storage references instead: the client uploads to a bucket, the function reads from and writes to the bucket, and returns a URL. Streaming responses, where the platform supports them, also help with large outputs.

Step 5 — measure cold and warm latency

Measure both in the platform’s metrics: Lambda reports “Init Duration” for cold starts separately from “Duration”. Typical findings: module-scope compilation adds tens to hundreds of milliseconds to init, warm invocations are fast, and cold starts are rare under steady load but frequent under bursty traffic. If cold starts matter, shrink the module, reduce initialisation work (snapshot it with Wizer), and consider provisioned concurrency or the platform’s snapshot features (such as Lambda SnapStart for supported runtimes), checking their current support for Node.

Linear memory across warm invocations

Because environments are reused, linear memory persists between invocations: a request that grew memory to 500 MB leaves the environment at 500 MB for all later requests. That is usually fine, but a leak inside the module grows across invocations until the environment is killed. Reset state between invocations (free per-request objects, clear caches with bounded size), and if memory cannot be bounded, re-instantiate the module from the compiled WebAssembly.Module when it passes a threshold — cheap compared with a cold start.

Edge functions are different

Edge platforms (Cloudflare Workers, Vercel Edge, Deno Deploy) run V8 isolates rather than full Node processes, with different packaging (.wasm imported as a module), smaller memory limits, and much faster cold starts. The same module often runs in both with different glue. See Serverless & Edge Deployment for edge-specific guidance.

Architecture and platform differences

Serverless platforms run on x86-64 and ARM64 (Lambda’s Graviton, for example). A WebAssembly module is architecture-independent, which is one of its advantages over native Node addons here: the same .wasm file works on both, so switching a function to ARM64 for its lower price needs no rebuild of the module. What does differ is speed — engines generate different machine code per architecture — so benchmark warm invocations on both before switching. SIMD-heavy modules benefit on both, since V8 maps Wasm SIMD to SSE/AVX on x86-64 and to NEON on ARM64. If the module uses threads, check that the platform allows worker_threads and that the function has more than one vCPU at the chosen memory setting; on Lambda, CPU share scales with memory, and below roughly 1,769 MB a function gets less than one full vCPU, which makes parallelism pointless.

Local testing that matches production

Many Wasm-in-serverless bugs appear only after deployment because local runs use the unbundled source, where the .wasm file sits in node_modules. Test the bundled artefact locally: run the platform’s local emulator (SAM’s sam local invoke, the Serverless Framework’s offline plugin, or a container built from the platform’s base image) against dist/, not src/. Add a smoke test to CI that invokes the bundled handler once with a small real input; it catches missing files, wrong paths and initialisation errors before they reach production.

Expected output

The deployed function finds codec_bg.wasm next to the bundle; module-scope compilation adds 130 ms to init duration; warm invocations resize a 2 MB image in 25 ms; memory is set to 1,024 MB after measuring a 410 MB peak linear memory; large images go through object storage; and a memory threshold re-instantiates the module before the environment runs out.

Gotchas

  • Missing .wasm after bundling. Copy it explicitly or inline it.
  • Compiling in the handler. Every request pays. Compile at module scope.
  • Memory sized for JavaScript only. Linear memory counts too. Measure the peak.
  • Base64 payload limits. Large binaries fail. Use object storage.
  • State leaking across warm invocations. Reset per request and bound caches.

Performance note

Compiling at module scope instead of per request cut warm-invocation latency from 160 ms to 25 ms; raising memory from 512 MB to 1,024 MB cut warm latency further to 17 ms because the function received more CPU.

Warm invocation latency for a Wasm image resize Milliseconds per warm invocation when compiling the module inside the handler, when compiling at module scope with 512 MB, and when compiling at module scope with 1,024 MB of memory. ms per warm invocation compile per request 160 ms module scope, 512 MB 25 ms module scope, 1024 MB 17 ms

Frequently Asked Questions

Should I use the nodejs or web target from wasm-pack? Either works; the web target with initSync and explicit bytes gives you control over file paths after bundling.

Can I keep the module in a Lambda layer? Yes — layers are mounted under /opt; load the file from there.

Do cold starts compile the whole module? V8 compiles lazily, so validation and baseline compilation dominate; smaller modules start faster.

Is Wasm faster than native Node addons here? Usually somewhat slower, but it avoids building native binaries for the platform’s architecture.

Does switching a function to ARM64 require rebuilding the Wasm module? No — the module is architecture-independent; benchmark both, since engine code generation differs.

← Back to Wasm in Node.js, Deno & Bun