Running Wasm on AWS Lambda

This page answers one task: you want to run WebAssembly code on AWS Lambda — a module compiled from Rust, C or Go that processes requests or events — and need to choose between running it inside a Node.js function and running WASI modules in a custom runtime, then package, size and measure it properly.

Prerequisites

  • [ ] An AWS account with Lambda access and a deployment tool (SAM, CDK, Serverless Framework or Terraform).
  • [ ] A Wasm module: a browser/Node-style module with JavaScript glue, or a WASI module.
  • [ ] Representative events and a way to measure cold and warm durations (CloudWatch metrics and logs).

Two ways to run Wasm on Lambda

Lambda runs functions in managed runtimes (Node.js, Python, Java, .NET, Ruby) or custom runtimes on its provided.al2023 OS-only runtime. That gives two practical options for WebAssembly.

Wasm inside Node.js uses V8’s WebAssembly support: package the .wasm file with the function, instantiate it at module scope, and call it from the handler. It is simple, uses a managed runtime, and suits modules with JavaScript glue (wasm-bindgen, Emscripten).

A custom runtime embedding a Wasm engine — a small native bootstrap binary (in Rust, for example) that embeds Wasmtime, implements the Lambda Runtime API loop, and invokes a WASI module or component per event. It runs WASI modules directly, gives you fuel and memory limits, and avoids a JavaScript layer, at the cost of building and maintaining the bootstrap.

Wasm in Node.js versus a custom Wasmtime runtime on Lambda Running Wasm inside the Node.js runtime is simplest and works with JavaScript glue, using V8 to compile the module. A custom runtime embedding Wasmtime runs WASI modules and components directly with fuel and memory limits and no JavaScript layer, but requires building and maintaining a bootstrap binary. Wasm in the Node.js runtime managed runtime, simple JS glue modules V8 compiles the module easiest start custom runtime + Wasmtime WASI modules and components fuel, memory limits, AOT you maintain the bootstrap WASI-native workloads

Step 1 — Wasm inside a Node.js function

// index.mjs
import { readFile } from "node:fs/promises";
import { initSync, resize } from "./pkg/imaging.js";

const bytes = await readFile(new URL("./pkg/imaging_bg.wasm", import.meta.url));
initSync({ module: new WebAssembly.Module(bytes) });          // once per execution environment

export const handler = async (event) => {
  const input = Buffer.from(event.body, "base64");
  const out = resize(input, Number(event.queryStringParameters?.w ?? 800));
  return { statusCode: 200, isBase64Encoded: true, headers: { "Content-Type": "image/jpeg" }, body: Buffer.from(out).toString("base64") };
};

Package the .wasm file next to the handler (bundlers do not copy it automatically). Packaging and memory considerations for Node functions are covered in using Wasm in serverless Node functions.

Step 2 — a custom runtime with Wasmtime

The bootstrap is a native executable named bootstrap that loops over the Lambda Runtime API: fetch the next event, run the Wasm module, post the response.

// bootstrap (sketch): Rust + lambda_runtime + wasmtime
use lambda_runtime::{run, service_fn, LambdaEvent, Error};
use serde_json::Value;

#[tokio::main]
async fn main() -> Result<(), Error> {
    let engine = wasmtime::Engine::new(wasmtime::Config::new().consume_fuel(true))?;
    let module = wasmtime::Module::deserialize_file(&engine, "/var/task/handler.cwasm")?;   // AOT-compiled at build time
    let linker = build_wasi_linker(&engine)?;
    run(service_fn(|ev: LambdaEvent<Value>| invoke(&engine, &module, &linker, ev))).await
}

Precompile the module at build time (wasmtime compile or Module::serialize) for the Lambda architecture, so cold starts load native code instead of compiling. Each invocation creates a fresh store and instance — microseconds — passing the event through stdin or a component’s exported function, and reading the response from stdout or the return value.

A WASI module behind a custom Lambda runtime At build time the WASI module is AOT-compiled for the target architecture and packaged with a Rust bootstrap that embeds Wasmtime. On cold start the bootstrap deserialises the precompiled module. For each event it creates a fresh store and instance with fuel and memory limits, runs the module, and posts the result to the Runtime API. build: AOT compile .cwasm for arm64 package bootstrap Rust + Wasmtime cold start: deserialize no compilation per event: new instance fuel + memory caps post response Runtime API

Step 3 — choose architecture and memory

Lambda offers x86_64 and ARM64 (Graviton). Wasm modules are architecture-independent, so switching to ARM64 for its lower price is easy in the Node.js approach; in the custom runtime, build the bootstrap and AOT artefacts for the chosen architecture. Memory setting determines CPU share: CPU-bound Wasm work runs faster with more memory, often enough that cost per invocation stays flat or drops. Measure duration across settings (Lambda Power Tuning automates this).

Step 4 — package and deploy

Zip packages suit most functions; container images (up to 10 GB) suit large modules or toolchains. Lambda layers can share a Wasm runtime or common modules across functions. Keep the deployment package small for faster cold starts: strip debug info, optimise the module, and avoid bundling unused files.

Step 5 — measure cold and warm starts

CloudWatch reports Init Duration for cold starts and Duration for every invocation. Compare approaches with the same module: the Node.js runtime’s init includes starting Node and compiling the module with V8 (lazily); the custom runtime’s init includes starting the bootstrap and deserialising precompiled code. For small modules both are fast; for large modules, AOT precompilation in the custom runtime often wins on cold start.

When Lambda fits, and when it does not

Lambda suits event-driven Wasm workloads with moderate duration: request handlers, file processing triggered by S3, queue consumers. Its cold starts, while modest, are slower than Wasm-native platforms that instantiate per request in microseconds; if per-request isolation and sub-millisecond startup matter, edge platforms or Wasm-native hosts (Spin, wasmCloud, Fastly Compute) may suit better. Lambda’s strengths are AWS integration, mature operations and generous limits.

Event sources and payload sizes

Lambda functions receive events from many sources — API Gateway or function URLs for HTTP, S3 for new objects, SQS and Kinesis for queues and streams — and each shapes how the Wasm code receives data. HTTP payloads are limited (6 MB for synchronous invocations) and binary bodies arrive base64-encoded, so image or document processing usually works better from S3 events: the function receives a bucket and key, streams the object with the AWS SDK, passes it to the module in chunks, and writes the result back to S3. Queue consumers receive batches of messages, which map naturally onto a Wasm function that processes a batch per call, amortising boundary costs. Design the module’s interface around the event shape you expect — streaming for objects, batches for queues — rather than a single “process all bytes” function that forces buffering.

Observability on Lambda

Lambda emits invocation metrics automatically, but Wasm-specific information needs instrumentation. Log the module version at init, record durations of the Wasm call separately from total handler time (so SDK and network time are distinguished from compute), and report traps and fuel exhaustion as structured log fields or custom metrics. With the custom runtime, emit these from the bootstrap, which sees every invocation; include the module’s content hash so a regression can be tied to a specific deployment. Distributed tracing through AWS X-Ray or OpenTelemetry works as for any function; add a span around the Wasm call so traces show where time is spent.

Concurrency and instance reuse

Each Lambda execution environment handles one invocation at a time, so a module instance created at init is never used concurrently — no locking needed. Scaling happens by adding environments, each with its own instance and its own cold start; provisioned concurrency keeps a pool warm when cold starts are unacceptable.

Expected output

An image-resizing function runs Wasm inside Node.js 22 on ARM64 with 1,024 MB, with warm durations around 40 ms; a WASI text-processing module runs in a custom runtime with an AOT-compiled module, cold-starting in about 60 ms; both deploy with SAM; and Power Tuning shows 1,024 MB as the cost-optimal setting for the image function.

Gotchas

  • Missing .wasm in the package. Bundlers do not copy it. Include it explicitly.
  • Compiling at cold start in custom runtimes. Slow inits. Precompile for the target architecture.
  • AOT artefacts for the wrong architecture. They fail to load. Build per architecture.
  • Too little memory. CPU is proportional. Measure settings.
  • Instantiating per invocation in Node. Wasted time. Instantiate at module scope.
  • Large binaries through HTTP events. Payload limits and base64 overhead. Use S3 events.

Performance note

For a 3 MB WASI module, cold start (Init Duration) was about 210 ms in the Node.js runtime and about 60 ms in a custom Wasmtime runtime with a precompiled module; warm invocations were similar in both.

Cold start for a 3 MB Wasm module on Lambda (ARM64) Milliseconds of init duration for a function running a 3 MB Wasm module inside the Node.js runtime and in a custom runtime embedding Wasmtime with an AOT-precompiled module. ms init duration Node.js runtime + V8 210 ms custom runtime + precompiled 60 ms

Frequently Asked Questions

Can Lambda run Wasm components? In a custom runtime with Wasmtime’s component support, yes.

Does SnapStart help? SnapStart applies to specific managed runtimes; check current support before relying on it for your runtime.

Is there a ready-made Wasm runtime for Lambda? Community projects exist; evaluate maintenance before depending on one, or build a small bootstrap yourself.

How do I limit untrusted modules? In the custom runtime, use fuel or epoch interruption and memory limits per invocation.

How should large files reach the Wasm code? Through S3 events: stream the object, process it in chunks with the module, and write results back to S3 instead of using HTTP payloads.

Does a Lambda instance need locks around the Wasm instance? No — each execution environment handles one invocation at a time, so the instance is never used concurrently.

How do I see Wasm time separately in traces? Add a span or timer around the Wasm call so compute time is distinguished from SDK and network time.

When is provisioned concurrency worth it? When cold starts are unacceptable for a latency-sensitive route; it keeps a pool of environments initialised.

← Back to Serverless & Edge Deployment