Deploying Wasm to Cloudflare Workers

This guide answers one task: take a compiled WebAssembly module, deploy it as part of a Cloudflare Worker, and call it from the request handler — with the instantiation done once rather than per request.

Prerequisites

  • [ ] wrangler 3.60 or later, authenticated against your account.
  • [ ] A module built for wasm32-unknown-unknown — Workers is not a WASI host by default.
  • [ ] Node 18+ for local development with wrangler dev.
  • [ ] A module under a few megabytes; the platform caps compiled artifact size.

The module is an import, not a fetch

On Workers you do not fetch the .wasm at runtime. It is part of the deployment bundle and is imported directly, arriving as an already-compiled WebAssembly.Module.

import wasmModule from './pkg/engine_bg.wasm';       // a WebAssembly.Module, compiled at deploy

let instance;                                          // per-isolate, reused across requests
function getInstance() {
  instance ??= new WebAssembly.Instance(wasmModule, { env: hostImports });
  return instance;
}

export default {
  async fetch(request, env, ctx) {
    const inst = getInstance();
    const body = new Uint8Array(await request.arrayBuffer());
    const out = runModule(inst, body);
    return new Response(out, { headers: { 'content-type': 'application/json' } });
  },
};

Two consequences follow. Compilation never appears in request latency, because the platform did it at deploy time. And new WebAssembly.Instance — the synchronous constructor — is permitted here, which it is not for a large module in a browser, precisely because the module is already compiled.

Compile at deploy, instantiate per isolate The module is compiled when the Worker is deployed. Each isolate creates one instance on its first request and reuses it for every subsequent request it handles, so instantiation cost is amortised. wrangler deploy module compiled here isolate starts first request instantiates every later request reuses the instance no compile, no instantiate, no allocation of the heap state persists — which you must account for Reusing the instance is fast and means linear memory carries over between requests handled by the same isolate. Reset any per-request state explicitly, and never leave one request's data where the next can read it.

Configure wrangler

The module needs to be recognised as a WebAssembly module rather than an opaque asset. Modern wrangler handles .wasm imports directly in module-format Workers.

# wrangler.toml
name = "engine-worker"
main = "src/index.js"
compatibility_date = "2026-09-01"

[[rules]]
type   = "CompiledWasm"
globs  = ["**/*.wasm"]
fallthrough = true

[vars]
LOG_LEVEL = "info"

If you are using wasm-pack output, build with --target bundler and import the generated JavaScript glue, which handles the instantiation and gives you typed functions. For a hand-written module with a small interface, importing the .wasm directly as above is simpler and avoids shipping glue you do not need.

Building the module for Workers

Workers is not a WASI environment. A module compiled for wasm32-wasip1 will fail to instantiate because its WASI imports are unsatisfied, which produces a LinkError naming a function you never wrote.

# right target for Workers
cargo build --release --target wasm32-unknown-unknown

# or, with wasm-pack, for a bundler-style import
wasm-pack build --release --target bundler

If your code genuinely needs WASI-style facilities — a clock, randomness, an environment variable — supply them yourself as imports. Both are one line each, and doing so explicitly keeps the module’s capability list honest.

const hostImports = {
  env: {
    now_ms: () => Date.now(),
    random_u32: () => (crypto.getRandomValues(new Uint32Array(1)))[0],
  },
};

Reusing the instance safely

An isolate handles many requests, and the instance you create on the first one persists. That is the behaviour you want for performance and a hazard for correctness: whatever the previous request left in linear memory is still there.

Two disciplines make it safe. Reset explicitly at the start of each request — a reset() export that clears the arena costs microseconds and removes an entire class of cross-request bug. And never keep request-scoped data in module globals that outlive the call, which is easy to do accidentally when porting code that assumed a process per request.

export default {
  async fetch(request, env) {
    const inst = getInstance();
    inst.exports.reset();                     // explicit, cheap, and the thing that keeps tenants apart
    return handle(inst, request, env);
  },
};

If the module cannot be reset cheaply, the alternative is a fresh instance per request. From an already-compiled module that costs well under a millisecond, which is usually affordable and always safer.

Bindings: where state comes from

Workers gives the runtime no filesystem and no ambient network. Everything external arrives as a binding on env, declared in configuration and injected per request.

[[kv_namespaces]]
binding = "CONFIG"
id = "…"

[[r2_buckets]]
binding = "ASSETS"
bucket_name = "engine-assets"
const flags = await env.CONFIG.get('flags', { type: 'json' });      // replicated, fast read
const blob  = await env.ASSETS.get('table.bin');                    // object storage
const table = new Uint8Array(await blob.arrayBuffer());
writeIntoWasm(inst, table);                                          // hand it to the module

The module itself never touches these. It receives bytes through its own memory, which keeps the capability surface at the JavaScript layer where you can review it, and keeps the module portable to a browser or another platform.

The handler mediates everything The request reaches a JavaScript handler which reads bindings for any external state, passes bytes into the module's memory, calls the export and builds a response. The module never touches platform APIs directly. request fetch handler reads bindings writes bytes into memory builds the response KV · R2 · queues Wasm instance pure computation Keeping platform access in the handler is what lets the same module run unchanged in a browser or under a standalone runtime.

Streaming requests and responses

A handler that buffers the entire request body before calling the module works and puts a ceiling on the payload you can accept. For larger bodies, or for a response the client should start receiving immediately, stream through the module instead.

export default {
  async fetch(request, env) {
    const inst = getInstance();
    inst.exports.reset();
    const { readable, writable } = new TransformStream();
    const writer = writable.getWriter();

    (async () => {
      const reader = request.body.getReader();
      const inView = memView(inst);
      for (;;) {
        const { value, done } = await reader.read();
        if (done) break;
        inView.set(value.subarray(0, Math.min(value.length, CHUNK)));
        const outLen = inst.exports.feed(Math.min(value.length, CHUNK));
        if (outLen > 0) await writer.write(readOut(inst, outLen));
      }
      const tailLen = inst.exports.finish();
      if (tailLen > 0) await writer.write(readOut(inst, tailLen));
      await writer.close();
    })().catch((e) => writer.abort(e));

    return new Response(readable, { headers: { 'content-type': 'application/octet-stream' } });
  },
};

Peak memory becomes one chunk rather than the whole body, and time to first byte drops to whatever the module needs to produce its first output. The module has to expose a streaming interface for this — feed and finish rather than a single run — which is worth designing in from the start if bodies may be large.

Note the catch that aborts the writer. Without it, an error inside the async block leaves the response stream open and the request hangs until the platform times it out, which presents as a mysterious latency spike rather than as the error it actually is.

Watching CPU time rather than wall clock

The platform’s limit is on compute, not on elapsed time, so waiting on a binding or a subrequest does not count against it. That distinction matters when interpreting your own measurements: a handler that takes 80 ms of wall clock may have used 3 ms of CPU, and only the second number is bounded.

Log both. The platform reports CPU time per request in its own analytics, and comparing it against your own wall-clock measurement tells you immediately whether a slow endpoint is slow because of your module or because of something it is waiting for. The two have completely different fixes, and guessing wrong wastes a day.

Expected output

A local run and a deploy both report the artifact size, which is the number to watch:

npx wrangler dev
# ⎔ Starting local server...
# [wrangler] Ready on http://localhost:8787

curl -s localhost:8787 --data-binary @fixtures/input.json | head -c 120
# {"ok":true,"items":[{"id":1,"score":0.94}, …
npx wrangler deploy
Total Upload: 812.44 KiB / gzip: 301.18 KiB
Uploaded engine-worker (2.41 sec)
Published engine-worker (0.35 sec)
  https://engine-worker.example.workers.dev

Watch the gzipped figure against the platform’s limit, and watch it over time — a dependency added without thought can double it in one commit.

Where the binary lives at the edge The binary ships as part of the worker bundle and is compiled once per isolate, not once per request, which is why per-request cost is small. worker bundle binary included deployed to edge every location isolate starts compiled once requests served instance reused Module-level instantiation is the pattern: create once outside the handler and reuse across requests. Watch the bundle size limit — the compressed binary counts, and a large model will not fit. Isolates are recycled without warning, so nothing durable may live in module state.

Gotchas

  • LinkError: import object field 'fd_write' is not a Function. The module was built for WASI. Rebuild for wasm32-unknown-unknown or supply the imports yourself.
  • Instance created per request. Works, but wastes time. Memoise it at module scope.
  • State leaking between requests. The isolate reuses the instance. Reset explicitly.
  • Exceeding the CPU limit. Workers measures actual compute; a tight loop over a large input will be terminated. Bound the input size you accept.
  • Artifact too large. Strip debug info and run wasm-opt -Oz; the platform’s cap is on the bundle, not just your source.
  • Assuming Date.now() advances during a request. Timers are coarsened and may not move within a single synchronous block, which breaks naive timing code.

Performance note

A 300 kB compressed module instantiates in roughly 0.3–0.8 ms from the already-compiled artifact, and subsequent requests on the same isolate pay nothing. A representative transform over a 20 kB payload ran in 1.9 ms of CPU, comfortably inside the platform’s budget. The dominant cost in practice was neither — it was the first request on a cold isolate, where module load and JavaScript initialisation together took about 6 ms, and which is exactly what the cold start page takes apart.

Frequently Asked Questions

Can I use wasm-bindgen output directly? Yes, with --target bundler. The glue expects a bundler that can handle the .wasm import, which wrangler’s build pipeline does. For very small interfaces, importing the module directly avoids the glue entirely.

Is WASI supported? Not as the default environment, though there are shims and the platform’s support continues to evolve. For portable code, keep host interaction behind a small interface and implement it per platform rather than depending on WASI being present.

How do I share one module between the Worker and the browser? Compile once for wasm32-unknown-unknown and write two thin wrappers: a fetch handler on the Worker and a wasm-bindgen or hand-written loader in the browser. The core logic is identical, which is much of the appeal.

← Back to Serverless & Edge Deployment