Deploying Wasm to Cloudflare Workers
This guide answers one task: take a compiled WebAssembly module, deploy it as part of a Cloudflare Worker, and call it from the request handler — with the instantiation done once rather than per request.
Prerequisites
- [ ]
wrangler3.60 or later, authenticated against your account. - [ ] A module built for
wasm32-unknown-unknown— Workers is not a WASI host by default. - [ ] Node 18+ for local development with
wrangler dev. - [ ] A module under a few megabytes; the platform caps compiled artifact size.
The module is an import, not a fetch
On Workers you do not fetch the .wasm at runtime. It is part of the deployment bundle and is imported
directly, arriving as an already-compiled WebAssembly.Module.
import wasmModule from './pkg/engine_bg.wasm'; // a WebAssembly.Module, compiled at deploy
let instance; // per-isolate, reused across requests
function getInstance() {
instance ??= new WebAssembly.Instance(wasmModule, { env: hostImports });
return instance;
}
export default {
async fetch(request, env, ctx) {
const inst = getInstance();
const body = new Uint8Array(await request.arrayBuffer());
const out = runModule(inst, body);
return new Response(out, { headers: { 'content-type': 'application/json' } });
},
};
Two consequences follow. Compilation never appears in request latency, because the platform did it at
deploy time. And new WebAssembly.Instance — the synchronous constructor — is permitted here, which it
is not for a large module in a browser, precisely because the module is already compiled.
Configure wrangler
The module needs to be recognised as a WebAssembly module rather than an opaque asset. Modern wrangler
handles .wasm imports directly in module-format Workers.
# wrangler.toml
name = "engine-worker"
main = "src/index.js"
compatibility_date = "2026-09-01"
[[rules]]
type = "CompiledWasm"
globs = ["**/*.wasm"]
fallthrough = true
[vars]
LOG_LEVEL = "info"
If you are using wasm-pack output, build with --target bundler and import the generated JavaScript
glue, which handles the instantiation and gives you typed functions. For a hand-written module with a
small interface, importing the .wasm directly as above is simpler and avoids shipping glue you do not
need.
Building the module for Workers
Workers is not a WASI environment. A module compiled for wasm32-wasip1 will fail to instantiate because
its WASI imports are unsatisfied, which produces a LinkError naming a function you never wrote.
# right target for Workers
cargo build --release --target wasm32-unknown-unknown
# or, with wasm-pack, for a bundler-style import
wasm-pack build --release --target bundler
If your code genuinely needs WASI-style facilities — a clock, randomness, an environment variable — supply them yourself as imports. Both are one line each, and doing so explicitly keeps the module’s capability list honest.
const hostImports = {
env: {
now_ms: () => Date.now(),
random_u32: () => (crypto.getRandomValues(new Uint32Array(1)))[0],
},
};
Reusing the instance safely
An isolate handles many requests, and the instance you create on the first one persists. That is the
behaviour you want for performance and a hazard for correctness: whatever the previous request left in
linear memory is still there.
Two disciplines make it safe. Reset explicitly at the start of each request — a reset() export that
clears the arena costs microseconds and removes an entire class of cross-request bug. And never keep
request-scoped data in module globals that outlive the call, which is easy to do accidentally when
porting code that assumed a process per request.
export default {
async fetch(request, env) {
const inst = getInstance();
inst.exports.reset(); // explicit, cheap, and the thing that keeps tenants apart
return handle(inst, request, env);
},
};
If the module cannot be reset cheaply, the alternative is a fresh instance per request. From an already-compiled module that costs well under a millisecond, which is usually affordable and always safer.
Bindings: where state comes from
Workers gives the runtime no filesystem and no ambient network. Everything external arrives as a binding
on env, declared in configuration and injected per request.
[[kv_namespaces]]
binding = "CONFIG"
id = "…"
[[r2_buckets]]
binding = "ASSETS"
bucket_name = "engine-assets"
const flags = await env.CONFIG.get('flags', { type: 'json' }); // replicated, fast read
const blob = await env.ASSETS.get('table.bin'); // object storage
const table = new Uint8Array(await blob.arrayBuffer());
writeIntoWasm(inst, table); // hand it to the module
The module itself never touches these. It receives bytes through its own memory, which keeps the capability surface at the JavaScript layer where you can review it, and keeps the module portable to a browser or another platform.
Streaming requests and responses
A handler that buffers the entire request body before calling the module works and puts a ceiling on the payload you can accept. For larger bodies, or for a response the client should start receiving immediately, stream through the module instead.
export default {
async fetch(request, env) {
const inst = getInstance();
inst.exports.reset();
const { readable, writable } = new TransformStream();
const writer = writable.getWriter();
(async () => {
const reader = request.body.getReader();
const inView = memView(inst);
for (;;) {
const { value, done } = await reader.read();
if (done) break;
inView.set(value.subarray(0, Math.min(value.length, CHUNK)));
const outLen = inst.exports.feed(Math.min(value.length, CHUNK));
if (outLen > 0) await writer.write(readOut(inst, outLen));
}
const tailLen = inst.exports.finish();
if (tailLen > 0) await writer.write(readOut(inst, tailLen));
await writer.close();
})().catch((e) => writer.abort(e));
return new Response(readable, { headers: { 'content-type': 'application/octet-stream' } });
},
};
Peak memory becomes one chunk rather than the whole body, and time to first byte drops to whatever the
module needs to produce its first output. The module has to expose a streaming interface for this —
feed and finish rather than a single run — which is worth designing in from the start if bodies may
be large.
Note the catch that aborts the writer. Without it, an error inside the async block leaves the response
stream open and the request hangs until the platform times it out, which presents as a mysterious latency
spike rather than as the error it actually is.
Watching CPU time rather than wall clock
The platform’s limit is on compute, not on elapsed time, so waiting on a binding or a subrequest does not count against it. That distinction matters when interpreting your own measurements: a handler that takes 80 ms of wall clock may have used 3 ms of CPU, and only the second number is bounded.
Log both. The platform reports CPU time per request in its own analytics, and comparing it against your own wall-clock measurement tells you immediately whether a slow endpoint is slow because of your module or because of something it is waiting for. The two have completely different fixes, and guessing wrong wastes a day.
Expected output
A local run and a deploy both report the artifact size, which is the number to watch:
npx wrangler dev
# ⎔ Starting local server...
# [wrangler] Ready on http://localhost:8787
curl -s localhost:8787 --data-binary @fixtures/input.json | head -c 120
# {"ok":true,"items":[{"id":1,"score":0.94}, …
npx wrangler deploy
Total Upload: 812.44 KiB / gzip: 301.18 KiB
Uploaded engine-worker (2.41 sec)
Published engine-worker (0.35 sec)
https://engine-worker.example.workers.dev
Watch the gzipped figure against the platform’s limit, and watch it over time — a dependency added without thought can double it in one commit.
Gotchas
LinkError: import object field 'fd_write' is not a Function. The module was built for WASI. Rebuild forwasm32-unknown-unknownor supply the imports yourself.- Instance created per request. Works, but wastes time. Memoise it at module scope.
- State leaking between requests. The isolate reuses the instance. Reset explicitly.
- Exceeding the CPU limit. Workers measures actual compute; a tight loop over a large input will be terminated. Bound the input size you accept.
- Artifact too large. Strip debug info and run
wasm-opt -Oz; the platform’s cap is on the bundle, not just your source. - Assuming
Date.now()advances during a request. Timers are coarsened and may not move within a single synchronous block, which breaks naive timing code.
Performance note
A 300 kB compressed module instantiates in roughly 0.3–0.8 ms from the already-compiled artifact, and subsequent requests on the same isolate pay nothing. A representative transform over a 20 kB payload ran in 1.9 ms of CPU, comfortably inside the platform’s budget. The dominant cost in practice was neither — it was the first request on a cold isolate, where module load and JavaScript initialisation together took about 6 ms, and which is exactly what the cold start page takes apart.
Frequently Asked Questions
Can I use wasm-bindgen output directly?
Yes, with --target bundler. The glue expects a bundler that can handle the .wasm import, which
wrangler’s build pipeline does. For very small interfaces, importing the module directly avoids the glue
entirely.
Is WASI supported? Not as the default environment, though there are shims and the platform’s support continues to evolve. For portable code, keep host interaction behind a small interface and implement it per platform rather than depending on WASI being present.
How do I share one module between the Worker and the browser?
Compile once for wasm32-unknown-unknown and write two thin wrappers: a fetch handler on the Worker and
a wasm-bindgen or hand-written loader in the browser. The core logic is identical, which is much of the
appeal.
Related
- Running Wasm on Fastly Compute — a platform built entirely on Wasm.
- Serving Wasm files with the right headers — the browser-facing side of delivery.
- Sharing validation logic between server and browser — the use case this enables.
← Back to Serverless & Edge Deployment