Serverless & Edge Deployment

Edge platforms adopted WebAssembly for one reason: a sandbox that starts in microseconds. A container takes hundreds of milliseconds to cold start and holds tens of megabytes; a WebAssembly instance created from an already-compiled module starts in well under a millisecond and holds whatever its memory requires. That difference is what makes it economical to run untrusted customer code in hundreds of locations, with one instance per request and no long-lived process at all.

Prerequisites

  • [ ] A module built for a server target — wasm32-wasip1, wasm32-wasip2, or a platform-specific target such as Cloudflare Workers’ bindings.
  • [ ] Familiarity with WASI, or a platform SDK that hides it.
  • [ ] A deployment path: wrangler, fastly, spin, or your own runtime host.
  • [ ] Measurements from the browser side of your application, if the same code runs in both.

What belongs at the edge

Proximity is the resource an edge platform sells, and the workloads that benefit are the ones where a round trip to an origin is the dominant cost. Four categories recur.

Request shaping — rewriting, routing, header manipulation, redirects, A/B assignment — is decided in microseconds and saves a full round trip when it prevents one. Personalisation that depends only on the request and a small data store: geolocation, currency, feature flags, locale selection. Validation and authentication, where rejecting a bad request at the edge saves the origin from seeing it at all. And transformation of responses already flowing through: image resizing, HTML rewriting, compression of a streaming body.

What does not belong is anything needing a large dataset, a long-lived connection to a database in one region, or more than a few hundred milliseconds of compute. An edge function that queries a database across an ocean has added a hop rather than removed one, and the platform’s per-request limits will usually stop you before the latency does.

The useful test is whether the function can answer from the request plus something small and replicated. If it can, running it in a hundred locations is a clear win. If it cannot, it belongs closer to its data, and the edge layer’s job is to route to it rather than to be it.

The execution model is different from a browser’s

In a browser, a module is compiled and instantiated once and lives for the page’s lifetime. On an edge platform it is compiled once at deploy time, cached in that form, and instantiated per request — sometimes tens of thousands of times a second on one machine.

That inverts which costs matter. Compilation happens off the request path entirely, so binary size affects deployment rather than latency. Instantiation is on the path, so anything the module does at startup — allocating a large heap, building a lookup table, parsing an embedded configuration — is paid on every single request. A module that takes 40 ms to initialise is fine in a browser and unusable at the edge.

One instance for a page, one instance per request A browser compiles and instantiates once for the lifetime of a page. An edge runtime compiles at deploy time and creates a fresh instance for each request, so per-instantiation work is multiplied by request volume. browser compile + instantiate one instance serves every interaction on the page edge runtime compile at deploy req 1 req 2 req 3 req 4 each with fresh memory Work done at startup is paid once in a browser and per request at the edge — a factor of millions in how much it matters. Move table building and configuration parsing to build time, or make it lazy so only requests that need it pay.

Capabilities, not a machine

A server-side module has no filesystem, no network and no environment unless the host grants them. WASI expresses those grants explicitly: directories are preopened and appear in the module as file descriptors it may use, environment variables are passed in, and outbound network access is a separate capability the host either provides or does not.

# nothing is available that is not named here
wasmtime run --dir=./data::/data --env LOG=info --allow-precompiled service.wasm

Platforms differ in how much of this they expose. Cloudflare Workers supplies its own bindings — key-value stores, object storage, queues — instead of a POSIX-like surface. Fastly Compute exposes backends declared in configuration. A self-hosted Wasmtime embedding gives you complete control and complete responsibility. In every case the principle is identical: the module can reach exactly what was named at instantiation, which is the property that makes multi-tenant hosting viable.

State, and why there is so little of it

An edge function is stateless by construction: a fresh instance per request, no shared memory between instances, and no guarantee that two requests from the same user reach the same machine. Every platform therefore offers external state, and each option has a different latency and consistency story.

Key-value stores replicated to every location give fast reads and eventually consistent writes, which suits configuration, feature flags and cached lookups. Object storage suits larger, less frequently changed artifacts. Coordinated objects — a single addressable instance per key, as several platforms now offer — give strong consistency for a specific entity at the cost of routing every request for that key to one location. And a regional database is the traditional answer for anything transactional, with the latency that implies from distant locations.

// typical shapes: fast replicated read, strongly consistent coordinated write
const flags = await env.CONFIG.get('flags', { type: 'json' });     // replicated KV, ~1 ms
const room  = env.ROOMS.get(env.ROOMS.idFromName(roomId));         // one instance, strong consistency

The design consequence is that caching is not an optimisation here but part of the architecture. A function that reads replicated state and writes rarely runs at edge speed; one that writes on every request is doing a round trip and has lost the advantage that brought it there.

For a compiled module, all of this arrives through host functions or platform bindings rather than through WASI, which is another reason to keep host interaction at the edges of the module and the pure logic in the middle.

Cold start is not one number

“Cold start” on an edge platform means at least three different things, and conflating them makes comparisons meaningless.

There is compilation, which happens at deploy time on these platforms and does not appear in request latency at all. There is instantiation, which is the real per-request cost and is usually tens to hundreds of microseconds for a small module. And there is module load into a fresh process, which happens when a machine has not served your code recently and has to fetch and map the compiled artifact.

The number people quote as a cold start is usually the third, and it is the one that varies most between platforms and with module size. The cold start guide separates them and shows how to measure each.

Limits you will meet

Every platform imposes limits, and they are tighter than a server’s. Knowing the shape of them before designing saves a rewrite.

CPU time per request is the one people meet first — typically measured in milliseconds of actual compute rather than wall-clock, so waiting on a subrequest does not usually count against it. Memory per instance is bounded at tens to low hundreds of megabytes, which rules out loading a large model or a sizeable dataset per request. Compiled artifact size is capped, often between one and a few tens of megabytes after compression. Subrequest counts are limited, so a function that fans out to a dozen services may not fit. And execution is normally single-threaded, so a threaded build is not an option.

typical edge limits (check your platform's current numbers)
  CPU per request      10–50 ms of compute
  memory per instance  128 MB
  artifact size        1–10 MB compressed
  subrequests          50 per request
  concurrency          platform-managed, not yours to tune

The productive response is to treat these as design inputs rather than obstacles. A function that fits comfortably inside them is also a function that is cheap, fast and easy to reason about; one that constantly brushes against them is usually doing work that belongs elsewhere in the system.

Where the same code runs on both sides

The strongest argument for WebAssembly at the edge, for a full-stack team, is that the same compiled logic runs in the browser. Validation, pricing, formatting, parsing and business rules can exist once and execute in both places, with no risk of the two implementations drifting.

That works when the logic is pure — input in, output out, no host interaction. The moment it needs a clock, a random number, a file or a network call, the two environments diverge and you need a thin per-host shim. Designing for that from the start means keeping host interaction at the edges of the module rather than threaded through it, which is good structure regardless.

// the core compiles for both targets unchanged
pub fn validate(order: &Order, rules: &Rules) -> Result<Validated, Vec<Violation>> { … }

#[cfg(target_arch = "wasm32")]
mod browser_shim { /* wasm-bindgen exports */ }

#[cfg(target_os = "wasi")]
mod server_shim { /* WASI entry point */ }

Choosing a runtime

If you are hosting WebAssembly yourself rather than using a platform, the runtime choice is real but not agonising. Wasmtime is the reference-quality implementation with the strongest support for the component model, fuel and epochs. Wasmer emphasises embedding breadth and a package registry. WasmEdge targets cloud-native and edge deployments with extensions for AI inference and networking.

All three execute the same standard, so the decision is about embedding ergonomics in your host language, the resource-control features you need, and which non-standard extensions you are willing to depend on. The comparison page goes through the specifics.

Three places server-side Wasm runs A managed edge platform handles compilation, distribution and limits. A self-hosted runtime embedded in your own service gives full control. A cluster node runs modules as workloads alongside containers. managed edge platform deploy and forget platform bindings, not WASI least control, least work embedded runtime your process, your limits fuel, epochs, custom imports full control, full responsibility cluster workload scheduled like a container far smaller and faster to start fits existing operations

Deployment and rollback

Because the artifact is compiled at deploy time, a deployment is slower than a container image push in one respect and faster in every other: there is nothing to schedule, no image to pull at the edge, and no warm-up. A new version becomes live across every location within seconds to a couple of minutes depending on the platform.

That speed makes gradual rollout both easy and necessary. Deploy to a fraction of traffic, watch error rate and CPU time per request, and promote or roll back from the same dashboard. Rolling back is typically as fast as deploying, which is the property that makes aggressive deployment safe — but only if you can tell quickly that something is wrong, which means the metrics have to exist before the rollout, not after the incident.

Version the module alongside the code that calls it, and include a build identifier the function reports in a response header during rollout. When two versions are live at once, that header is what turns “some requests are failing” into “the new version is failing”, which is the difference between a five-minute rollback and an afternoon.

Observability on a platform you do not own

Debugging server-side WebAssembly is harder than debugging a service you can attach to. There is no process to ssh into, often no filesystem to write to, and stack traces from a compiled module are addresses unless you kept the name section.

Three things make it tractable. Structured logging through whatever the platform provides, with a request identifier on every line, is the primary tool and should be in place before the first deploy. Keeping the custom name section in your release build costs a few kilobytes and turns unreadable traces into function names — a trade worth making on the server even where you strip it for the browser. And a local harness that runs the same module under wasmtime with the same inputs lets you reproduce most failures on your own machine, which is where you want to be debugging.

Four things that change off the browser The binary can be identical, but everything around it differs: who supplies the imports, how long an instance lives, what the limits are, and how failures are seen. the host interface WASI or a platform API instead of JavaScript glue instance lifetime often one request, sometimes thousands resource limits imposed by the platform, and usually much tighter observability logs and metrics rather than a console in front of you Startup cost matters far more here: a page instantiates once, a request handler may instantiate every time. Test against the platform's real limits early — a module that fits locally can exceed a size cap on deploy.

Gotchas and failure modes

  • Startup work on the request path. Table building or configuration parsing at instantiation is multiplied by request volume.
  • Assuming a filesystem. Nothing is available unless preopened, and paths inside the module are the guest paths, not the host’s.
  • Assuming threads. Most edge platforms run single-threaded per request; a threaded build will fail or silently degrade.
  • Large modules on a platform with a size limit. Several platforms cap compiled artifact size; check before designing around a large dependency.
  • Non-deterministic behaviour from the host. Clocks and randomness come from the host and differ between environments, which breaks tests that assumed the browser’s behaviour.
  • Logging without a request identifier. With thousands of concurrent instances, unattributed log lines are noise.

Verifying a deployment

Test the module under a standalone runtime before it reaches a platform, with the same capabilities the platform will grant:

# reproduce the platform's capability set locally
wasmtime run --dir=./fixtures::/data --env MODE=test service.wasm < fixtures/request.json

# and assert the boundaries hold
wasmtime run service.wasm --invoke handle_request < fixtures/hostile.json   # no --dir: must fail cleanly

The second test is the valuable one: running without the capability and asserting a clean error proves that the module fails safe when something is missing, rather than producing wrong output. Deployments that skip this discover the behaviour during an incident, when a binding was misconfigured.

Guides in this topic

Frequently Asked Questions

Is server-side Wasm faster than a container? Not at steady-state compute — native code in a container is faster than the same code in a sandbox. It is dramatically faster to start and far smaller, which is what matters for per-request isolation and for running many tenants on one machine.

Can I use my existing libraries? If they compile to wasm32-wasip1 or wasm32-wasip2, yes. Anything using threads, raw sockets, subprocesses or dynamic loading will need work, and some of it will need a different library.

How do I handle secrets? Through the platform’s own mechanism — environment variables or a secret store binding injected at instantiation. Never compile a secret into the module: a .wasm file is a distributable artifact and the string is plainly visible in it.

Does the component model change any of this? It makes the interface between host and guest typed and generated rather than hand-written, which removes a category of marshalling bugs. The deployment and capability story is unchanged.

How do I test an edge function properly? In three layers. Unit-test the pure logic natively, where it is fastest to iterate. Run the compiled module under a local runtime with the same capability set the platform grants, which catches capability and WASI assumptions. Then run the platform’s own local emulator for binding behaviour, and keep a smoke test that runs against a real deployment, because emulators diverge from production in small ways that matter.

What happens when a request exceeds the CPU limit? The platform terminates the instance and returns an error, usually without your code getting a chance to clean up or log. Design for it: keep per-request work well inside the limit, and treat a rising rate of these terminations as the signal that an input distribution has changed.

Can one artifact target several platforms? The pure logic can. The entry point cannot, because each platform expects a different shape — a fetch handler, a WASI command, a component export. Keep the logic in a library crate and write a thin platform-specific wrapper per target; it is usually fewer than a hundred lines each.

← Back to Production Wasm: Workloads & Deployment