Production Wasm: Workloads & Deployment

Compiling a module is the easy part. This area is about the workloads that justify the compile — video transcoding, model inference, analytical queries, cryptography, physics, untrusted plugins — and about what changes when that module has to run on somebody else’s phone, inside a serverless request, or on a cluster node with a 5 ms budget. Every page here starts from a shipped-system constraint rather than a tutorial: payload size the user pays for on first load, memory a tab is allowed to hold, latency a request is allowed to add, and the blast radius of code you did not write.

Engineering takeaways

  • Pick the workloads where WebAssembly genuinely wins — dense compute over bytes you already have in memory — and recognise the ones where the boundary crossing eats the speedup.
  • Budget a module the way you budget an image: transfer size, decode/compile time, resident memory and steady-state throughput are four separate numbers, and optimising one routinely worsens another.
  • Run the same compiled artifact in the browser, on an edge runtime and on a server, and know which parts of that story are real portability and which parts need a per-host shim.
  • Treat a third-party module as untrusted code with a capability list, not as a library — the sandbox is the feature you are actually buying.
  • Instrument before you optimise: first byte, compile, instantiate, first result and steady-state rate are all measurable in a few lines, and they usually disagree about where the time went.

The shape of a production Wasm system

A shipped WebAssembly feature is almost never a single module called from a click handler. It is a pipeline: bytes arrive from the network, a .wasm payload is fetched and compiled, a worker takes ownership of the instance so the main thread stays free, and the interesting data — pixels, samples, rows, tensors — is moved into linear memory in the largest chunks the format allows. The module computes, writes results back into the same buffer, and JavaScript reads them out with a typed-array view. The value of the design is that the expensive inner loop never crosses the boundary; the cost is that you own the memory layout on both sides.

Where a shipped Wasm feature spends its life Bytes arrive over the network, are compiled once and cached, then an instance owned by a worker processes data held in linear memory. The page only sends pointers and reads results, so the hot loop never crosses the JavaScript boundary. one-time cost fetch .wasm compressed transfer compile + cache streaming, reusable module instantiate in worker main thread stays free per-item cost copy bytes in once per frame or batch compute in linear memory no boundary crossings here read result view pointer plus length The one-time row is what users feel on first paint; the per-item row is what they feel afterwards. Optimise the two rows separately — shrinking the binary rarely speeds up the loop, and a faster loop never pays back a 6 MB download.

The pipeline above is why the first question about a candidate workload is not “is Wasm faster?” but “how many bytes cross the boundary per unit of work?”. A JPEG decoder that receives a 2 MB buffer and returns a 12 MB pixel surface crosses twice per image and computes for milliseconds in between: an excellent ratio. A string-formatting helper called once per table cell crosses twice per cell and computes for microseconds: a terrible one, and it will lose to plain JavaScript even though the compiled code is faster in isolation. The throughput measurement guide walks the arithmetic; the practical rule is that the work per crossing has to be large enough that the crossing disappears into the noise.

Deciding whether a workload belongs in Wasm

The decision is mechanical once you frame it as a ratio. Estimate the bytes that must cross the boundary for one unit of work, estimate the arithmetic performed on those bytes, and compare. Image convolution over a 4-megapixel surface performs tens of operations per pixel on data that crosses once — a ratio in the hundreds. Parsing a 2 kB JSON document performs a handful of operations per byte and returns a graph of objects that has to be rebuilt on the JavaScript side — a ratio near zero, and a guaranteed regression.

Four properties reliably predict a win:

  1. The data is already bytes. Pixels, samples, matrices, compressed frames, columnar batches. If the input is a JavaScript object graph, you pay to flatten it before the module sees it, and that cost is usually larger than the saving.
  2. The inner loop is arithmetic. Integer and float math, branch-light, cache-friendly. Engines optimise this shape aggressively, and the compiled version keeps its advantage across engines rather than depending on a JIT warming up.
  3. The work is batched. One call that processes ten thousand items beats ten thousand calls that process one item, by orders of magnitude, regardless of what the code inside does.
  4. An implementation already exists in C, C++ or Rust. Porting a mature codec, solver or database engine to JavaScript is months of work and a permanent maintenance tax; compiling it is an afternoon and a size budget.

The inverse list is just as useful. DOM manipulation loses, because every operation is a boundary crossing into JavaScript anyway — measured in detail in is Wasm faster than JavaScript for DOM manipulation. String-heavy transformation usually loses, because UTF-8 encoding and decoding at the boundary can cost more than the transformation. Anything dominated by network waits obviously gains nothing. And “we want the code to be harder to read” is not a performance argument — a minified bundle and a .wasm file are equally reverse-engineerable to anyone who cares.

// a quick, honest triage before anyone writes Rust
const bytesPerUnit = frame.byteLength;          // in + out
const opsPerUnit   = width * height * kernelOps;
const ratio = opsPerUnit / bytesPerUnit;
// ratio > ~20 and batched: strong candidate
// ratio < ~2 or called per item: keep it in JavaScript

Memory is the budget you forget

Transfer size gets all the attention because it is visible in a network panel. Resident memory is what actually kills production features: a tab that holds 400 MB gets killed on a mid-range Android device, an edge runtime rejects an instance that exceeds its per-request limit, and a cluster node running a hundred instances multiplies whatever each one reserves.

A WebAssembly instance’s memory is a single growable buffer measured in 64 KiB pages. Three numbers control it — the initial size, the maximum, and how your allocator behaves in between:

# what the module actually asks for, before it runs a single instruction
wasm-objdump -x pipeline.wasm | grep -A2 "Memory\["
# memory[0] pages: initial=256 max=4096      # 16 MB initial, 256 MB ceiling

Setting initial too low means a burst of memory.grow calls during warm-up, each of which may reallocate and copy the whole heap, and each of which detaches every typed-array view JavaScript is holding. Setting it too high means every instance reserves that much whether it needs it or not. The workable pattern for batch workloads is an arena: allocate one region sized for the largest expected item, reuse it for every item, and never free — covered in implementing a bump allocator in Wasm. For long-lived interactive modules, a real allocator plus explicit lifetime management is worth the code, and memory profiling tells you which one you have.

Three ways to size an instance A fixed arena reserves once and reuses the region for every item. A growing heap starts small and reallocates under load, detaching views each time. A per-request instance is created and discarded, trading instantiation cost for a hard memory ceiling. fixed arena reserved once no grow, no detach views stay valid best for batch pipelines growing heap grows under load every grow detaches views rebuild views after each call instance per request discarded after the response hard ceiling per tenant pay instantiation each time Reuse a compiled module across instances — compilation is the expensive half, and a cached WebAssembly.Module instantiates in well under a millisecond. Whichever strategy you pick, write down the ceiling: an instance with no maximum is a crash waiting for an unusual input.

Getting an artifact to the host that will run it

The deployment surface is broader than “put the file on a CDN”. The same module can end up in four places with four different loading contracts, and the differences are not cosmetic — they decide whether you can use threads, whether compilation is cached between requests, and whether the file is even allowed to grow past a size limit.

# browser: one artifact, served with the right type so streaming compilation engages
curl -sI https://example.com/pipeline.wasm | grep -i content-type
# content-type: application/wasm

# edge/serverless: the module is part of the deployment bundle, compiled at deploy time
wrangler deploy                      # Cloudflare Workers
fastly compute publish               # Fastly Compute

# server / cluster: a WASI binary run by a standalone runtime
wasmtime run --dir=./data pipeline.wasm

Each of those hosts imposes a different budget. A browser cares about transfer size and the time to first result, and rewards caching the compiled module. An edge runtime typically compiles ahead of time and cares about instantiation cost per request, which is where a module that allocates an 80 MB heap at startup becomes visible in your p99. A cluster node running a WASI binary cares about the host calls you make, because every one of them is a syscall the runtime has to broker. The pages under serverless and edge deployment work through each host in turn with the numbers that matter for it.

One artifact, three hosts, three budgets A browser is dominated by transfer and compile time, an edge runtime by per-request instantiation, and a standalone server runtime by host call overhead and memory footprint. The binary is the same; the constraint that decides your design is not. one .wasm artifact browser tab transfer size is the budget compile once, cache it threads need COOP/COEP measure: first result edge runtime compiled at deploy time instance per request startup allocation shows up measure: cold start standalone runtime WASI host calls are syscalls memory is a hard limit capabilities granted, not assumed measure: throughput Portability is real at the instruction level and thin at the edges: imports, clocks, randomness, filesystem and threads all need a per-host answer.

Crossing the boundary with real payloads

Production payloads are rarely four integers. They are frames, sample blocks, row batches and tensors, and the interop question becomes how to get megabytes in and out without a copy you did not intend. The answer is nearly always the same shape: allocate inside the module, hand JavaScript a pointer, and fill that region through a typed-array view — the pattern covered in reading linear memory with typed arrays.

// allocate inside the module, fill from JS, process in place, read back
const ptr = mod.exports.alloc(frame.byteLength);
new Uint8Array(mod.exports.memory.buffer, ptr, frame.byteLength).set(frame);
const outLen = mod.exports.process(ptr, frame.byteLength);
const out = new Uint8Array(mod.exports.memory.buffer, ptr, outLen);   // rebuild after every call

Two details in those four lines cause most production bugs. The view is rebuilt after the call because any allocation inside process may have grown memory and detached the old buffer — the failure mode explained in why memory.grow invalidates pointers. And the result is described by a pointer and a length rather than a copy, so nothing is duplicated until you decide to keep it. Workloads that ignore the second point spend more time in slice() than in the algorithm.

Performance, size and the tradeoffs that bite

Four numbers describe a production module, and they trade against each other:

Number Typical lever What it costs you
Transfer size -Oz, wasm-opt, strip debug info, Brotli Slower code, no stack traces
Compile time Smaller binary, streaming compile Usually follows size
Resident memory Smaller initial memory, arena reuse More memory.grow calls, more churn
Steady-state rate -O3, SIMD, threads Bigger binary, stricter headers

The honest version of the size/speed tradeoff is that -Oz and -O3 differ by a factor of two or more on both axes for numeric code, and the right choice depends entirely on whether your users pay the download once per session or once per interaction. A dashboard that loads a module at startup and runs it for an hour should take the fast, fat build; a marketing page that uses a module for a single animation should take the small one. The optimization flags guide has the measurements, and SIMD is the one lever that often improves both at once, because a vectorised loop does more work per instruction fetched.

Choosing a build profile from the usage pattern If the module runs once per page view, optimise for size. If it runs continuously, optimise for speed. If the workload is embarrassingly parallel and the headers allow cross-origin isolation, add threads and accept the larger payload. how often does it run? once a visit all session batch, parallel small build -Oz, strip, Brotli lazy, on interaction target: bytes shipped fast build -O3, SIMD, cached module warm it during idle time target: steady-state rate threaded build shared memory, worker pool needs cross-origin isolation target: wall clock per batch Ship more than one profile when the split is real: a baseline build for everyone and a SIMD or threaded build behind a capability check. The wrong default is a single build tuned for the machine you develop on.

Shipping updates without breaking a cached page

A .wasm file is a deployment artifact with the same versioning problems as any other, plus one of its own: the glue JavaScript and the binary must agree on the ABI, and a CDN is perfectly capable of serving you last week’s binary with this week’s glue. The symptom is a LinkError about a missing import, or — far worse — a silent mismatch where a struct gained a field and every offset after it is now wrong.

Three rules keep that from happening:

  • Hash the filename. pipeline.a91c3f.wasm with a long immutable cache lifetime, referenced from glue that is itself hashed. Never ship pipeline.wasm with a short TTL and hope.
  • Version the ABI explicitly. Export a constant the loader checks before the first real call, so a mismatch fails loudly at startup instead of corrupting data an hour later.
  • Invalidate compiled-module caches on the same key. If you store a compiled module in IndexedDB, the cache key must include the binary’s hash, or a returning user gets last week’s code forever.
const ABI = 3;
const { instance } = await WebAssembly.instantiateStreaming(fetch(WASM_URL), imports);
if (instance.exports.abi_version() !== ABI) {
  throw new Error(`wasm ABI ${instance.exports.abi_version()} != loader ABI ${ABI}`);
}

The same discipline applies on the server, where the failure is quieter because there is no user to report it. An edge deployment that bundles the module with the handler gets this for free — the two ship together or not at all — which is a genuine argument for bundling over fetching on runtimes that allow both.

Rolling out: fallbacks are part of the feature

Something in the chain will be unavailable for some fraction of your users: SharedArrayBuffer behind missing headers, SIMD on an old engine, an extension that blocks application/wasm, or a corporate proxy that mangles the response. A production feature therefore has two implementations far more often than teams expect — the compiled path and a slower path that still produces a correct answer.

async function loadEngine() {
  try {
    if (!crossOriginIsolated) throw new Error('no shared memory');
    return await loadThreadedWasm();          // best case
  } catch {
    try { return await loadBaselineWasm(); }  // single-threaded, no SIMD
    catch { return loadJsFallback(); }        // correct, slower, always works
  }
}

The fallback does not have to be fast; it has to exist and be tested. Route a small percentage of traffic through it deliberately so it does not rot, and log which path each session took — the ratio is one of the more interesting numbers you will collect, and it is usually not what the team guessed. The polyfill and fallback pages cover detection order and the traps in each check.

Observability: five numbers, logged from day one

Every production Wasm problem is easier with these five numbers, and all of them cost a few lines:

  1. Transfer bytes and time — from the resource timing entry for the .wasm request.
  2. Compile time — around WebAssembly.compileStreaming, separately from instantiate.
  3. Instantiate time — around WebAssembly.instantiate, which is where import-object mistakes and large data segments show up.
  4. Time to first result — from the user’s action to a value they can see.
  5. Steady-state rate — items per second once warm, which is the only number that reflects the compiled code itself.
const t0 = performance.now();
const { module, instance } = await WebAssembly.instantiateStreaming(fetch(url), imports);
const t1 = performance.now();
const first = run(instance, sample);            // one warm-up item
const t2 = performance.now();
for (let i = 0; i < 200; i++) run(instance, sample);
const t3 = performance.now();
report({
  loadCompileMs: +(t1 - t0).toFixed(1),
  firstResultMs: +(t2 - t1).toFixed(1),
  steadyItemsPerSec: +(200 / ((t3 - t2) / 1000)).toFixed(1),
});

The reason to separate them is that they move independently and mislead when merged. A change that cuts the binary in half improves the first two numbers and may slightly worsen the last. Enabling SIMD improves the last and worsens the first. A single “how long did it take” metric hides both, and teams that only have that metric regularly conclude that WebAssembly “didn’t help” when what actually happened is that they measured a cold start two hundred times. The benchmark harness guide turns this into something you can run in CI and compare across commits.

Security: the sandbox is the product

Most of this area’s workloads involve code or data you did not write — a codec from a vendor, a model from a bucket, a plugin from a customer. WebAssembly’s value there is not speed; it is that a module cannot reach anything you did not hand it. There is no ambient filesystem, no network, no DOM, and no way to forge a pointer outside linear memory. Everything a module can do arrives through its import object, which makes the import object your security policy.

// the module gets exactly two capabilities, and neither is a general-purpose escape hatch
const instance = await WebAssembly.instantiate(module, {
  host: {
    log: (ptr, len) => console.log(readUtf8(ptr, len)),   // no raw console access
    now: () => Math.floor(performance.now()),             // coarse clock, no fingerprinting
  },
});

That said, the boundaries are specific and worth knowing precisely. A module can exhaust memory, spin forever, and read every byte of its own heap including data you put there for a different tenant. It cannot escape the process, but “cannot escape” is not “cannot harm”. The browser sandbox boundaries pages cover what the platform guarantees, and sandboxing untrusted code with Wasm covers the guarantees you have to build yourself: fuel limits, memory caps, per-tenant instances and a deliberate, minimal import surface.

Multi-tenant hosts need one more layer: a way to stop a module that never returns. Browsers give you a worker you can terminate; server runtimes give you fuel metering or epoch interruption, which let the host pre-empt a module after a bounded amount of work. Neither is optional once customers can upload code. The same applies to memory — an instance with an unbounded maximum will happily take the whole address space when handed a crafted input, so set a maximum at instantiation and treat a failed memory.grow as a normal, recoverable error rather than a crash.

What this area covers

Frequently Asked Questions

Is WebAssembly worth it for a feature that takes 30 ms in JavaScript? Almost never. The compile and instantiate cost alone is typically 5–40 ms for a small module, and you add a boundary crossing plus a memory copy on every call. The wins start when a JavaScript implementation takes hundreds of milliseconds, when the work is numeric and cache-friendly, or when you already have a C/C++/Rust implementation you would otherwise have to port and keep in sync.

Can I ship one module for the browser and the server? The core is portable; the edges are not. Compile against WASI for the server and against the browser’s import object for the page, and keep host interactions behind a thin interface you implement twice. Code that only computes over linear memory moves without change — anything touching files, clocks, randomness or threads needs a per-host implementation.

How big is too big for a browser module? Treat 1 MB compressed as the point where you need a plan: lazy loading, a cached compiled module, or a split between a small always-loaded core and an on-demand extension. Modules of 5–10 MB ship in production (Pyodide and some codec bundles are far larger) but only when the page can defer them past first paint and cache them aggressively afterwards.

Do threads actually help in the browser? Yes, for parallel numeric work, and only if you can serve the page cross-origin isolated with COOP and COEP headers. Without those, SharedArrayBuffer is unavailable and a threaded build will fail at instantiation. Check the headers before you plan the architecture, not after.

What is the single most common production mistake? Caching memory.buffer in a variable. Any allocation inside the module can grow memory, which detaches every view built on the old buffer, and the symptom — zero-length reads, silently wrong results — shows up far from the cause and usually only under load.

How do I stop a third-party module from hanging the page? Run it in a worker and terminate the worker on a deadline; there is no way to interrupt a running instance from the same thread. On a server runtime use fuel metering or epoch interruption instead, which pre-empt the module without killing the host process. Either way, decide the deadline before you accept the module, not after a customer trips it.

Which language should I compile from? Rust if you want the smallest self-contained binaries and the best browser tooling; C or C++ if the code already exists; Go or .NET if the team’s expertise dominates and payload size is negotiable; Zig when you want C-level output with a modern build. The payload comparison puts numbers on the difference, which is roughly two orders of magnitude between the extremes.

← Back to the site home