When Porting to Wasm Makes Code Slower

This page answers one task: a WebAssembly port was supposed to speed something up, the benchmark says it is slower (or no faster), and you need to know why — and whether the port can be rescued or should be abandoned.

Prerequisites

  • [ ] The original JavaScript and the Wasm port, both runnable on the same inputs.
  • [ ] A benchmark harness with warm-up, as in avoiding JIT warm-up errors in Wasm benchmarks.
  • [ ] A profile of the Wasm version, to see where its time goes.

Why “compiled” does not mean “faster”

Modern JavaScript engines are optimising compilers. Hot, type-stable JavaScript — a loop over a typed array with consistent integer arithmetic — is compiled to machine code that is often within a small factor of what a WebAssembly compiler produces for the same algorithm. WebAssembly’s advantages are real but specific: predictable performance without warm-up or deoptimisation, efficient 64-bit integers, SIMD, manual memory layout, and freedom from garbage-collection pauses. When the workload does not use those advantages, a port adds costs that JavaScript did not have: crossing the boundary, copying data into linear memory and back, converting strings, and creating JavaScript objects for results. If those costs exceed the computational gain, the port is slower.

Ports lose in a recognisable set of situations. Recognising them early — ideally before porting — saves weeks.

Code shapes where a Wasm port tends to lose Chatty boundaries pay per-call overhead. DOM-heavy code must call back into JavaScript for every operation. String-heavy code pays encoding costs. Object-graph code must marshal structures. Work already done by native built-ins gains nothing. Short scenarios pay compile and startup costs. code shape why Wasm loses rescue many tiny calls per-call boundary cost batch the calls DOM and Web API heavy every operation is an import keep in JS string processing UTF-16 to UTF-8 copies operate on bytes object graphs marshalling every field flat buffers native built-ins already compiled code keep in JS one-off short tasks compile + instantiate cost cache or skip

Step 1 — measure the boundary separately

Profile the Wasm version and separate three costs: time inside Wasm functions, time in the glue (encoding, copying, object creation), and time in JavaScript callbacks the module calls. If glue and callbacks dominate, the algorithm itself may be fast, and the problem is the interface. A quick test is to call the Wasm entry point with pre-encoded input and discard the output: if that is fast and the full call is slow, the boundary is the bottleneck. The per-call overheads are measured in measuring JS-to-Wasm call overhead.

Step 2 — fix chatty interfaces by batching

The most common rescue is changing the call pattern. A function called once per item becomes a function called once per array:

// slow: 1,000,000 boundary crossings
for (const p of points) out.push(wasm.transform(p.x, p.y));

// fast: one crossing, data in typed arrays
const xs = Float64Array.from(points, (p) => p.x), ys = Float64Array.from(points, (p) => p.y);
const result = wasm.transform_all(xs, ys);              // returns Float64Array of pairs

Even better, keep the data in linear memory across calls so it is never copied, as described in double buffering data between JavaScript and Wasm. Batching often turns a 2× slowdown into a 3× speed-up without changing the algorithm.

Step 3 — leave DOM and Web API work in JavaScript

Code that creates elements, reads layout, sets styles or calls Web APIs cannot run faster in WebAssembly: every such operation is an import call into JavaScript, which then performs the same work it would have done anyway, plus marshalling. Frameworks that write UIs in Rust or C# succeed by doing their computation (diffing, state management) in Wasm and batching DOM mutations, not by making individual DOM calls faster. For a port, split the function: the computation moves to Wasm and returns a compact description of what changed, and JavaScript applies it to the DOM. The general argument is covered in is WebAssembly faster than JavaScript for DOM manipulation?.

A slower port and its rescued version The slow port calls into Wasm once per item, converts strings each time and builds result objects through the boundary. The rescued version passes typed arrays or bytes once, keeps the algorithm in Wasm, and returns a flat result, leaving DOM work to JavaScript. slower port one call per item strings converted every call objects built via the glue 1.8× slower than JS rescued port one call per batch bytes and typed arrays in flat result, JS renders 3.1× faster than JS

Step 4 — avoid string and object marshalling

Strings must be encoded from UTF-16 to UTF-8 and copied into linear memory on the way in, and decoded on the way out. For workloads that do a little work per string — trimming, splitting, simple matching — that conversion can cost more than the work. Either operate on bytes end to end (read files as Uint8Array, not strings), or keep the string work in JavaScript, whose string operations are native code. Similarly, returning results as nested JavaScript objects through generated glue creates them one boundary call at a time; return flat buffers or a single JSON string instead, as in passing nested objects efficiently.

Step 5 — account for startup and know when to stop

For short tasks, compile and instantiation time can dominate: a 1 MB module may take 30–100 ms to compile on a phone, more than a one-off computation it replaces. Cache compiled modules, load them before they are needed, or keep short tasks in JavaScript. And if after batching, byte-level input and flat output the port is still not faster, accept the result. The JavaScript was already well optimised, and maintaining a second language for no gain is a cost. Document the measurement so the idea is not revisited without new evidence.

Built-ins you cannot beat

Some JavaScript operations are thin wrappers over highly optimised native code, and a Wasm reimplementation competes with C++ or hand-written assembly inside the engine. TypedArray.prototype.sort without a comparator, Array.prototype.indexOf on packed arrays, String.prototype.indexOf, regular expressions (compiled by the engine), JSON.parse and JSON.stringify, TextEncoder/TextDecoder, crypto.subtle digests and Math functions are all native. Code that spends most of its time inside these does not benefit from porting; replacing JSON.parse with a Rust JSON parser compiled to Wasm, for example, is usually slower once the result has to become JavaScript objects. Look at the bottom-up profile: if the self time sits in built-ins rather than in your functions, the opportunity is to call them less, not to port them.

Garbage collection is not always the enemy

Ports are often motivated by GC pauses, and WebAssembly’s linear memory does avoid them. But modern engines’ generational collectors handle short-lived objects cheaply, and a port that replaces many small JavaScript objects with a Rust Vec of structs then has to convert those structs back into JavaScript objects for callers — reintroducing the allocations. The gain is real only when the data can stay in linear memory for most of its life and only small results cross. Measure GC time before and after: if it barely changes, the port moved the allocation rather than removing it.

Debug builds and unoptimised modules

A surprising number of “Wasm is slower” reports come from measuring the wrong build. A Rust module built without --release is often 10–30× slower than the release build, and an Emscripten build at -O0 is similarly slow; wasm-pack’s --dev profile and some framework dev servers produce exactly those. JavaScript, meanwhile, is always optimised by the engine. Before drawing conclusions, confirm the benchmark uses a release build with wasm-opt applied, and that debug assertions, overflow checks and logging are off. Check the module’s size too: a release build is typically several times smaller than a debug build, so an unexpectedly large .wasm in the benchmark is a strong hint.

Expected output

A short report: the original port was 1.8× slower due to one boundary call per item and string conversions; after batching into typed arrays and returning a flat result it is 3.1× faster; a second candidate, DOM-heavy formatting, was left in JavaScript after a prototype showed no gain.

Gotchas

  • Benchmarking only the Wasm function. Measure the full call including glue and conversion.
  • Porting code dominated by built-ins. They are already native. Call them less instead.
  • Per-item calls. Batch work into arrays.
  • Converting strings for trivial work. Operate on bytes or keep it in JavaScript.
  • Ignoring startup. Compile cost can outweigh a short computation.

Performance note

A geometry transform over 1 million points took 38 ms in JavaScript, 69 ms in a per-point Wasm port, and 12 ms in a batched Wasm port with typed-array input and output. The algorithm was identical in all three; only the interface changed.

Transforming one million points three ways Milliseconds to transform one million 2D points with the original JavaScript, a Wasm port called once per point, and a Wasm port called once with typed arrays. ms per run original JavaScript 38 ms Wasm, one call per point 69 ms Wasm, one batched call 12 ms

Frequently Asked Questions

Is WebAssembly ever slower for pure computation? Rarely by much; JavaScript’s JIT can match it for simple loops. Wasm’s advantage grows with integer-heavy, SIMD-friendly or GC-sensitive code.

Does WasmGC change the picture for object-heavy code? It lets compiled languages use the engine’s GC, reducing marshalling for some languages, but crossing to JavaScript objects still costs.

Should I port to get predictable performance rather than peak speed? That is a valid reason: Wasm avoids deoptimisation cliffs, which matters for latency-sensitive code.

How do I know if my JavaScript is already optimal? Check for deoptimisations in the profiler and compare with simple, type-stable rewrites before porting.

Can I mix: Wasm for some steps, JS for others? Yes — that is usually the best design, with each step where it is fastest and data crossing as few times as possible.

Why is my Wasm port 20× slower in development? Development builds are unoptimised. Benchmark release builds with wasm-opt applied.

← Back to Porting JavaScript Hot Paths to Wasm