Porting JavaScript Hot Paths to Wasm

Most teams meet WebAssembly not on a new project but on an existing JavaScript application with a slow spot: an import that takes seconds, a filter that stutters, a parser that blocks the page. The question is not “should we rewrite in Wasm” but “which few hundred lines, if moved, would make users wait less — and how do we move them without breaking anything”. Done well, a port of a small, well-chosen hot path delivers a large, measurable improvement while the rest of the application stays in JavaScript. Done badly, it adds a second language, a build pipeline and a binary download, and makes the slow spot no faster — or slower.

This topic covers the whole path: finding candidates with a profiler, predicting the gain before writing the port, designing the interface so data crosses the boundary rarely, keeping results identical to the JavaScript, rolling the port out behind a flag, and recognising the cases where the JavaScript should stay. The guides below take each step in depth, with worked examples.

Prerequisites

  • [ ] A JavaScript application with a reproducible slow scenario and production-sized test data.
  • [ ] Chrome DevTools or Node’s --cpu-prof for profiling.
  • [ ] A compiled language toolchain — usually Rust with wasm-bindgen, or C/C++ with Emscripten, or AssemblyScript for teams close to TypeScript.
  • [ ] Real-user monitoring or at least a benchmark harness that runs on target devices.

The porting workflow

Successful ports follow a short, disciplined sequence. Each step produces evidence for the next, and each can end the project early if the evidence says the port will not pay off. That is a feature: the cheapest port is the one you decided not to do because a profile and a four-hour prototype showed it would save 6%.

The porting workflow Profile the real scenario to find candidates, estimate the gain with a prototype, port the region behind the original API with a one-call interface, prove parity with differential tests, roll out behind a flag with shadow comparison, then remove the old path. profile where CPU actually goes estimate prototype the inner loop port behind the API one call in, one result out prove parity differential tests roll out flag, shadow, ramp

Choosing what to port

A profile of the real scenario, on the devices users have, is the only reliable starting point. Sort by self time to see where the CPU goes, find the entry point that encloses the hot region, and score it on four properties: its share of total time, whether it works on numbers and bytes rather than objects and DOM, how much it allocates, and how often it is called. Strong candidates are parsers, decoders, compressors, hashing, image and audio kernels, geometry and physics, search and matching over large arrays — compute-bound code with data that can live in linear memory and an interface that can be called a few times with large inputs. Weak candidates are code dominated by DOM manipulation, Web API calls, built-ins that are already native (regular expressions, JSON.parse, sorting without comparators), or millions of tiny calls.

Before porting, optimise the JavaScript. Profiles often reveal deoptimised functions, accidental quadratic loops and repeated work that a few hours of conventional optimisation fixes. Compare any port against the best JavaScript, not the first version.

Predicting the gain

Amdahl’s law sets the ceiling: if the candidate is a fraction p of the scenario and becomes s times faster, the scenario takes (1 − p) + p/s of its original time, plus whatever the port adds in copying, conversion and module startup. Measure p from the profile, measure s with a small prototype of just the inner loop, and measure the added costs directly. Together they predict the outcome within a few tens of percent, which is enough to decide. A written estimate also creates accountability: comparing the prediction with the measured result after shipping improves the next estimate.

How the candidate's share limits the possible gain Best possible reduction in scenario time for a candidate that becomes four times faster, depending on its share of the scenario: ten, thirty, sixty and ninety percent. scenario time saved (%) candidate is 10% of time 7.5 % candidate is 30% of time 22.5 % candidate is 60% of time 45 % candidate is 90% of time 67.5 %

Designing the interface

The interface between JavaScript and the ported code decides more of the outcome than the language does. Every call crosses the boundary; every string is encoded and copied; every result object is built through glue code. A port that mirrors the JavaScript structure — an object with a method called per token or per pixel — pays those costs millions of times and often loses. A port designed for the boundary hands over the whole input once, does all the work inside WebAssembly, and returns a compact result: a typed array, a flat binary buffer, a single JSON string, or offsets into the input. A thin JavaScript wrapper keeps the original public API, so callers do not change and both implementations can coexist during rollout.

For repeated operations on large data — filters applied as a slider moves, frames processed in a loop — keep the data in linear memory across calls and let JavaScript write into and read from views over it, removing copies entirely.

Keeping behaviour identical

A port is a refactoring: callers must not see a difference unless you decided they should. JavaScript and compiled languages differ in predictable places — floating-point operation order, integers above 2⁵³ and overflow, UTF-16 versus UTF-8 lengths and positions, sort stability, number formatting, and implicit inputs like the time zone or locale. A differential harness that runs both implementations on a production corpus and on property-based generated inputs finds these differences quickly. Each one gets a decision: match the JavaScript, or accept a deliberate, documented behaviour change with a regression test.

Port design decisions and their usual answers The interface should take the whole input and return a compact result. The public API should stay unchanged behind a wrapper. Parity should be proven with a differential harness. Rollout should use a feature flag with shadow comparison. The old code should be removed once the port is stable. decision usual answer why interface granularity whole input per call boundary costs public API unchanged callers and rollout correctness differential harness hidden divergences rollout flag + shadow mode production inputs old implementation remove after stable maintenance cost

Rolling out and cleaning up

Ship the port behind a feature flag with both implementations available. Start in shadow mode, where the JavaScript serves results and the port runs alongside on a sample of calls for comparison; then serve the port to a growing percentage of users, bucketed by a stable id, watching latency, error and fallback rates per cohort. Make the fallback to JavaScript automatic for users whose browsers cannot load the module, and record why it happened. Test the kill switch before you need it. When the port has served everyone for a few weeks, remove the old implementation or demote it to a documented fallback with its own tests — two actively maintained implementations of the same logic are a cost that compounds.

Where the speed-up really comes from

Worked examples consistently show that a straight translation of well-typed JavaScript into a compiled language gives a modest gain — often 1.2–2× — because modern JavaScript engines already optimise such code well. The large gains come from three other sources: better algorithms (which usually also help the JavaScript), removing data movement and allocation, and capabilities JavaScript lacks — SIMD, efficient 64-bit integers, explicit memory layout and freedom from garbage-collection pauses. Measuring each change separately prevents crediting WebAssembly for gains that were really algorithmic, keeps the JavaScript fallback fast by backporting algorithmic improvements, and focuses the port on the parts only WebAssembly can accelerate.

Choosing the source language

Rust is the most common choice for ports: wasm-bindgen produces convenient glue, the tooling is mature, binaries are small, and the type system catches many memory bugs that would otherwise surface as traps. C and C++ make sense when a mature native library already implements the algorithm — porting a JavaScript image decoder is pointless when libjpeg-turbo exists — and Emscripten handles the build. AssemblyScript is attractive for teams that want to stay close to TypeScript; its performance is good for numeric loops, though its ecosystem is smaller and its semantics differ from TypeScript in places. Zig and Go (via TinyGo) are viable too. The language matters less than the interface design and the measurement discipline described above; pick the one your team can maintain, and the one with the libraries the problem needs.

Moving the data, not just the code

The input to a hot path rarely starts life in the form the port wants. A file arrives as a Blob, network data as a ReadableStream, canvas pixels as ImageData, user text as a JavaScript string. Each conversion costs time and memory, and the port’s benefit depends on keeping them few. Prefer bytes over strings: read files and responses as Uint8Array and hand them straight to the module, rather than decoding to a string in JavaScript and re-encoding to UTF-8 for the port. Stream large inputs in chunks into a fixed buffer in linear memory, so peak memory stays bounded and processing overlaps loading. Produce outputs in the form the next consumer needs — an ImageData view over linear memory for a canvas, a typed array for a chart, a compact buffer for a worker transfer — rather than JavaScript objects that the next step immediately converts again. Mapping the data’s journey end to end, before writing the port, often reveals that the biggest saving is in conversions the original JavaScript performed, not in the computation itself.

Building the test infrastructure first

The port’s test infrastructure is worth building before the port, because it pays off at every later step. A corpus of real inputs with recorded outputs from the current JavaScript defines correctness. A differential harness that runs any two implementations over that corpus and over generated inputs turns parity into a number. A benchmark harness that runs on target devices, with warm-up and repetition, turns speed into a number. With these in place, each optimisation step is a quick measure-and-compare loop, and the rollout’s shadow mode reuses the same comparison code in production. Without them, every change is a debate about whether it helped and whether it broke something.

Ports on the server

Server-side JavaScript has the same hot paths — parsing uploads, compressing responses, validating payloads, rendering templates — and the same workflow applies, with friendlier conditions. There is no download cost, module startup is amortised over the life of the process, devices do not vary, and inputs can be compared in-process without privacy concerns. The main benefit is often capacity rather than latency: a port that halves CPU per request halves the number of instances needed. Measure CPU time per request and instance count alongside latency when deciding, and roll out per request or per tenant rather than per user.

Team and maintenance considerations

A port introduces a second language into a JavaScript codebase, with its own toolchain, build steps and review skills. Keep the ported code small, focused and well tested, and document the boundary contract — inputs, outputs, error behaviour — so JavaScript developers can use it without reading Rust or C. Pin toolchain versions, run the compiled code’s tests in the same CI as the JavaScript, and assign clear ownership. Ports that stay small and self-contained remain easy to maintain for years; ports that grow into a parallel application with its own conventions become the next migration problem.

Using existing native libraries instead of porting

Sometimes the best port is not a port at all. Many hot paths implement something a mature native library already does better — image codecs, compression, cryptography, geometry, text shaping, regular-expression engines, SQL. Compiling such a library to WebAssembly with Emscripten or wasi-sdk, or adopting an existing Wasm build of it, replaces hand-written JavaScript with code that has years of optimisation and testing behind it. The work then shifts from writing an algorithm to designing the interface and managing the binary’s size, which is usually less risky. Check package registries for existing Wasm builds before starting; for common formats and algorithms, one almost certainly exists, and the remaining question is whether its size and API fit.

Measuring success after launch

The port’s success is measured by users, not benchmarks. Before the rollout, define the metric that motivated it — import duration at the 95th percentile on mid-range phones, frames dropped while dragging a slider, server CPU per request — and record its baseline. After full rollout, compare the same metric over the same kind of period, segmented by device class and browser. Keep the comparison in the project’s documentation alongside the original estimate. If the gain is smaller than predicted, the stage timings usually show why, and that analysis is valuable for the next port even if this one is kept as it is.

Gotchas and failure modes

  • Porting without a profile. Intuition about hot spots is often wrong. Measure first.
  • Chatty interfaces. Per-item calls can make a port slower than JavaScript. Batch.
  • Benchmarking debug builds. Unoptimised Wasm is many times slower. Measure release builds.
  • Comparing with unoptimised JavaScript. Optimise the JavaScript first, or the port gets credit it did not earn.
  • Unexamined behaviour changes. Differential tests must cover floats, integers, text and ordering.
  • Ignoring existing libraries. A mature native library compiled to Wasm often beats a fresh port.
  • Never removing the old path. Set a removal date at the start of the rollout.

Verification

A port is verified when three things hold. The differential harness reports no unexplained differences on the corpus and on generated inputs. The rollout dashboard shows the target improvement on the devices that matter, at the median and the 95th percentile, with a negligible fallback rate. And the measured gain is recorded next to the prediction made before the port, closing the loop on the estimate.

Guides in this topic

Frequently Asked Questions

How much faster will my code be in WebAssembly? It depends on the code shape. Straight translations of typed JavaScript gain 1.2–2×; ports that add SIMD, remove allocation and batch boundary crossings often gain 3–10× on the ported region. Measure your own loop.

Should we rewrite the whole application? Almost never. Port the hot region behind its existing API and keep the rest in JavaScript.

Which language should the port use? The one your team can maintain and that has the libraries the problem needs — usually Rust, sometimes C/C++ or AssemblyScript.

Will the port work in all browsers? WebAssembly itself is universal; features like SIMD and threads need detection and fallback builds.

What about server-side JavaScript? The same workflow applies to Node, Deno and Bun, often with easier measurement and no download cost.

How long does a typical port take? For a focused hot path, days to a few weeks, including tests and rollout — most of it in parity testing and integration, not the core algorithm.

Can a port make the bundle much larger? A focused Rust port is often 50–300 KB compressed; whole libraries compiled from C can be megabytes. Load the module lazily where it is used.

What if only some users benefit? That is common — slow devices benefit most. Segment rollout metrics by device class before concluding.

How many functions should a first port include? One hot path with a clear interface; widen the scope only after it ships and the gain is confirmed in production.

What if the profile has no single dominant hot path? Then WebAssembly is unlikely to help much; spread-out cost usually comes from architecture, allocation or rendering, which a port does not fix.

Can a port be reverted easily? Yes, if the JavaScript implementation stays behind a flag until the Wasm version has run in production long enough to trust it.

Who should own the ported module? The team that owns the JavaScript feature, so behaviour, tests and releases stay with the people who understand the workload.

← Back to Production Wasm: Workloads & Deployment