Estimating the Speed-Up Before Porting

This page answers one task: before committing weeks to a WebAssembly port, estimate how much faster the user-visible operation will actually become — with numbers good enough to decide go or no-go, and cheap enough to produce in a day or two.

Prerequisites

  • [ ] A profile of the slow scenario, as in profiling JavaScript to find Wasm candidates.
  • [ ] The input sizes and call counts for the candidate code.
  • [ ] Enough Rust, C or AssemblyScript to write a small prototype of the inner loop.

The arithmetic that decides most ports

The upper bound on any speed-up is set by how much of the scenario the candidate code accounts for — Amdahl’s law. If the candidate is a fraction p of the total time and becomes s times faster, the scenario’s new duration is (1 − p) + p / s of the original, plus whatever new costs the port adds: copying input into linear memory, converting results back, loading and compiling the module. A port of code that is 60% of the time and becomes 4× faster saves 45%; the same port of code that is 10% of the time saves 7.5%, which users will not notice.

So the estimate needs three numbers: p from the profile, s from a prototype, and the added costs from a quick measurement of data movement and module startup. Each can be measured in hours, and together they predict the outcome far better than intuition or published benchmarks of other workloads.

From profile to a go/no-go estimate The profile gives the candidate's share of the scenario. A micro-prototype of the inner loop gives the speed-up factor. A measurement of copying, conversion and module startup gives the added costs. Combining them predicts the new scenario time and the decision. share p from the profile speed-up s from a prototype added costs copy, convert, startup (1−p) + p/s + costs predicted time go / no-go against a target

Step 1 — take p from a realistic profile

Measure the candidate region’s share of the scenario on the devices that matter — usually a mid-range phone, where the gap between JavaScript and Wasm and the cost of garbage collection are both larger than on a laptop. Include the garbage-collection time attributable to the region if its allocations would disappear in a port. Write down the absolute numbers too: 60% of 2 seconds is a different decision from 60% of 50 milliseconds.

Step 2 — prototype only the inner loop

Port only the innermost hot loop — the part where most of the self time is — with fixed, pre-loaded input, no error handling and no API polish. It should take hours, not days:

#[no_mangle]
pub extern "C" fn checksum(ptr: *const u8, len: usize) -> u32 {
    let data = unsafe { core::slice::from_raw_parts(ptr, len) };
    let (mut a, mut b) = (1u32, 0u32);
    for &x in data { a = (a + x as u32) % 65521; b = (b + a) % 65521; }   // the inner loop, as in the JS
    (b << 16) | a
}

Benchmark it against the JavaScript inner loop on the same input, after warm-up, both in a worker or both on the main thread. The ratio is s. Try the obvious Wasm-specific improvement too — SIMD, 64-bit arithmetic, avoiding a modulo per iteration — because that is often where the real gain lies; report both the straight port and the improved one.

Step 3 — measure the added costs

Measure, separately, what the full port would add: encoding and copying the input into linear memory, converting results back to the shapes callers expect, and compiling and instantiating the module the first time. A 20 MB input copied with TypedArray.set costs a few milliseconds; a million result objects built from Wasm data can cost tens of milliseconds; compiling a 300 KB module on a phone costs 10–30 ms once. For long-running pages the startup cost is amortised; for one-shot tasks it is not.

Predicted versus measured import time after a port Seconds for a large file import on a mid-range phone: the original JavaScript, the estimate computed from the profile share, prototype speed-up and measured added costs, and the time measured after the full port shipped. seconds per import original JavaScript 2.1 s estimate before porting 0.9 s measured after the port 0.9 s

Step 4 — compute and compare with a target

Put the numbers together. Example: an import takes 2.1 s; the parsing region is p = 0.62; the prototype runs 3.4× faster (s = 3.4); copying and conversion add 0.06 s; compilation is cached after first use. Predicted time: 2.1 × (0.38 + 0.62 / 3.4) + 0.06 = 1.24 s. With the SIMD variant at s = 6, it is 2.1 × (0.38 + 0.103) + 0.06 = 1.07 s. Compare with a target the product cares about — “imports under 1.5 s on mid-range phones” — and the decision follows: go, with SIMD.

Step 5 — write the decision down

Record the estimate, the measurements behind it and the decision in the issue or design document, using a short template: scenario, devices, p, s (plain and improved), added costs, predicted time, target, decision. When the port ships, compare the measured result with the prediction. Over a few ports, this builds a team-specific sense of how accurate estimates are, and the written record stops the same rejected idea from being proposed again every quarter.

Common estimation mistakes

Three mistakes account for most bad predictions. The first is using a laptop for p and a phone for the target — the profile shape changes between devices, sometimes dramatically. The second is prototyping a different algorithm: if the prototype uses a smarter algorithm than the JavaScript, part of the measured speed-up would also be available by improving the JavaScript, and the honest comparison is against that improved JavaScript. The third is ignoring how data reaches the code: a prototype that reads pre-loaded linear memory does not pay for decoding a string or building objects, and in real use those costs can be as large as the loop itself. Each mistake makes the estimate optimistic, so a sensible habit is to apply a safety margin — treat the predicted improvement as an upper bound and require it to clear the target comfortably, not narrowly.

Non-speed reasons to port

Sometimes the estimate shows a modest speed-up and porting is still the right call. Predictability can matter more than average speed: WebAssembly code does not deoptimise, so tail latencies for interactive features can drop even when the median barely moves. Sharing one implementation between the browser and a Rust or C server removes a class of disagreement bugs. Access to mature native libraries — codecs, cryptography, geometry — can be worth more than speed. Record such reasons explicitly in the decision, so the port is judged on the right criteria afterwards, and so a later “it is only 20% faster” does not lead to an unfounded rollback.

Estimating the cost side

A decision weighs benefit against cost, and porting has costs beyond the initial work. Estimate them as concretely as the speed-up: the person-days to port and test; the build changes (a Rust or C toolchain in CI, wasm-bindgen or Emscripten, binary size checks); the bytes added to the download; the second language every future maintainer must read; and the period during which both implementations must be kept in sync. A port that saves 300 ms on a feature used by every user is worth a lot of that; a port that saves 50 ms on an admin screen is not. Writing both sides down — predicted benefit, predicted cost — makes the decision easy to explain and easy to revisit.

Expected output

A one-page estimate concluding “go: predicted 1.07 s against a 1.5 s target on mid-range phones, using the SIMD variant”, backed by the profile and the prototype benchmark; after the port ships, the measured time of 0.9 s is recorded next to the prediction.

Gotchas

  • Estimating from published benchmarks. Other workloads tell you little. Prototype your own loop.
  • Comparing against unoptimised JavaScript. Improve the JavaScript first or compare with an improved version.
  • Forgetting data movement. Copying and conversion can erase the gain. Measure them.
  • Laptop-only numbers. Measure where users are.
  • No written prediction. You cannot learn from estimates you did not record.

Performance note

The micro-prototype took about four hours to write and benchmark; the full port took eight days. The estimate was within 15% of the measured outcome, which is typical when the profile, prototype and added costs are all measured on the target device.

An estimate worksheet for one candidate The worksheet records the scenario time, the candidate share from the profile, the prototype speed-up for a straight and a SIMD port, the measured added costs, the predicted time, the target and the decision. input value source scenario time 2.1 s profile on mid-range phone candidate share p 0.62 bottom-up self time speed-up s (plain / SIMD) 3.4 / 6.0 prototype benchmark added costs 0.06 s copy + conversion timing prediction vs target 1.07 s vs 1.5 s go

Frequently Asked Questions

How accurate are these estimates? Usually within 10–25% when all inputs are measured on the target device; worse when any are guessed.

What if the prototype is slower than JavaScript? Check build settings first (release mode, wasm-opt). If it is still slower, the JavaScript is already near optimal; stop.

Should I include compile time? Yes for one-shot tasks; for long-running pages, amortise it or cache compiled modules.

Can I estimate SIMD gains without writing SIMD? Roughly — 2–4× for suitable loops — but a short SIMD prototype is more reliable.

Does the estimate change for server-side Node? The same method applies; startup is amortised and device variation disappears.

Who should review the estimate? Someone who will maintain the code afterwards; they weigh the long-term cost of a second language most realistically.

← Back to Porting JavaScript Hot Paths to Wasm