Profiling JavaScript to Find Wasm Candidates
This page answers one task: an application is slow somewhere, someone suggests “rewrite it in WebAssembly”, and you need evidence about which code — if any — would actually get faster, before anyone spends weeks porting it.
Prerequisites
- [ ] A reproducible slow scenario: a user action, a page load, a batch job.
- [ ] Chrome DevTools (Performance panel) or Node with
--cpu-prof. - [ ] A production-like build with source maps, so profiles show real function names.
What makes code a good candidate
WebAssembly speeds up code by giving the engine predictable types and a compact, statically typed instruction stream that compiles to efficient machine code without speculation. JavaScript engines already optimise hot, type-stable code very well, so the gap is largest for specific shapes of work: tight numeric loops over typed arrays, bit manipulation, byte parsing, hashing, compression, image and audio processing, and algorithms with many integer operations — especially where SIMD or 64-bit integers help. The gap is small or negative for code that mostly manipulates JavaScript objects, calls Web APIs, touches the DOM, or allocates many short-lived objects, because WebAssembly has to cross back into JavaScript for all of those.
So a good candidate has three properties: it accounts for a large share of CPU time, it does its work on data that can live in linear memory, and it can be called with few boundary crossings — one call that processes a megabyte, not a million calls that each process a byte. Profiling measures the first property directly and gives strong hints about the other two.
Step 1 — record a profile of the real scenario
In Chrome, open DevTools → Performance, enable “Screenshots” off and “JavaScript samples” on, click record, perform the slow action, and stop. For Node:
node --cpu-prof --cpu-prof-dir=./profiles scripts/import-large-file.js
# open the .cpuprofile in Chrome DevTools (Performance → Load profile)
Profile a realistic workload — production-sized inputs on a mid-range device, not a toy example on a fast laptop. Record several runs; the first run includes JIT warm-up that a long-running page amortises, so compare steady-state runs when deciding.
Step 2 — rank by self time, then check total time
Switch to the Bottom-Up view and sort by self time: the time spent in each function’s own code, excluding callees. Functions at the top are where the CPU actually goes. Then look at the Call Tree for total time, to find the entry point that encloses a hot region — that entry point is the natural unit to port, because porting it moves the whole region behind one boundary call.
A typical result looks like this: parseRecord 34% self, decodeVarint 18% self, Object.assign 9%, garbage collection 12%, everything else spread
thin. parseRecord and decodeVarint together are a strong candidate region with an entry point parseFile; the garbage-collection share is a hint that
the current code allocates heavily, which a Wasm port with reusable buffers would also reduce.
Step 3 — measure call frequency and data movement
Count how often the candidate’s entry point is called and how much data crosses per call. Add temporary instrumentation:
let calls = 0, bytes = 0;
const original = parseFile;
parseFile = (buf) => { calls++; bytes += buf.byteLength; return original(buf); };
// after the scenario
console.log({ calls, bytesPerCall: bytes / calls });
A handful of calls moving megabytes each is ideal. Thousands of calls per second moving a few bytes each will spend most of their time crossing the boundary; such code needs restructuring into batches before a port can help. The costs are quantified in measuring JS-to-Wasm call overhead.
Step 4 — check what the hot code touches
Read the candidate’s code with one question: what does it touch besides numbers and buffers? Every document access, Map lookup, regular expression,
JSON.parse, Date, fetch or callback becomes an import call from WebAssembly or has to be reimplemented. A function that is 90% arithmetic and 10%
Map lookups can be ported if the map becomes a hash table in Rust; a function that is mostly regular expressions on strings will gain little, since
JavaScript’s regex engine is already compiled native code and strings must be copied into linear memory first. Note also any parts that already run in
native code — TextDecoder, crypto.subtle, Array.prototype.sort with a comparator does not count, but built-ins without callbacks often do.
Step 5 — estimate before committing
Write down the expected gain: the candidate’s share of the scenario times a plausible speed-up, minus the new boundary and copy costs. If parseFile
takes 60% of a 2-second import and a Wasm version is three times faster, the import drops to roughly 1.2 seconds — a meaningful change. If the hot
function is 8% of the time, even an infinitely fast port saves only 8%. That arithmetic, and a small prototype to check the speed-up assumption, is
described in
estimating the speed-up before porting.
Optimise the JavaScript first
Profiles often reveal problems that are cheaper to fix in JavaScript than to port: a function that deoptimises because it receives objects of different
shapes, an accidental quadratic loop, repeated string concatenation, JSON.parse of the same data in a loop, a missing cache. Chrome’s profiler marks
deoptimised functions, and a few hours of conventional optimisation sometimes produce a larger gain than a port would. Do that first, re-profile, and only
then decide; a fair comparison measures Wasm against the best JavaScript, not the first version. If the optimised JavaScript is fast enough, you have
saved the cost of maintaining two languages. If it is still the bottleneck, the optimised version becomes the reference implementation for the port’s
correctness tests, as in
keeping JavaScript and Wasm results identical.
Profiling on the devices that matter
Developer laptops hide problems. The same hot function can take 40 ms on a laptop and 300 ms on a mid-range phone, and garbage-collection pressure is far more visible on devices with less memory. Profile on a representative phone with remote debugging, and use Chrome’s CPU throttling only as a rough approximation. Field data helps decide which scenarios matter at all: real-user monitoring that records long tasks and slow interactions per feature tells you where users actually wait, which may not be where internal benchmarks point. A port that speeds up a rarely used export dialog matters less than one that speeds up the search box everyone types into, even if the dialog’s code is a better fit for WebAssembly.
Expected output
A short written assessment: “parseFile region — 58% of import time on a mid-range phone, pure byte parsing over an ArrayBuffer, 3 calls per import
with 4–20 MB each, allocation-heavy in JS (12% GC) — strong candidate; formatRows — 9%, DOM-heavy — not a candidate.”
Gotchas
- Profiling toy inputs. Small inputs hide the real hot spots. Use production-sized data.
- Sorting by total time only. Entry points dominate total time; self time shows where work happens.
- Ignoring call frequency. Many tiny calls make ports slower. Measure calls and bytes per call.
- Skipping JavaScript optimisation. Compare against the best JavaScript, not the first draft.
- Profiling only on a fast laptop. Profile where users are.
- Counting idle time as work. Network and timer waits do not benefit from Wasm. Look at CPU samples only.
Performance note
In one import pipeline, the profile showed 58% self time in two parsing functions. After porting the enclosing parseFile region, the import dropped
from 2.1 s to 0.9 s on a mid-range phone; porting the DOM-heavy formatting code as well would have saved under 0.1 s.
Frequently Asked Questions
Can I profile Wasm and JavaScript together? Yes — the same Performance panel shows both, which makes before-and-after comparisons straightforward.
What if the hot code is spread across many small functions? Look for a common entry point in the call tree; port the region, not individual functions.
Is garbage-collection time a good signal? High GC time in a hot region suggests heavy allocation, which a Wasm port with reused buffers often removes.
Should I port code that is already fast enough? No. Port where users wait; everywhere else, simplicity wins.
Does Node’s profiler show the same information?
Yes — --cpu-prof output opens in DevTools with the same views.
What about code that waits on the network? Waiting is not CPU time. Profiles show it as idle; Wasm cannot speed it up.
How long should a profile run? Long enough to cover a realistic session, usually tens of seconds of the workload you want to speed up.
Related
- When porting to Wasm makes code slower — the counter-examples.
- Profiling Wasm with the Chrome Performance panel — profiling after the port.
- Measuring Wasm vs JavaScript throughput — fair comparisons.
- Is WebAssembly faster than JavaScript for DOM manipulation? — why DOM code rarely benefits.
← Back to Porting JavaScript Hot Paths to Wasm