Generating Flame Graphs for Wasm
This page answers one task: produce a flame graph of where a WebAssembly workload spends its time — with readable function names — from whichever environment it runs in, and use it to compare two builds.
Prerequisites
- [ ] A module built with its
namesection intact (release optimizations are fine; stripping is not). - [ ] A profile source: Chrome or Edge DevTools, Node’s
--cpu-prof, or wasmtime with--profile. - [ ] Brendan Gregg’s FlameGraph scripts,
inferno(cargo install inferno), or the speedscope web app for viewing.
Why flame graphs suit Wasm profiles
A profiler records thousands of call stacks — samples of what the program was doing at regular intervals. Tables of self time per function answer “which function is hottest”, but not “how did we get there”. A flame graph answers both at once: each stack is drawn as a column of frames, identical prefixes are merged, and the width of each frame is the fraction of samples in which it appeared. Wide towers are where time goes; the frames beneath them are the call paths that lead there.
That matters for WebAssembly because hot spots in Wasm code are frequently reached through more than one path. A memcpy that takes 12%
of the time is uninteresting until the flame graph shows that 10 of those points come from one caller copying a buffer it did not need
to copy. And because browsers and Node record JavaScript and Wasm frames in the same stacks, a flame graph shows the boundary too —
how much of a task is the export, how much is the JavaScript around it.
Step 1 — get a profile from the browser
Record in Chrome’s Performance panel as usual, then save the profile (the download icon in the panel). The saved file is a Chrome trace in JSON. For a pure CPU profile that tools can read more easily, use the JavaScript Profiler in DevTools (enable it under More tools), or record programmatically:
// run in DevTools console via the Profiler domain, or use console.profile in Chrome
console.profile("filter");
for (let i = 0; i < 20; i++) applyFilter(image); // the workload
console.profileEnd("filter");
The result appears in the JavaScript Profiler panel and can be saved as a .cpuprofile file. That format — nodes with call frames and
sample counts — is what Node writes too, and what most flame graph tools accept directly.
Step 2 — or from Node and wasmtime
Node writes the same format without any UI:
node --cpu-prof --cpu-prof-dir=profiles --cpu-prof-interval 200 scripts/run.mjs
ls profiles/ # CPU.20261002.104212.41822.0.001.cpuprofile
The interval is in microseconds; 200 µs gives five thousand samples per second, enough to resolve short functions. For server-side modules
under wasmtime, use perf with wasmtime’s perfmap support and fold the output, as covered in
profiling Wasm hot paths with perf:
perf record -g -k mono wasmtime run --profile=perfmap app.wasm
perf script | inferno-collapse-perf > stacks.folded
wasmtime can also write a profile in the Firefox Profiler format directly with --profile=guest, which needs no perf and works on any
operating system.
Step 3 — fold and filter the stacks
Flame graph tools work on folded stacks: one line per unique stack, frames separated by semicolons, followed by a sample count. Convert a
.cpuprofile with a few lines of Node:
// fold-cpuprofile.mjs — usage: node fold-cpuprofile.mjs profile.cpuprofile > stacks.folded
import { readFileSync } from "node:fs";
const p = JSON.parse(readFileSync(process.argv[2], "utf8"));
const byId = new Map(p.nodes.map((n) => [n.id, n]));
const parent = new Map();
for (const n of p.nodes) for (const c of n.children ?? []) parent.set(c, n.id);
const counts = new Map();
for (const id of p.samples) counts.set(id, (counts.get(id) ?? 0) + 1);
for (const [id, count] of counts) {
const stack = [];
for (let cur = id; cur !== undefined; cur = parent.get(cur)) {
const f = byId.get(cur).callFrame;
if (f.functionName && f.functionName !== "(root)") stack.push(f.functionName);
}
console.log(stack.reverse().join(";") + " " + count);
}
Then filter out frames that obscure your code — (program), (garbage collector), engine-internal wrappers — or keep them if they are the
point of the investigation:
node fold-cpuprofile.mjs profiles/*.cpuprofile | grep -v '^(program)' > stacks.folded
Step 4 — render and read it
inferno-flamegraph --title "filter: 20 iterations" < stacks.folded > flame.svg
# or: flamegraph.pl stacks.folded > flame.svg
Open the SVG in a browser — any browser, since the file is self-contained; it is interactive — click a frame to zoom into its subtree, search for a function name to highlight every occurrence. Read it bottom-up: the bottom row is the entry point, each row above is a callee, and the plateaus at the top are the functions where samples actually landed. A wide plateau with a single caller beneath is a straightforward hot function; a function that appears in several towers is reached from several paths, and the paths tell you which caller to optimize.
Step 5 — compare two builds with a differential flame graph
The most useful flame graph is often a difference. Profile the same workload with the old and new build, fold both, and render the difference: frames that grew are drawn in red, frames that shrank in blue.
inferno-diff-folded before.folded after.folded | inferno-flamegraph --title "after vs before" > diff.svg
This turns “the new release is 15% slower” into “the new release spends twice as long in utf8_validate under parse_header”, which is
usually enough to find the commit responsible. Differential graphs need comparable workloads — same inputs, same iteration count — and
enough samples on both sides for the colours to mean something.
Choosing where to profile
The same module can be profiled in a browser, in Node, and in a standalone runtime, and the three do not always agree. Browser profiles show the real user’s engine and include the JavaScript boundary and browser work, but they carry DevTools overhead and sampling limits. Node profiles use the same V8 engine without the UI, which makes them convenient for scripted, repeatable runs. wasmtime profiles show a different compiler — Cranelift instead of TurboFan — so hot spots that come from code generation may move. For browser-shipped modules, treat the browser profile as ground truth and use Node for quick iteration; for server-side modules, profile in the runtime you deploy.
Expected output
A flame graph SVG where the widest top-level frames carry Wasm function names such as convolve and blur_rows, with the JavaScript
caller visible at the base. Hovering a frame shows its sample count and percentage:
convolve (2,914 samples, 61.2%)
Gotchas
- Frames named
wasm-function[42]. The module’snamesection was stripped. Profile a build that keeps names. - One giant
(program)frame. The profile includes idle time or startup. Profile only the workload, or filter that frame. - Too few samples. A short workload at a coarse interval produces a noisy graph. Repeat the workload or lower the sampling interval.
- Async work split across stacks. Work scheduled through promises or
postMessageappears as separate stacks without a common caller. Profile the worker or the task that does the work, and read the trigger from the timeline instead. - Differential graph is mostly noise. The two runs did different amounts of work. Use identical inputs and iteration counts.
Performance note
Sampling at 200 µs added about 4% overhead to the Node workload, and folding a ten-second profile took under a second. For the filter above, the differential graph between two releases pointed directly at a dependency upgrade that had replaced a SIMD path with a scalar one — a regression of 22% that had been attributed to “general slowness” for two weeks.
Frequently Asked Questions
Is speedscope a good alternative?
Yes — it opens .cpuprofile files and Chrome traces directly in a browser and offers flame graph, icicle and sandwich views. It is the
quickest way to look at a profile without installing anything.
Can I make flame graphs of memory allocation? Yes, if you record allocation stacks instead of time samples. For Wasm, that means instrumenting the allocator; see profiling allocation hot spots in a Wasm module.
How many samples do I need? A few thousand samples resolve functions that take more than about one percent of the time. For smaller effects, run the workload longer rather than sampling faster, which adds overhead.
Why are inlined functions missing? The engine inlines small functions into callers, and samples land in the caller. That reflects how the code actually ran.
Can I put flame graphs in CI? Generating them per build is cheap; storing differential graphs alongside benchmark results makes regressions much faster to diagnose.
Related
- Profiling Wasm with the Chrome Performance panel — recording in the browser.
- Debugging Wasm in Node.js with the Inspector — Node-side profiling setup.
- Tracking benchmark results in CI — detecting the regressions you then explain with a diff.
- Reading the name custom section — where the frame names come from.
← Back to Debugging & Profiling Wasm Modules