Profiling Wasm with the Chrome Performance Panel

This page answers one task: a page that uses WebAssembly is slow somewhere, and you want to find out where — which Wasm function, which JavaScript around it, or which browser work in between — using Chrome DevTools’ Performance panel.

Prerequisites

  • [ ] Chrome or Edge with DevTools.
  • [ ] A module built with function names kept (the name section); a release build with names is ideal.
  • [ ] A reproducible slow interaction: a button that takes too long, a frame that drops, a slow page load.

What the panel shows for WebAssembly

The Performance panel records a timeline of everything the page’s main thread and workers did: JavaScript, WebAssembly, style and layout, painting, garbage collection, compilation, network events. Its sampling profiler interrupts execution thousands of times per second and records the call stack, so every function that ran long enough appears in the flame chart in proportion to its time. Since Chrome treats Wasm frames as first-class, a Wasm function appears in that chart next to the JavaScript that called it, under its name from the module’s name section.

That combined view is what makes the panel useful for Wasm apps. A slow interaction is often not slow inside the module at all — it is slow because of copying data into linear memory, converting results back into objects, or layout work triggered by applying the result. The flame chart shows all three on one time axis, so you can see which part dominates before deciding what to optimize.

Reading one slow interaction in the flame chart A click handler's flame chart broken into layers from top to bottom — the event handler in JavaScript, the copy into linear memory, the Wasm export and the functions it calls, then the result conversion and the layout and paint that follow. Each layer's width shows its share of the time. click handler (JS) the whole task: 182 ms copy into linear memory Uint8Array.set of 24 MB: 21 ms wasm export → inner functions apply_filter → blur_rows → convolve: 96 ms result → ImageData copy out and putImageData: 18 ms style, layout, paint canvas paint and compositing: 47 ms

Step 1 — record a focused trace

Open DevTools, select the Performance panel, and click the record button. Perform the slow interaction once, then stop. Keep traces short — a few seconds around the problem — so the timeline is readable.

Before recording, open the panel’s capture settings and enable Advanced paint instrumentation only if paint is suspect; it adds overhead. Set CPU throttling to 4× or 6× slowdown to approximate a mid-range phone, which makes problems that hide on a fast laptop visible. Close other tabs to reduce noise.

For page-load problems, use the reload-and-record button instead, which captures navigation, module download and compilation together.

Step 2 — find the Wasm frames

In the Main track, zoom into the long task (marked with a red corner when it exceeds 50 ms). Wasm functions appear as frames with their names from the module — apply_filter, blur_rows — and a source location of wasm://wasm/<hash>. Clicking one shows its self time and total time in the Summary tab.

If frames show as wasm-function[123] instead of names, the module was stripped of its name section. Rebuild keeping names — for Rust, wasm-bindgen --keep-debug or wasm-opt --debuginfo; for Emscripten, --profiling-funcs — and record again. Names cost size, so profile a build that keeps them even if production strips them; the trade-off is discussed in reading the name custom section.

Step 3 — use Bottom-Up to rank functions

The flame chart shows when; the Bottom-Up tab shows how much. It aggregates self time per function across the whole recording, which answers “which function costs the most” directly:

Self Time     Total Time    Activity
  71.4 ms 39%   71.4 ms 39%   convolve          wasm://wasm/8c1e2f1a
  21.0 ms 12%   21.0 ms 12%   (anonymous) set   main.js:44
  18.2 ms 10%   18.2 ms 10%   putImageData      main.js:61
  14.8 ms  8%   96.2 ms 53%   blur_rows         wasm://wasm/8c1e2f1a
   9.9 ms  5%    9.9 ms  5%   Layout

convolve dominates self time inside the module, which makes it the candidate for SIMD or algorithmic work. The 21 ms in set is the copy into linear memory — a boundary cost that no change to the Wasm code will fix, but that can be avoided entirely, as discussed in avoiding copies when passing image buffers.

Step 4 — look for compilation and GC

Not all time in a Wasm app is execution. Two other kinds of work show up in traces and are easy to overlook.

Compilation appears as tasks named v8.wasm.compile* or CompileLazy on the main thread or on background threads. On first load the whole module is compiled by the baseline tier, and hot functions are recompiled by the optimizing tier later; a trace that includes the first use of a feature may show its first call slower than later ones for this reason, as explained in why the first call into Wasm is slow.

Garbage collection appears as Minor GC and Major GC slices. WebAssembly linear memory is not garbage-collected, but the JavaScript around it is: typed array views created per call, result objects built from Wasm output, closures. Frequent minor GCs inside a hot loop are a sign that the JavaScript side of the boundary allocates too much.

What to look for in a Wasm trace and what it means Common patterns in a Chrome Performance trace of a WebAssembly page, how to recognise each, and what it usually indicates. pattern in trace where usually means one Wasm function dominates self time Bottom-Up algorithm or SIMD work large Uint8Array.set or slice around the export copying at the boundary v8.wasm.compile on main thread first use compile in a worker or earlier many Minor GC slices during the loop JS allocations per call Layout / Paint after the call after the export DOM work, not Wasm

Step 5 — profile workers too

Heavy Wasm work usually runs in workers, and each worker has its own track in the Performance panel, below Main. Expand it to see the worker’s flame chart; the same Bottom-Up view is available by selecting a range on the worker track. A common finding is a worker that is busy for most of the trace while the main thread waits on postMessage — which is correct — or the reverse, a main thread blocked because work that should have gone to the worker did not.

Interpreting a trace before optimizing

Read a trace in this order and you will rarely optimize the wrong thing. First, how long is the task as a whole, and does it matter — a 40 ms task during a page transition may be fine, the same task on every keystroke is not. Second, how is that time split between the Wasm call itself, the boundary work around it, and browser work after it; the flame chart answers that at a glance. Only then, third, which function inside the module is hottest. Many performance investigations of Wasm apps end at step two, because the module was fast and the time went into copies or layout — and no amount of SIMD in convolve would have helped.

Expected output

A short annotated finding, which is the useful output of a profiling session:

filter click, 182 ms total at 4× CPU throttling
  96 ms in Wasm (convolve 71 ms self)    → SIMD candidate
  39 ms copying in/out of linear memory  → keep image resident in Wasm memory
  47 ms paint                            → draw with a single putImageData per frame

Gotchas

  • Function names missing. The module is stripped. Profile a build with the name section.
  • DevTools changes timings. Recording adds overhead, and an open DevTools can affect optimization. Compare relative costs within a trace rather than absolute numbers with production.
  • Warm-up not captured. A trace of the first interaction includes baseline-tier code and compilation. Record a second interaction to see steady-state behaviour.
  • Extensions in the trace. Browser extensions inject scripts that appear in recordings. Profile in a clean profile or a guest window.
  • Profiling a debug build. Unoptimized builds have different hot spots. Profile optimized code with names kept.

Performance note

Acting on the trace above — keeping the image resident in linear memory and vectorizing convolve — took the interaction from 182 ms to 71 ms at 4× throttling. Most of the improvement came from removing the copies, not from the SIMD work, which is typical and is why the boundary deserves a look first.

The same interaction before and after acting on the trace Task duration for the filter click at 4× CPU throttling: as recorded, after removing the copies into and out of linear memory, and after also vectorizing the hot convolution. ms per click at 4× CPU throttling as recorded 182 ms image kept in Wasm memory 118 ms + SIMD convolve 71 ms

Frequently Asked Questions

Can I see source lines rather than function names? With DWARF debug information and the C/C++ DevTools Support extension, the Sources panel maps Wasm to source lines; the Performance panel stays at function granularity. See using the C/C++ DevTools Support extension.

Why do some Wasm functions not appear even though they run? Very short functions inlined by the engine or taking less than a sampling interval may not be sampled. They are accounted to their callers.

Is the profiler accurate enough for small differences? It is a sampling profiler, so differences of a few percent are noise. Use it to find where time goes and a benchmark harness to measure changes precisely.

Does throttling affect Wasm differently from JavaScript? CPU throttling slows the whole renderer uniformly, so the relative split between Wasm, JavaScript and browser work stays representative. Real phones also differ in memory bandwidth and caches, so confirm important findings on a device.

How do I share a trace? Save the profile from the panel’s menu as a JSON file; it opens in any Chrome DevTools, preserving names and timings.

← Back to Debugging & Profiling Wasm Modules