Profiling Wasm in Production with Sampling

This page answers one task: a WebAssembly feature is fast on developers’ machines, yet field data shows it is slow for some users, and nobody can reproduce it. You want CPU profiles from production — which functions take time on real devices with real inputs — collected at low overhead from a small sample of sessions.

Prerequisites

  • [ ] A Chromium-based audience for in-browser profiling (the JS Self-Profiling API is Chromium-only at present).
  • [ ] The ability to set response headers on your HTML (Document-Policy).
  • [ ] Function names for your module: a name section in production builds or archived symbol files.

What field profiles add

Lab profiling shows where time goes on the machines and inputs you chose. Field profiling shows where time goes for users: low-end devices where a function that is negligible on a laptop dominates, inputs ten times larger than test files, browser versions with different tier-up behaviour, contention with extensions and other tabs. Aggregated across thousands of sessions, field profiles reveal the hot paths that matter in practice, and comparing them between releases shows whether an optimisation helped real users.

In browsers, the JS Self-Profiling API lets a page sample its own call stacks — including WebAssembly frames — at a set interval, returning a compact trace. On servers, sampling profilers attached to the host process (Wasmtime with its profiling support, Node with --cpu-prof or continuous profilers) do the same for server-side Wasm.

Field profiling of a Wasm feature A small random share of sessions start a JS Self-Profiler when the Wasm feature runs. Samples of call stacks, including Wasm frames, are collected at a 10 ms interval, then stopped and sent with the module hash. The backend symbolicates Wasm frames with archived symbols and aggregates samples across sessions into flame graphs per release and device class. sample 1% of sessions Document-Policy enabled profile around feature 10 ms interval send trace + module hash compact JSON symbolicate Wasm frames archived names aggregate flame graphs per release, device

Step 1 — enable the API with Document-Policy

The page must opt in with a response header on the HTML document:

Document-Policy: js-profiling

Without it, constructing a Profiler throws. Setting the header enables the capability; profiling only happens when your code starts a profiler. In some Chromium versions, enabling the policy has a small cost even without active profiling (it can affect script compilation); measure, and consider sending the header only for the sampled sessions if your server can decide per request.

Step 2 — profile a sample of sessions around the feature

async function withProfiling(label, fn) {
  if (!("Profiler" in globalThis) || Math.random() > 0.01) return fn();   // 1% of eligible sessions
  const profiler = new Profiler({ sampleInterval: 10, maxBufferSize: 10_000 });
  try {
    return await fn();
  } finally {
    const trace = await profiler.stop();
    navigator.sendBeacon("/profiles", JSON.stringify({
      label, module: MODULE_HASH, app: APP_VERSION, device: deviceClass(), trace,
    }));
  }
}

await withProfiling("export-pdf", () => exportPdf(doc));

Profile around specific operations rather than whole sessions: traces are smaller, and aggregation by operation is more meaningful. The browser may enforce a minimum sample interval; ask for 10 ms and accept what it grants.

Step 3 — symbolicate Wasm frames

The trace format lists frames with names, resource URLs and line or column information. Wasm frames appear with function names if the module carries a name section; otherwise they appear as function indices (wasm-function[412]). Two options: keep the name section in production (it adds size, often 10–20% uncompressed before Brotli) or strip it and symbolicate server-side using archived name maps keyed by module hash. The second keeps downloads small and is the same machinery used for crash reports; see symbolicating Wasm stack traces in production.

Shipping names versus server-side symbolication Keeping the name section in production makes Wasm frames readable in every trace immediately, at the cost of a larger download. Stripping names and symbolicating on the server keeps the module small but requires archiving name maps for every released build. ship the name section readable traces everywhere larger module no symbol infrastructure simple strip + symbolicate server-side smaller download names from archived maps shared with crash reporting scalable

Step 4 — aggregate into flame graphs

Individual traces are noisy. Aggregate samples by stack across all sessions for an operation, release and device class, and render flame graphs or top-function tables. Normalise by operation count, not by total samples, so a release with more users does not look slower. The most actionable views are “top self-time functions in the export operation on low-end Android” and “difference between release N and N−1 for the same operation”.

Step 5 — profile server-side Wasm too

For Wasm running in servers, use the host’s tools. Wasmtime supports profiling integrations (for example emitting perf maps so Linux perf can attribute JIT code to Wasm functions, or its built-in sampling profiler in recent versions); Node exposes V8’s CPU profiler via --cpu-prof and the inspector; continuous profilers that sample production processes at low frequency can include Wasm frames when the runtime provides symbol information. Sample a fraction of instances, keep the overhead budget explicit, and tag profiles with the module version.

Privacy and overhead

Profiles contain function names and URLs of your code, not user data — but URLs can include query strings, and labels you add could. Strip query strings and review what you attach. Overhead of sampling at 10 ms is typically low (a few percent while active), and sampling 1% of sessions makes the aggregate cost negligible. Respect consent requirements for performance data in your jurisdiction.

Lightweight timing instrumentation as a complement

Sampling profilers show where time goes inside an operation, but only in browsers that support them and only for sampled sessions. Timing marks cover every browser and every session at almost no cost: wrap major phases of a Wasm operation — parse, layout, render, encode — with performance.now() around the calls (or performance.measure for visibility in DevTools), and report the phase durations with the same release and device context as the profiles. The phases tell you which part of an operation got slower for Safari users after a release; profiles from Chromium users then tell you which functions within that phase to look at. Inside the module, a few exported counters of work done (pages laid out, glyphs shaped, bytes encoded) let you normalise durations by workload size, distinguishing “slower code” from “users processing larger documents”.

Turning field profiles into lab reproductions

A field profile points at a hot function, but optimising it needs a reproducible case. Use the context attached to the profile — device class, operation, input size bucket — to build a lab benchmark with a similar input, run it on a matching device, and confirm it shows the same hot spot before changing code. After the fix, verify twice: in the lab benchmark, and in the next release’s aggregated field profiles for the same device class, because lab improvements do not always translate to the field. Recording call traces with consent provides the most faithful lab reproductions when inputs are the cause.

Retention and storage

Profiles are small but numerous. Keep raw traces for a short period (days to weeks) for investigation and store aggregates — per-release, per-operation flame graphs — longer for trend comparison.

Expected output

One percent of sessions in Chromium browsers send profiles of the PDF export; aggregated flame graphs for low-end Android show 41% of export time in a font-subsetting function that takes 4% on desktops; after optimising it, the next release’s aggregated export time on those devices falls by a third; and server-side Wasm plugins are profiled with Wasmtime’s sampler on 2% of instances.

Gotchas

  • Forgetting Document-Policy: js-profiling. The Profiler constructor throws. Set the header.
  • Profiling whole sessions. Large, unfocused traces. Profile around operations.
  • Stripped names without archived maps. Frames are indices. Archive symbols by module hash.
  • Comparing raw sample counts across releases. Traffic differs. Normalise per operation.
  • Query strings in trace URLs. Possible data leakage. Strip them.
  • Profiles from one browser family only. Conclusions skew to Chromium. Add phase timings for all browsers.

Performance note

With a 10 ms sample interval, active profiling added about 2–4% CPU during the profiled operation; at a 1% session sampling rate the fleet-wide overhead was unmeasurable.

Share of export time in the top function by device class Percentage of PDF export time spent in a font-subsetting function according to aggregated field profiles, on desktop, mid-range Android and low-end Android devices. share of export time (%) desktop 4 % mid-range Android 19 % low-end Android 41 %

Frequently Asked Questions

Does the API work in Firefox or Safari? Not at present; field profiles therefore represent Chromium users only.

Can I profile inside workers? Support for profiling in workers has been limited; profile on the thread where the API is available or instrument workers with timing marks.

Is a 10 ms interval precise enough? For aggregated profiles across many sessions, yes; individual traces are coarse.

Do I need a vendor tool? No — a small endpoint and an aggregation script work; vendors add convenience and UI.

What covers browsers without the profiling API? Phase timing marks around Wasm calls, reported with release and device context, work in every browser at negligible cost.

How do I distinguish slower code from larger inputs? Report workload counters such as pages or bytes processed and normalise durations by them.

How long should raw profiles be kept? Days to weeks for investigation; keep per-release aggregated flame graphs longer for trends.

How do I confirm a fix helped real users? Compare the next release’s aggregated profiles and phase timings for the same operation and device class.

Should profiling be enabled for every user? No — a small random share is enough for aggregate flame graphs and keeps overhead and data volume negligible.

← Back to Observability & Error Reporting