Profiling Wasm in Production with Sampling
This page answers one task: a WebAssembly feature is fast on developers’ machines, yet field data shows it is slow for some users, and nobody can reproduce it. You want CPU profiles from production — which functions take time on real devices with real inputs — collected at low overhead from a small sample of sessions.
Prerequisites
- [ ] A Chromium-based audience for in-browser profiling (the JS Self-Profiling API is Chromium-only at present).
- [ ] The ability to set response headers on your HTML (
Document-Policy). - [ ] Function names for your module: a name section in production builds or archived symbol files.
What field profiles add
Lab profiling shows where time goes on the machines and inputs you chose. Field profiling shows where time goes for users: low-end devices where a function that is negligible on a laptop dominates, inputs ten times larger than test files, browser versions with different tier-up behaviour, contention with extensions and other tabs. Aggregated across thousands of sessions, field profiles reveal the hot paths that matter in practice, and comparing them between releases shows whether an optimisation helped real users.
In browsers, the JS Self-Profiling API lets a page sample its own call stacks — including WebAssembly frames — at a set interval, returning a compact trace.
On servers, sampling profilers attached to the host process (Wasmtime with its profiling support, Node with --cpu-prof or continuous profilers) do the same
for server-side Wasm.
Step 1 — enable the API with Document-Policy
The page must opt in with a response header on the HTML document:
Document-Policy: js-profiling
Without it, constructing a Profiler throws. Setting the header enables the capability; profiling only happens when your code starts a profiler. In some
Chromium versions, enabling the policy has a small cost even without active profiling (it can affect script compilation); measure, and consider sending the
header only for the sampled sessions if your server can decide per request.
Step 2 — profile a sample of sessions around the feature
async function withProfiling(label, fn) {
if (!("Profiler" in globalThis) || Math.random() > 0.01) return fn(); // 1% of eligible sessions
const profiler = new Profiler({ sampleInterval: 10, maxBufferSize: 10_000 });
try {
return await fn();
} finally {
const trace = await profiler.stop();
navigator.sendBeacon("/profiles", JSON.stringify({
label, module: MODULE_HASH, app: APP_VERSION, device: deviceClass(), trace,
}));
}
}
await withProfiling("export-pdf", () => exportPdf(doc));
Profile around specific operations rather than whole sessions: traces are smaller, and aggregation by operation is more meaningful. The browser may enforce a minimum sample interval; ask for 10 ms and accept what it grants.
Step 3 — symbolicate Wasm frames
The trace format lists frames with names, resource URLs and line or column information. Wasm frames appear with function names if the module carries a name
section; otherwise they appear as function indices (wasm-function[412]). Two options: keep the name section in production (it adds size, often 10–20%
uncompressed before Brotli) or strip it and symbolicate server-side using archived name maps keyed by module hash. The second keeps downloads small and is
the same machinery used for crash reports; see
symbolicating Wasm stack traces in production.
Step 4 — aggregate into flame graphs
Individual traces are noisy. Aggregate samples by stack across all sessions for an operation, release and device class, and render flame graphs or top-function tables. Normalise by operation count, not by total samples, so a release with more users does not look slower. The most actionable views are “top self-time functions in the export operation on low-end Android” and “difference between release N and N−1 for the same operation”.
Step 5 — profile server-side Wasm too
For Wasm running in servers, use the host’s tools. Wasmtime supports profiling integrations (for example emitting perf maps so Linux perf can attribute JIT
code to Wasm functions, or its built-in sampling profiler in recent versions); Node exposes V8’s CPU profiler via --cpu-prof and the inspector; continuous
profilers that sample production processes at low frequency can include Wasm frames when the runtime provides symbol information. Sample a fraction of
instances, keep the overhead budget explicit, and tag profiles with the module version.
Privacy and overhead
Profiles contain function names and URLs of your code, not user data — but URLs can include query strings, and labels you add could. Strip query strings and review what you attach. Overhead of sampling at 10 ms is typically low (a few percent while active), and sampling 1% of sessions makes the aggregate cost negligible. Respect consent requirements for performance data in your jurisdiction.
Lightweight timing instrumentation as a complement
Sampling profilers show where time goes inside an operation, but only in browsers that support them and only for sampled sessions. Timing marks cover every
browser and every session at almost no cost: wrap major phases of a Wasm operation — parse, layout, render, encode — with performance.now() around the calls
(or performance.measure for visibility in DevTools), and report the phase durations with the same release and device context as the profiles. The phases tell
you which part of an operation got slower for Safari users after a release; profiles from Chromium users then tell you which functions within that phase
to look at. Inside the module, a few exported counters of work done (pages laid out, glyphs shaped, bytes encoded) let you normalise durations by workload
size, distinguishing “slower code” from “users processing larger documents”.
Turning field profiles into lab reproductions
A field profile points at a hot function, but optimising it needs a reproducible case. Use the context attached to the profile — device class, operation, input size bucket — to build a lab benchmark with a similar input, run it on a matching device, and confirm it shows the same hot spot before changing code. After the fix, verify twice: in the lab benchmark, and in the next release’s aggregated field profiles for the same device class, because lab improvements do not always translate to the field. Recording call traces with consent provides the most faithful lab reproductions when inputs are the cause.
Retention and storage
Profiles are small but numerous. Keep raw traces for a short period (days to weeks) for investigation and store aggregates — per-release, per-operation flame graphs — longer for trend comparison.
Expected output
One percent of sessions in Chromium browsers send profiles of the PDF export; aggregated flame graphs for low-end Android show 41% of export time in a font-subsetting function that takes 4% on desktops; after optimising it, the next release’s aggregated export time on those devices falls by a third; and server-side Wasm plugins are profiled with Wasmtime’s sampler on 2% of instances.
Gotchas
- Forgetting
Document-Policy: js-profiling. TheProfilerconstructor throws. Set the header. - Profiling whole sessions. Large, unfocused traces. Profile around operations.
- Stripped names without archived maps. Frames are indices. Archive symbols by module hash.
- Comparing raw sample counts across releases. Traffic differs. Normalise per operation.
- Query strings in trace URLs. Possible data leakage. Strip them.
- Profiles from one browser family only. Conclusions skew to Chromium. Add phase timings for all browsers.
Performance note
With a 10 ms sample interval, active profiling added about 2–4% CPU during the profiled operation; at a 1% session sampling rate the fleet-wide overhead was unmeasurable.
Frequently Asked Questions
Does the API work in Firefox or Safari? Not at present; field profiles therefore represent Chromium users only.
Can I profile inside workers? Support for profiling in workers has been limited; profile on the thread where the API is available or instrument workers with timing marks.
Is a 10 ms interval precise enough? For aggregated profiles across many sessions, yes; individual traces are coarse.
Do I need a vendor tool? No — a small endpoint and an aggregation script work; vendors add convenience and UI.
What covers browsers without the profiling API? Phase timing marks around Wasm calls, reported with release and device context, work in every browser at negligible cost.
How do I distinguish slower code from larger inputs? Report workload counters such as pages or bytes processed and normalise durations by them.
How long should raw profiles be kept? Days to weeks for investigation; keep per-release aggregated flame graphs longer for trends.
How do I confirm a fix helped real users? Compare the next release’s aggregated profiles and phase timings for the same operation and device class.
Should profiling be enabled for every user? No — a small random share is enough for aggregate flame graphs and keeps overhead and data volume negligible.
Related
- Measuring Wasm performance with real-user monitoring — timing metrics.
- Symbolicating Wasm stack traces in production — names for frames.
- Profiling Wasm hot paths with perf — server-side profiling.
- Benchmarking Wasm on mobile devices — lab follow-up.
← Back to Observability & Error Reporting