Monitoring Wasm Memory in Production

This page answers one task: WebAssembly memory problems — leaks, fragmentation, oversized peaks — usually show up as slow tabs, reloaded pages on phones, or out-of-memory errors in the field, and you want to see them in your monitoring before users report them.

Prerequisites

  • [ ] A module whose memory you can read (memory.buffer.byteLength) and, ideally, allocator statistics exported from the module.
  • [ ] A real-user monitoring pipeline, as in measuring Wasm performance with real-user monitoring.
  • [ ] For server-side guests, a metrics system (Prometheus, OpenTelemetry metrics).

Why memory needs its own monitoring

Wasm memory behaves differently from JavaScript memory. Linear memory only grows — it never shrinks while the instance lives — so a single large operation leaves the module at its peak for the rest of the session, as explained in why Wasm memory never shrinks. Leaks inside the module are invisible to the JavaScript heap profiler. And the consequences differ by platform: on desktop, a large module mostly costs RAM; on phones, especially iOS, the operating system may silently kill the tab, which looks to monitoring like a session that simply ended.

So monitor three things: the size of linear memory over a session, the live bytes the allocator reports, and failures — growth that failed, out-of-memory traps, and sessions that ended abruptly after memory climbed.

Interpreting memory telemetry If live bytes grow steadily over a session, suspect a leak. If size jumps once and stays flat, a large operation set the peak. If size keeps rising while live bytes stay flat, suspect fragmentation. If sessions end abruptly at high memory on phones, the operating system is killing tabs. What does the session's memory curve look like? live bytes climbing leak — find what is not freed one jump, then flat peak from a large operation size up, live flat fragmentation

Step 1 — sample memory size and allocator statistics

Linear memory size is free to read. Allocator statistics need an export from the module, as shown in measuring allocator fragmentation in Wasm:

function memorySample() {
  const [heap, live] = wasm.heap_stats?.() ?? [0, 0];
  return {
    t: Math.round(performance.now() / 1000),
    size: wasm.memory.buffer.byteLength,
    heap, live,
  };
}

const samples = [];
const timer = setInterval(() => { samples.push(memorySample()); if (samples.length > 120) samples.shift(); }, 30_000);

Sample every 30–60 seconds; memory changes slowly, and more frequent sampling adds cost without adding insight. Also sample immediately after known heavy operations — opening a document, decoding a large image — with the operation name, which makes peaks attributable to the feature that caused them.

Step 2 — report summaries, not raw series

Send a compact summary at the end of the session (on visibilitychange to hidden) and periodically for long sessions:

function memorySummary() {
  const sizes = samples.map((s) => s.size), lives = samples.map((s) => s.live);
  return {
    build: WASM_BUILD_ID,
    peakSize: Math.max(...sizes),
    endSize: sizes.at(-1),
    liveSlopePerHour: slope(samples.map((s) => [s.t / 3600, s.live])),   // bytes/hour, least squares
    sessionMinutes: Math.round(performance.now() / 60000),
    device: { memory: navigator.deviceMemory, cores: navigator.hardwareConcurrency },
  };
}

The live-bytes slope over time is the best single leak indicator: near zero in healthy sessions, consistently positive when something leaks. The peak size tells you how close users get to device limits.

Step 3 — record failures explicitly

Out-of-memory conditions should be reported as events, not just inferred. Catch RangeError from memory.grow in JavaScript, out-of-memory traps and allocation-failure errors from the module, and report them with the current size and the operation in progress, as described in handling out-of-memory in Wasm. For tab kills, which cannot be caught, record a “session alive” heartbeat with the current memory size in localStorage; on the next page load, if the previous session ended without a clean pagehide and its last heartbeat showed high memory, report a probable out-of-memory termination.

Detecting tabs killed for memory While the page runs, a heartbeat stores the current memory size and time. A clean page hide marks the session as ended normally. On the next load, a session without a clean end whose last heartbeat showed high memory is reported as a probable out-of-memory kill. heartbeat every 30 s size + time → localStorage pagehide mark clean exit next page load read last session no clean exit + high memory? probable OOM kill report with build + device aggregate per release

Step 4 — monitor server-side guests

On the server, the host can read each instance’s memory size directly — instance.get_memory(&mut store, "memory")?.data_size(&store) in wasmtime — and record it after every invocation as a histogram labelled by guest. With a ResourceLimiter, the host also sees every growth request and can count denials. Long-lived instances that serve many requests should export their peak and current size as gauges, which reveals guests whose memory creeps upward across requests. Per-request instances need only the per-invocation peak.

Alert on changes, not absolute values, since normal memory use differs widely between features. Useful alerts: the 95th-percentile peak size per release rising more than 20% over the previous release; the share of sessions with a positive live-bytes slope rising; out-of-memory failures or probable tab kills exceeding a small rate on any device segment. Because memory regressions usually come from code changes, comparing releases catches most of them within a day of deployment.

Building the dashboard

A memory dashboard that people actually use answers a few questions at a glance. Show, per release, the distribution of peak linear memory as percentiles, so regressions stand out as a step between releases. Show the share of sessions whose live bytes grow by more than a threshold per hour, which is the leak signal. Show out-of-memory events and probable tab kills per thousand sessions, split by device memory class, because phones with 2–4 GB behave very differently from laptops. Add a table of the operations most often recorded right before a peak, which points engineers at the code to look at. Keep raw per-session data for a short window so that, when a chart moves, someone can drill into individual sessions with the same build and device to see their memory curves. Annotate the charts with deploys, so the first question — “what changed?” — answers itself.

Choosing sampling rates and retention

Memory telemetry is low-volume compared with performance traces, so most applications can afford to collect it from every session as a single summary beacon, with full time series from a small sample — one or two percent — for drill-downs. Keep summaries for months, to compare releases and seasons, and detailed series for a couple of weeks. Avoid collecting anything that could identify users or reveal their content; memory sizes, operation names and device classes are enough.

Linking memory to user impact

Memory numbers matter when they change what users experience. Correlate them with outcomes you already measure: sessions that end abruptly, features that fail, input latency that rises as memory grows (allocation and garbage collection pressure), and conversion or task completion on memory-heavy flows. Segment by device memory: a peak of 600 MB is unremarkable on a 16 GB laptop and fatal on a 3 GB phone. When telemetry shows a problem, reproduce it with the same operation and input sizes in a lab session and use the techniques in tracking linear memory growth over time to find the cause. Often the fix is not a leak at all but an operation that holds two copies of a large buffer, or a feature that should run in a disposable worker so its memory is returned when it finishes.

Expected output

The dashboard shows per release the median and 95th-percentile peak Wasm memory, the share of sessions with growing live bytes, and out-of-memory events by device class; a release that introduced a leak in the undo history appears as a jump in positive slopes within hours, and the probable tab-kill rate on iOS returns to baseline after the fix.

Gotchas

  • Watching only the JavaScript heap. Linear memory is separate. Sample memory.buffer.byteLength.
  • Treating size as usage. Size never shrinks. Use allocator live bytes for leaks.
  • Missing tab kills. They cannot be caught. Infer them from heartbeats.
  • Absolute thresholds. Normal use varies by feature. Alert on changes between releases.
  • No operation names on peaks. A peak without context is hard to act on. Sample after heavy operations and name them.
  • Ignoring device memory. The same peak means different things on different devices. Segment.

Performance note

Sampling memory size and allocator statistics every 30 seconds cost under 0.1 ms per sample. The session summary added about 400 bytes to the existing RUM beacon. A leak of 2 KB per edit operation was visible in field data as a live-bytes slope of about 3 MB per hour in active sessions.

Peak Wasm memory at the 95th percentile by release Megabytes of peak linear memory at the 95th percentile of sessions for three consecutive releases, showing a regression in the middle release and the fix in the next. MB peak memory (p95) release 41 180 MB release 42 (leak) 410 MB release 43 (fixed) 185 MB

Frequently Asked Questions

Can I read the memory of a worker’s module from the page? Not directly unless the memory is shared. Have the worker sample and post its numbers.

Is performance.measureUserAgentSpecificMemory better? It gives a whole-page total in isolated Chromium pages; use it alongside, not instead of, per-module sampling.

How do I monitor memory in long-running server instances? Export current and peak size as gauges per guest, and recycle instances that exceed a threshold.

Does monitoring itself use much memory? No — a bounded array of samples and a summary object are a few kilobytes.

How do I monitor memory in Electron or other embedded browsers? The same in-page sampling works; Electron additionally exposes process memory APIs to the main process.

Should memory data trigger automatic actions? On servers, yes — recycle instances above a threshold. In browsers, a module can offer to reload a heavy feature in a fresh worker.

← Back to Observability & Error Reporting