Counting Allocations with a Wrapping Allocator

This page answers one task: you want to know how many allocations a Rust WebAssembly function makes and how many bytes stay live afterwards — to find leaks, to cut allocation in hot paths, and to stop regressions — without a sanitizer or an external profiler.

Prerequisites

  • [ ] A Rust crate targeting wasm32-unknown-unknown (or a WASI target).
  • [ ] The ability to add a #[global_allocator] in a feature-gated or test build.
  • [ ] A way to call the module from JavaScript or wasm-bindgen-test.

The idea: intercept every allocation

Every heap allocation in a Rust program — Box, Vec, String, HashMap, everything in the standard library — goes through the global allocator, a type implementing the GlobalAlloc trait. Rust lets a program replace it with #[global_allocator]. A replacement does not have to implement allocation itself: it can forward every call to the real allocator and, on the way, update counters. That gives exact numbers with negligible overhead: how many times alloc was called, how many times dealloc, how many bytes are live right now, and the peak.

These four counters answer the most common memory questions. A function whose live-bytes count is higher after a call than before has leaked or cached something. A hot function with thousands of allocations per call is a candidate for buffer reuse. A test that asserts an allocation count turns “we accidentally started cloning the input” from a slow production regression into a failing CI check.

A counting allocator between Rust code and the real allocator Rust code allocates through the global allocator. The counting wrapper increments counters for calls and bytes, then forwards to the real allocator. JavaScript reads the counters through an exported function. Rust code Vec, String, Box … counting wrapper #[global_allocator] update counters allocs, frees, live, peak real allocator dlmalloc / System JavaScript reads stats()

Step 1 — write the wrapper

use std::alloc::{GlobalAlloc, Layout, System};
use std::sync::atomic::{AtomicUsize, Ordering::Relaxed};

pub struct Counting;

static ALLOCS: AtomicUsize = AtomicUsize::new(0);
static FREES: AtomicUsize = AtomicUsize::new(0);
static LIVE: AtomicUsize = AtomicUsize::new(0);
static PEAK: AtomicUsize = AtomicUsize::new(0);

unsafe impl GlobalAlloc for Counting {
    unsafe fn alloc(&self, layout: Layout) -> *mut u8 {
        let p = System.alloc(layout);
        if !p.is_null() {
            ALLOCS.fetch_add(1, Relaxed);
            let live = LIVE.fetch_add(layout.size(), Relaxed) + layout.size();
            PEAK.fetch_max(live, Relaxed);
        }
        p
    }
    unsafe fn dealloc(&self, p: *mut u8, layout: Layout) {
        System.dealloc(p, layout);
        FREES.fetch_add(1, Relaxed);
        LIVE.fetch_sub(layout.size(), Relaxed);
    }
    unsafe fn realloc(&self, p: *mut u8, layout: Layout, new_size: usize) -> *mut u8 {
        let q = System.realloc(p, layout, new_size);
        if !q.is_null() {
            ALLOCS.fetch_add(1, Relaxed);
            FREES.fetch_add(1, Relaxed);
            let live = LIVE.fetch_add(new_size, Relaxed) + new_size - layout.size();
            LIVE.fetch_sub(layout.size(), Relaxed);
            PEAK.fetch_max(live, Relaxed);
        }
        q
    }
}

#[cfg(feature = "alloc-stats")]
#[global_allocator]
static GLOBAL: Counting = Counting;

System on wasm32-unknown-unknown is the standard library’s bundled allocator (a dlmalloc port), so the wrapper adds counting without changing allocation behaviour. Atomics with Relaxed ordering are cheap and remain correct if the module is later built with threads. Counting realloc as one allocation and one free keeps ALLOCS - FREES equal to the number of live blocks.

Step 2 — expose the counters

#[wasm_bindgen]
pub fn alloc_stats() -> Vec<u32> {
    [&ALLOCS, &FREES, &LIVE, &PEAK].iter().map(|c| c.load(Relaxed) as u32).collect()
}

#[wasm_bindgen]
pub fn reset_peak() { PEAK.store(LIVE.load(Relaxed), Relaxed); }

Feature-gating the allocator keeps release builds untouched: build with --features alloc-stats for profiling and tests, without it for production. The overhead is small enough that some teams leave it on in production and report the numbers as telemetry, but gating is the conservative default.

Step 3 — measure one operation

From JavaScript, read the counters before and after a call:

function measure(fn) {
  const [a0, f0, l0] = wasm.alloc_stats();
  wasm.reset_peak();
  const result = fn();
  const [a1, f1, l1, peak] = wasm.alloc_stats();
  return { result, allocs: a1 - a0, frees: f1 - f0, liveDelta: l1 - l0, peak: peak - l0 };
}

console.table(measure(() => wasm.parse_document(bytes)));
// allocs 18,442   frees 18,409   liveDelta 2,112   peak 6,291,456

That one line tells a lot: the parse made eighteen thousand allocations, temporarily needed 6 MB above the starting level, and left 2 KB and 33 blocks live afterwards. If the returned document accounts for those 2 KB, fine; if the document has been freed and liveDelta is still positive, something leaked or was cached.

Reading the counters after an operation High allocation counts suggest reusing buffers. A positive live delta after everything is freed suggests a leak or cache. A high peak relative to the result suggests intermediate copies. Balanced counts with a flat live total indicate a clean operation. observation likely cause next step allocs in the thousands per call many small temporaries reuse buffers live delta > 0 after free leak or hidden cache repeat and compare peak far above result size intermediate copies stream or reserve allocs = frees and live flat clean operation keep as a test

Step 4 — find leaks by repetition

A single positive liveDelta might be a lazily initialised cache. A leak grows with every repetition. Run the operation, including its cleanup, many times and watch live bytes:

for (let i = 0; i < 5; i++) {
  for (let j = 0; j < 100; j++) wasm.parse_document(bytes).free();
  console.log(`round ${i}: live = ${wasm.alloc_stats()[2]}`);
}
// round 0: live = 41,280   round 1: live = 251,280   round 2: live = 461,280 …

Live bytes rising by 210,000 per hundred iterations — 2,100 bytes per call — is a leak of about that size per document. Bisect by measuring the operation’s internal phases separately, or add the same counters per phase with a scoped guard, until the leaking step is isolated. JavaScript-side leaks of wasm-bindgen objects show up the same way, and are covered in detecting forgotten free calls in wasm-bindgen.

Step 5 — lock in allocation budgets with tests

Once an operation is tuned, assert its allocation count in a test so a later change cannot silently regress it:

#[wasm_bindgen_test]
fn render_frame_allocates_at_most_16_times() {
    let (a0, _) = (ALLOCS.load(Relaxed), ());
    render_frame(0.0);
    assert!(ALLOCS.load(Relaxed) - a0 <= 16, "render_frame allocation budget exceeded");
}

Budgets work best for hot paths — per-frame rendering, per-message parsing — where allocation counts are stable and meaningful. Allow headroom so that harmless changes do not fail, and update the budget deliberately when a change justifies it.

Attributing allocations to call sites

Counters say how much; they do not say where. To find which code allocates most, extend the wrapper to record a small histogram by size class, and during investigation capture a stack trace for a sample of allocations — every thousandth, say — through an imported JavaScript function that calls new Error().stack. With names in the binary (keep the name section in the profiling build), those stacks include the Rust function names, and a few seconds of sampling usually points straight at the hottest allocation sites. Aggregating the sampled stacks into counts per function gives a crude but effective allocation profile without any special tooling. The browser’s own sampling profiler can show the same picture from a different angle, as in profiling allocation hot spots in a Wasm module. Keep sampling out of release builds: capturing stacks is expensive, and the imported function adds a boundary call on the allocation path.

Reducing what the counters reveal

The counters are only a means; the goal is fewer, cheaper allocations in the paths that matter. A handful of changes account for most reductions. Reuse buffers across calls: keep a Vec in a struct or a thread-local and clear() it instead of allocating a new one each time — capacity is retained, so steady-state calls allocate nothing. Pre-size collections with with_capacity when the final size is known or can be estimated, turning a dozen doubling reallocations into one. Borrow instead of cloning: functions that take &str or &[u8] rather than String or Vec<u8> avoid copies at every call site. Replace short-lived String formatting in hot loops with writes into a reused buffer. And at the boundary, accept &[u8] views from JavaScript instead of owned vectors where wasm-bindgen allows, which avoids an allocation for the copied-in data. After each change, the same measure() call shows the new count, so improvements are verified rather than assumed.

Expected output

measure(() => wasm.parse_document(bytes)) reports the allocation count, live delta and peak for one call; after fixing a leaked attribute map, the repetition test shows live bytes constant across rounds, and the allocation-budget test passes in CI.

Gotchas

  • Forgetting realloc. The default realloc calls alloc and dealloc, but overriding it without counting skews the numbers.
  • Counters overflowing u32 in JavaScript. Long runs exceed 4 billion allocations. Return f64 or BigInt if needed.
  • Lazy statics looking like leaks. First-call initialisation raises live bytes once. Warm up before measuring.
  • Allocator in release builds by accident. Gate it behind a feature.
  • Allocations from JavaScript glue. wasm-bindgen’s __wbindgen_malloc goes through the global allocator too — counts include boundary copies.

Performance note

The counting wrapper added about 2% to an allocation-heavy parsing benchmark in Chrome and nothing measurable to compute-bound code. Using it to find and remove 17,000 temporary allocations per parse cut parse time by 31%.

Parse time before and after acting on the counters Milliseconds to parse a 2 MB document with the counting wrapper disabled, enabled, and after removing temporary allocations found with it. ms per parse original build 48 ms with counting wrapper 49 ms after removing temporaries 33 ms

Frequently Asked Questions

Can I wrap a different allocator, such as talc or wee_alloc? Yes — forward to that allocator instead of System. The counting logic is independent of the allocator underneath.

Does this work for Emscripten C code? Not directly; for C, wrap malloc and free with --wrap linker options or use mallinfo().

Is Relaxed ordering safe? For counters that are only read for statistics, yes. Exact cross-thread consistency is not needed.

Can I count allocations per thread? Use thread-local counters alongside the global ones in threaded builds.

What about allocations that fail? Count failures separately if you need them; they return null and should not change live bytes.

← Back to Memory Profiling & Leak Detection