Measuring Allocator Fragmentation in Wasm
This page answers one task: a WebAssembly module’s memory keeps growing even though the amount of live data seems stable, and you need to know whether the allocator is fragmenting — holding free memory it cannot reuse — before deciding how to fix it.
Prerequisites
- [ ] A module built with Emscripten (dlmalloc or emmalloc) or Rust (the default allocator or a replacement such as talc).
- [ ] The ability to add an export that reports allocator statistics.
- [ ] A reproducible workload that shows memory growth over time.
Three numbers that tell the story
From the outside, the only memory number is the size of linear memory: memory.buffer.byteLength. That number only ever grows, so on its own it
cannot distinguish a leak from fragmentation from a legitimately large working set. Inside the module, the allocator knows more. Three figures are
enough to diagnose almost every case:
- Heap size — how much memory the allocator has claimed from linear memory.
- In use — the bytes currently allocated to live objects.
- Largest free block — the biggest single allocation that could succeed right now without growing memory.
The difference between heap size and in-use bytes is free memory. If free memory is large but the largest free block is small, the free memory is
scattered in pieces too small for the requests the program makes: that is fragmentation. A common single figure is the external fragmentation
ratio: 1 − largest_free / total_free. Near zero, free memory is mostly one usable block; near one, it is shattered.
Step 1 — export statistics from the allocator
With Emscripten’s default dlmalloc, mallinfo() reports the totals. Export a function that writes them to a small struct:
#include <malloc.h>
#include <emscripten.h>
typedef struct { size_t heap, in_use, free_total, largest_free; } HeapStats;
static HeapStats stats;
EMSCRIPTEN_KEEPALIVE HeapStats *heap_stats(void) {
struct mallinfo mi = mallinfo();
stats.heap = mi.arena; /* bytes obtained from memory */
stats.in_use = mi.uordblks; /* bytes in allocated blocks */
stats.free_total = mi.fordblks; /* bytes in free blocks */
stats.largest_free = largest_free_block();
return &stats;
}
dlmalloc’s mallinfo does not report the largest free block directly. Find it empirically: binary-search the largest malloc that succeeds without
growing memory, by comparing emscripten_get_heap_size() before and after, then free it. It is a diagnostic, not a hot-path call, so the cost is fine.
emmalloc offers emmalloc_compute_free_dynamic_memory_fragmentation_map() and emmalloc_dynamic_heap_size() for the same purpose.
Step 2 — do the same in Rust
Rust’s default wasm32 allocator exposes no statistics, so wrap it. A global allocator wrapper that counts live bytes gives in use; the memory size
gives an upper bound on heap:
use std::alloc::{GlobalAlloc, Layout, System};
use std::sync::atomic::{AtomicUsize, Ordering::Relaxed};
pub struct Stats;
static LIVE: AtomicUsize = AtomicUsize::new(0);
unsafe impl GlobalAlloc for Stats {
unsafe fn alloc(&self, l: Layout) -> *mut u8 {
let p = System.alloc(l);
if !p.is_null() { LIVE.fetch_add(l.size(), Relaxed); }
p
}
unsafe fn dealloc(&self, p: *mut u8, l: Layout) { System.dealloc(p, l); LIVE.fetch_sub(l.size(), Relaxed); }
}
#[global_allocator] static A: Stats = Stats;
#[wasm_bindgen]
pub fn heap_stats() -> Vec<u32> {
let mem = core::arch::wasm32::memory_size(0) * 65536;
vec![mem as u32, LIVE.load(Relaxed) as u32, largest_alloc_without_growth() as u32]
}
Allocators written for Wasm, such as talc, expose counters directly. Whatever the source, the goal is the same three numbers. The counting technique is expanded in counting allocations with a wrapping allocator.
Step 3 — sample over a realistic session
Fragmentation is a property of history: it develops as objects of different sizes and lifetimes are allocated and freed. A single snapshot after a unit test says little. Drive the real workload — open and close documents, run a game for ten minutes, process a stream of requests — and sample the statistics periodically:
const samples = [];
setInterval(() => {
const [heap, inUse, largest] = wasm.heap_stats();
const free = heap - inUse;
samples.push({ t: performance.now(), heap, inUse, free, frag: free ? 1 - largest / free : 0 });
}, 5000);
Plot heap and in-use bytes over time. Three shapes are common. If in-use grows steadily, it is a leak, not fragmentation — see finding Wasm memory leaks in the browser. If in-use is flat but heap keeps stepping up and the fragmentation ratio climbs, it is fragmentation. If both rise to a plateau and stay there, the module is simply using what it needs.
Step 4 — find the allocation pattern behind it
Fragmentation usually has an identifiable cause: a long-lived small object allocated between two short-lived large ones, pinning the space between them. Typical culprits are caches that allocate entries interleaved with temporary buffers, strings built incrementally by repeated reallocation, and vectors that grow by doubling while other allocations land behind them. Log allocation sizes and lifetimes for a short period — a histogram of sizes, and for each size class how long objects live — and look for small, long-lived objects allocated in the same phase as large, short-lived ones.
Step 5 — reduce it
The remedies follow from the cause. Separate lifetimes: put per-frame or per-request data in an arena, as in
using an arena allocator for per-frame data,
so it never interleaves with long-lived data. Pre-size growing containers with with_capacity or reserve so they do not reallocate repeatedly.
Allocate long-lived objects early, at startup, before temporary churn begins. Pool objects of a fixed size. And, if the pattern is inherent, try a
different allocator: size-class allocators fragment less for many small objects. Re-measure after each change, with the same workload, so the effect
is visible in the same three numbers.
Fragmentation versus overhead
Not all of the gap between heap size and live bytes is fragmentation. Every allocator spends memory on block headers, on rounding requests up to size classes or alignment, and on per-chunk bookkeeping. This internal overhead is typically 8–16 bytes per allocation, which becomes significant for programs with millions of tiny objects: a million 16-byte allocations can consume 32 MB rather than 16. It shows up as in-use bytes (as the allocator reports them) exceeding the sum of requested sizes. Measure it by comparing the allocator’s in-use figure with the wrapper’s count of requested bytes. High internal overhead is fixed differently from external fragmentation — by allocating fewer, larger objects (structure-of-arrays layouts, pooled nodes, small-string optimisation) rather than by separating lifetimes. Knowing which of the two you have saves trying the wrong fix.
Making the measurement part of CI
Once a fragmentation problem has been found and fixed, keep it fixed. A scripted session — the same one used for diagnosis, shortened to a minute or two of simulated use — can run headlessly in CI with Playwright or in Node, sampling the three statistics at the end. Assert two things: that the heap size after the session stays under a budget, and that the fragmentation ratio stays below a threshold such as 0.5. Regressions then show up as failed checks on the change that caused them, rather than as user reports of a sluggish tab weeks later. Record the numbers as build artefacts as well, so trends are visible even while the checks still pass.
Expected output
A ten-minute session of the document editor shows in-use bytes flat at about 38 MB while the heap grows from 64 MB to 142 MB, with the fragmentation ratio rising from 0.1 to 0.83 — fragmentation confirmed. After moving undo-snapshot scratch buffers to an arena, the heap plateaus at 71 MB with the ratio below 0.2.
Gotchas
- Reading only
memory.buffer.byteLength. It cannot distinguish leaks from fragmentation. Export allocator statistics. - Measuring after a short test. Fragmentation develops over time. Measure a realistic, long session.
- Probing the largest block in a hot path. The binary search allocates. Call it from diagnostics only.
- Counting requested bytes as in-use. Allocator overhead is extra. Compare both figures to see it.
Performance note
The statistics export cost about 0.2 ms per sample with the largest-block probe, and microseconds without it. Sampling every 5 seconds added nothing measurable to the workload. The fix found through it — an arena for scratch buffers — reduced peak memory by half.
Frequently Asked Questions
Can I defragment the heap? Not in general — pointers to live objects cannot be moved without the program’s cooperation. Compaction requires handles instead of raw pointers.
Does a fresh instance remove fragmentation? Yes. Re-instantiating, or reloading, starts with an empty heap; long-lived workers sometimes recycle periodically for this reason.
Is fragmentation worse in Wasm than natively? The allocators are the same, but Wasm cannot return pages to the OS, so fragmented memory stays allocated; see why Wasm memory never shrinks.
Which allocator fragments least? It depends on the workload. Measure two or three with your real session rather than relying on general benchmarks.
Related
- Implementing a free-list allocator in Wasm — how free lists fragment.
- Tracking linear memory growth over time — the outside view.
- Monitoring Wasm memory in production — sampling the same numbers from users.
- Profiling allocation hot spots in a Wasm module — finding where allocations come from.
← Back to Linear Memory Management & Allocators