Benchmarking Allocation-Heavy Wasm Code
This page answers one task: a WebAssembly function is slower than expected, and it allocates a lot — strings, vectors, tree nodes, temporary buffers — so you want measurements that show how much time goes to allocation versus the algorithm, before deciding whether to change the allocator, the data structures, or neither.
Prerequisites
- [ ] A benchmark harness for the function (Criterion natively, or a browser or Node harness for the built module).
- [ ] The ability to rebuild with a different global allocator (Rust) or
mallocimplementation (C/C++). - [ ] A representative input set.
Why allocation dominates more often in Wasm
WebAssembly has no built-in heap allocator: every module ships its own, compiled into the module — dlmalloc by default in Rust’s
wasm32-unknown-unknown, dlmalloc or emmalloc in Emscripten, or a size-optimised allocator chosen to save bytes. Size-optimised allocators are often
much slower than the allocators native programs get from the operating system, and none have thread-local caches on single-threaded Wasm. On top of
that, when the heap runs out, the allocator calls memory.grow, which can be expensive — the engine may need to reserve or commit memory, and in some
configurations copy it.
So code that allocates freely, and runs well natively, may spend a much larger share of its time in the allocator once compiled to Wasm. Measuring that share is the first step; the fix might be a faster allocator, fewer allocations, or pre-sizing memory, and each has different costs.
Step 1 — count allocations
Before timing, count: how many allocations per call, of which sizes. In Rust, wrap the global allocator with a counting allocator in a benchmark build:
use std::alloc::{GlobalAlloc, Layout, System};
use std::sync::atomic::{AtomicUsize, Ordering::Relaxed};
pub struct Counting;
pub static ALLOCS: AtomicUsize = AtomicUsize::new(0);
pub static BYTES: AtomicUsize = AtomicUsize::new(0);
unsafe impl GlobalAlloc for Counting {
unsafe fn alloc(&self, l: Layout) -> *mut u8 { ALLOCS.fetch_add(1, Relaxed); BYTES.fetch_add(l.size(), Relaxed); System.alloc(l) }
unsafe fn dealloc(&self, p: *mut u8, l: Layout) { System.dealloc(p, l) }
}
#[cfg(feature = "count-allocs")]
#[global_allocator]
static A: Counting = Counting;
Export a function that returns and resets the counters, and log them per benchmark iteration. Hundreds of thousands of small allocations per call is a strong signal; a handful of large ones usually is not. The counting technique is covered in more depth in counting allocations with a wrapping allocator.
Step 2 — measure the allocator’s share by swapping it
The most direct way to measure allocator cost is to change the allocator and see what moves. Build the same code with two or three allocators and benchmark each:
[features]
alloc-talc = ["dep:talc"]
alloc-lol = ["dep:lol_alloc"]
#[cfg(feature = "alloc-talc")]
#[global_allocator]
static ALLOC: talc::TalckWasm = unsafe { talc::TalckWasm::new_global() };
If switching from a size-optimised allocator to a faster one cuts run time by 30%, allocation was at least 30% of the cost. If nothing changes, the
allocator is not the problem. For Emscripten, compare -sMALLOC=dlmalloc with -sMALLOC=emmalloc and mimalloc where available. Record module size
too, since faster allocators are larger.
Step 3 — remove memory.grow from steady-state measurements
The first iterations of a benchmark grow memory from its initial size to its working size; later iterations reuse it. If you measure from a fresh
instance, growth costs land in the measurement. Decide what you want to measure: steady state (warm up until memory stops growing, then measure) or cold
behaviour (measure from a fresh instance and report growth separately). Track memory.buffer.byteLength before and after each iteration — if it is still
growing during measured runs, the numbers include growth:
const mem = instance.exports.memory;
for (let i = 0; i < 50; i++) run(); // warm-up until growth stops
const before = mem.buffer.byteLength;
const t = measure(run, 30);
if (mem.buffer.byteLength !== before) console.warn("memory grew during measurement");
Setting a larger initial memory at link time (--initial-memory or Emscripten’s INITIAL_MEMORY) removes growth for workloads with a known working set,
and is a legitimate optimisation in its own right.
Step 4 — try structural fixes and compare
With the allocator’s share known, test structural changes against the same benchmark: reuse buffers across calls instead of allocating per call;
pre-size vectors with with_capacity; replace many small heap objects with an arena (bumpalo in Rust) that is reset after each call; store tree nodes in
a Vec and refer to them by index. Arenas are particularly effective in Wasm, because a bump allocation is a pointer increment and freeing the whole
arena is one reset, avoiding per-object bookkeeping entirely.
use bumpalo::Bump;
pub fn parse_document(src: &str, arena: &mut Bump) -> usize {
arena.reset(); // reuse memory from the previous call
let nodes = parse_into(src, arena); // all nodes allocated in the arena
count_words(&nodes)
}
Step 5 — check long-running behaviour
Short benchmarks miss fragmentation. A function that allocates and frees varied sizes can fragment the heap over thousands of calls, so later calls search longer free lists or trigger growth. Run a long benchmark — tens of thousands of iterations with realistic input variety — and plot time per call and memory size over the run. A rising trend in either indicates fragmentation or a leak, both of which short benchmarks hide.
Comparing with native numbers
The native build of the same code usually allocates faster, because the system allocator is highly optimised and multithreaded. Comparing native and Wasm with allocation counts in hand tells you whether the Wasm slowdown is mostly allocation: if the ratio of Wasm to native time is much worse for the allocation-heavy benchmark than for a compute-only one, the allocator is the difference. This also guards against wasted effort — if Wasm is only 1.2× slower than native on the allocation-heavy path, a new allocator will not change much.
Reporting results that lead to the right fix
Report, per benchmark: allocations and bytes per call; time with each allocator; time with structural fixes; module size per allocator; and memory high-water mark. That set makes trade-offs visible — a faster allocator costing 15 KB, an arena that removes 90% of allocations but needs an API change, a larger initial memory that removes growth at the cost of reserving memory up front.
Allocation at the JavaScript boundary
Some allocation cost is not in your algorithm at all but in the glue. Passing a string or array into a wasm-bindgen function allocates a buffer in linear
memory for it, copies the data, and frees it afterwards; returning a Vec or String allocates on the Rust side and frees after JavaScript copies it out.
For a function called thousands of times per frame with small arguments, those boundary allocations can dominate. Count them separately by benchmarking
an empty function with the same signature — the time it takes is pure boundary cost — and compare with the real function. If the boundary is the
problem, the fixes are different from allocator tuning: pass larger batches per call, reuse a buffer that JavaScript writes into directly through a
typed-array view, or keep data inside Wasm between calls so it crosses the boundary once. The techniques are covered in
zero-copy data transfer patterns.
Making allocation benchmarks part of the suite
Once allocation has been identified as a cost, protect the improvement. Add the allocation count per call as a tracked metric next to time — counts are deterministic, so a budget of “at most 50 allocations per parse” can be enforced tightly in CI without noise — and keep the long-running fragmentation benchmark in the nightly job. A future change that reintroduces per-node allocation then fails the count budget immediately, long before it shows up as a timing regression.
Expected output
The parser benchmark shows 182,000 allocations per call averaging 38 bytes; switching from a size-optimised allocator to talc cut time per call from 31 ms
to 22 ms at a cost of 6 KB; an arena cut it to 12 ms with 40 allocations per call; and a 20,000-iteration run shows stable time and memory with the arena.
Gotchas
- Measuring growth as allocator cost. Warm up until memory is stable.
- Choosing a tiny allocator for a hot path. Size-optimised allocators can be slow. Measure both.
- Short runs only. Fragmentation appears over thousands of calls.
- Counting allocators in timing runs. Counting adds overhead. Count and time in separate builds.
- Ignoring size. Faster allocators add bytes. Report both numbers.
Performance note
Across the three fixes, time per parse fell from 31 ms to 12 ms. Most of the gain came from removing allocations rather than speeding them up.
Frequently Asked Questions
Is dlmalloc slow?
It is a reasonable general-purpose allocator; size-optimised alternatives are usually slower.
Can I profile allocation in DevTools? Wasm allocations are invisible to the JavaScript heap profiler; use counting allocators and CPU profiles of allocator functions.
Does Emscripten’s mimalloc help single-threaded code?
Sometimes; its main advantage is multithreaded scaling. Measure.
Should benchmarks preallocate memory? Only if production does; otherwise measure growth separately.
How do I measure the boundary’s share of the cost? Benchmark an empty function with the same signature; its time is pure argument and result marshalling.
Related
- Replacing the default allocator to save bytes — the size side.
- Profiling allocation hot spots in a Wasm module — finding where allocations come from.
- Building a reproducible Wasm benchmark harness — the harness.
- Benchmarking memory bandwidth in Wasm — memory-bound code.
← Back to Wasm Performance Benchmarking