Benchmarking Allocation-Heavy Wasm Code

This page answers one task: a WebAssembly function is slower than expected, and it allocates a lot — strings, vectors, tree nodes, temporary buffers — so you want measurements that show how much time goes to allocation versus the algorithm, before deciding whether to change the allocator, the data structures, or neither.

Prerequisites

  • [ ] A benchmark harness for the function (Criterion natively, or a browser or Node harness for the built module).
  • [ ] The ability to rebuild with a different global allocator (Rust) or malloc implementation (C/C++).
  • [ ] A representative input set.

Why allocation dominates more often in Wasm

WebAssembly has no built-in heap allocator: every module ships its own, compiled into the module — dlmalloc by default in Rust’s wasm32-unknown-unknown, dlmalloc or emmalloc in Emscripten, or a size-optimised allocator chosen to save bytes. Size-optimised allocators are often much slower than the allocators native programs get from the operating system, and none have thread-local caches on single-threaded Wasm. On top of that, when the heap runs out, the allocator calls memory.grow, which can be expensive — the engine may need to reserve or commit memory, and in some configurations copy it.

So code that allocates freely, and runs well natively, may spend a much larger share of its time in the allocator once compiled to Wasm. Measuring that share is the first step; the fix might be a faster allocator, fewer allocations, or pre-sizing memory, and each has different costs.

Where time goes in an allocation-heavy function Total time splits into the algorithm's own work, allocation and deallocation calls, memory.grow calls when the heap expands, and cache misses from scattered allocations. Measurements should isolate each part so the right one is fixed. algorithm work the code you meant to run malloc / free calls allocator bookkeeping memory.grow heap expansion during run cache misses scattered small objects glue copies data in and out

Step 1 — count allocations

Before timing, count: how many allocations per call, of which sizes. In Rust, wrap the global allocator with a counting allocator in a benchmark build:

use std::alloc::{GlobalAlloc, Layout, System};
use std::sync::atomic::{AtomicUsize, Ordering::Relaxed};

pub struct Counting;
pub static ALLOCS: AtomicUsize = AtomicUsize::new(0);
pub static BYTES: AtomicUsize = AtomicUsize::new(0);

unsafe impl GlobalAlloc for Counting {
    unsafe fn alloc(&self, l: Layout) -> *mut u8 { ALLOCS.fetch_add(1, Relaxed); BYTES.fetch_add(l.size(), Relaxed); System.alloc(l) }
    unsafe fn dealloc(&self, p: *mut u8, l: Layout) { System.dealloc(p, l) }
}
#[cfg(feature = "count-allocs")]
#[global_allocator]
static A: Counting = Counting;

Export a function that returns and resets the counters, and log them per benchmark iteration. Hundreds of thousands of small allocations per call is a strong signal; a handful of large ones usually is not. The counting technique is covered in more depth in counting allocations with a wrapping allocator.

Step 2 — measure the allocator’s share by swapping it

The most direct way to measure allocator cost is to change the allocator and see what moves. Build the same code with two or three allocators and benchmark each:

[features]
alloc-talc = ["dep:talc"]
alloc-lol = ["dep:lol_alloc"]
#[cfg(feature = "alloc-talc")]
#[global_allocator]
static ALLOC: talc::TalckWasm = unsafe { talc::TalckWasm::new_global() };

If switching from a size-optimised allocator to a faster one cuts run time by 30%, allocation was at least 30% of the cost. If nothing changes, the allocator is not the problem. For Emscripten, compare -sMALLOC=dlmalloc with -sMALLOC=emmalloc and mimalloc where available. Record module size too, since faster allocators are larger.

Step 3 — remove memory.grow from steady-state measurements

The first iterations of a benchmark grow memory from its initial size to its working size; later iterations reuse it. If you measure from a fresh instance, growth costs land in the measurement. Decide what you want to measure: steady state (warm up until memory stops growing, then measure) or cold behaviour (measure from a fresh instance and report growth separately). Track memory.buffer.byteLength before and after each iteration — if it is still growing during measured runs, the numbers include growth:

const mem = instance.exports.memory;
for (let i = 0; i < 50; i++) run();                 // warm-up until growth stops
const before = mem.buffer.byteLength;
const t = measure(run, 30);
if (mem.buffer.byteLength !== before) console.warn("memory grew during measurement");

Setting a larger initial memory at link time (--initial-memory or Emscripten’s INITIAL_MEMORY) removes growth for workloads with a known working set, and is a legitimate optimisation in its own right.

Measuring with and without heap growth Measuring from a fresh instance mixes memory.grow calls into the timing and inflates early iterations. Warming up until memory stops growing measures steady-state allocator and algorithm cost, while growth is reported separately as a startup cost. fresh instance each run growth inside measured time early runs much slower mixes two costs cold behaviour warmed-up instance memory stable during runs allocator + algorithm only growth reported separately steady state

Step 4 — try structural fixes and compare

With the allocator’s share known, test structural changes against the same benchmark: reuse buffers across calls instead of allocating per call; pre-size vectors with with_capacity; replace many small heap objects with an arena (bumpalo in Rust) that is reset after each call; store tree nodes in a Vec and refer to them by index. Arenas are particularly effective in Wasm, because a bump allocation is a pointer increment and freeing the whole arena is one reset, avoiding per-object bookkeeping entirely.

use bumpalo::Bump;

pub fn parse_document(src: &str, arena: &mut Bump) -> usize {
    arena.reset();                                   // reuse memory from the previous call
    let nodes = parse_into(src, arena);              // all nodes allocated in the arena
    count_words(&nodes)
}

Step 5 — check long-running behaviour

Short benchmarks miss fragmentation. A function that allocates and frees varied sizes can fragment the heap over thousands of calls, so later calls search longer free lists or trigger growth. Run a long benchmark — tens of thousands of iterations with realistic input variety — and plot time per call and memory size over the run. A rising trend in either indicates fragmentation or a leak, both of which short benchmarks hide.

Comparing with native numbers

The native build of the same code usually allocates faster, because the system allocator is highly optimised and multithreaded. Comparing native and Wasm with allocation counts in hand tells you whether the Wasm slowdown is mostly allocation: if the ratio of Wasm to native time is much worse for the allocation-heavy benchmark than for a compute-only one, the allocator is the difference. This also guards against wasted effort — if Wasm is only 1.2× slower than native on the allocation-heavy path, a new allocator will not change much.

Reporting results that lead to the right fix

Report, per benchmark: allocations and bytes per call; time with each allocator; time with structural fixes; module size per allocator; and memory high-water mark. That set makes trade-offs visible — a faster allocator costing 15 KB, an arena that removes 90% of allocations but needs an API change, a larger initial memory that removes growth at the cost of reserving memory up front.

Allocation at the JavaScript boundary

Some allocation cost is not in your algorithm at all but in the glue. Passing a string or array into a wasm-bindgen function allocates a buffer in linear memory for it, copies the data, and frees it afterwards; returning a Vec or String allocates on the Rust side and frees after JavaScript copies it out. For a function called thousands of times per frame with small arguments, those boundary allocations can dominate. Count them separately by benchmarking an empty function with the same signature — the time it takes is pure boundary cost — and compare with the real function. If the boundary is the problem, the fixes are different from allocator tuning: pass larger batches per call, reuse a buffer that JavaScript writes into directly through a typed-array view, or keep data inside Wasm between calls so it crosses the boundary once. The techniques are covered in zero-copy data transfer patterns.

Making allocation benchmarks part of the suite

Once allocation has been identified as a cost, protect the improvement. Add the allocation count per call as a tracked metric next to time — counts are deterministic, so a budget of “at most 50 allocations per parse” can be enforced tightly in CI without noise — and keep the long-running fragmentation benchmark in the nightly job. A future change that reintroduces per-node allocation then fails the count budget immediately, long before it shows up as a timing regression.

Expected output

The parser benchmark shows 182,000 allocations per call averaging 38 bytes; switching from a size-optimised allocator to talc cut time per call from 31 ms to 22 ms at a cost of 6 KB; an arena cut it to 12 ms with 40 allocations per call; and a 20,000-iteration run shows stable time and memory with the arena.

Gotchas

  • Measuring growth as allocator cost. Warm up until memory is stable.
  • Choosing a tiny allocator for a hot path. Size-optimised allocators can be slow. Measure both.
  • Short runs only. Fragmentation appears over thousands of calls.
  • Counting allocators in timing runs. Counting adds overhead. Count and time in separate builds.
  • Ignoring size. Faster allocators add bytes. Report both numbers.

Performance note

Across the three fixes, time per parse fell from 31 ms to 12 ms. Most of the gain came from removing allocations rather than speeding them up.

Time per parse with each change Milliseconds per call for an allocation-heavy parser with a size-optimised allocator, with a faster allocator, and with an arena reused across calls. ms per call (median) size-optimised allocator 31 ms faster allocator 22 ms arena reused per call 12 ms

Frequently Asked Questions

Is dlmalloc slow? It is a reasonable general-purpose allocator; size-optimised alternatives are usually slower.

Can I profile allocation in DevTools? Wasm allocations are invisible to the JavaScript heap profiler; use counting allocators and CPU profiles of allocator functions.

Does Emscripten’s mimalloc help single-threaded code? Sometimes; its main advantage is multithreaded scaling. Measure.

Should benchmarks preallocate memory? Only if production does; otherwise measure growth separately.

How do I measure the boundary’s share of the cost? Benchmark an empty function with the same signature; its time is pure argument and result marshalling.

← Back to Wasm Performance Benchmarking