Replacing the Default Allocator to Save Bytes

This page answers one question: the memory allocator is one of the largest fixed costs in a small Rust WebAssembly module — when is it worth replacing, with what, and what do you give up?

Prerequisites

  • [ ] A Rust module built for wasm32-unknown-unknown with a release profile already tuned for size.
  • [ ] twiggy to confirm how much the allocator costs in your build.
  • [ ] A benchmark or test that exercises allocation the way your real workload does.

Why the allocator matters for small modules

Any Rust code that uses Vec, String, Box or a collection needs a global allocator, and on wasm32-unknown-unknown the standard library provides a port of dlmalloc. It is a good general-purpose allocator: it reuses freed memory efficiently, handles a wide range of allocation sizes, and is reasonably fast. It is also around 10 KB of code before compression, roughly 4 KB after.

For a module that is hundreds of kilobytes, that is noise. For a small module — a validator, a hash function, a parser for one format, a module embedded inline in a page — the allocator can be a fifth of the total. That is when replacing it is worth considering, and the trade-off is always the same: a smaller allocator is simpler, and simpler means it either reuses memory less well or does less work to find a good block.

Rust Wasm allocators compared Approximate code size, memory reuse behaviour, speed and suitability for dlmalloc, talc, lol_alloc's free-list allocator, a leaking bump allocator and a no-allocation design. allocator code size reuses freed memory suited to dlmalloc (default) ~10 KB yes, well general workloads talc ~3-4 KB yes small modules, long-lived lol_alloc FreeListAllocator ~1 KB yes, slower search tiny modules bump / leaking allocator < 0.5 KB never one-shot, short-lived calls no allocation (no_std) 0 n/a fixed buffers only

A note on wee_alloc, which older guides recommend: it is unmaintained, has known memory leaks with some allocation patterns, and should not be used in new code.

Step 1 — measure what the allocator costs you

wasm-bindgen target/wasm32-unknown-unknown/release/app.wasm --out-dir pkg --target web --keep-debug
twiggy top -n 40 pkg/app_bg.wasm | grep -iE 'dlmalloc|alloc'
          4108 ┊     9.12% ┊ dlmalloc::dlmalloc::Dlmalloc::malloc
          2215 ┊     4.92% ┊ dlmalloc::dlmalloc::Dlmalloc::free
          1382 ┊     3.07% ┊ dlmalloc::dlmalloc::Dlmalloc::dispose_chunk
           894 ┊     1.99% ┊ dlmalloc::dlmalloc::Dlmalloc::memalign

Add up the allocator’s entries. If they are a small fraction of the module — under 5% — stop here: the saving will not be noticed and the risks are not worth it. If they are 15% or more, keep going.

Step 2 — try a smaller general-purpose allocator

talc is a modern allocator with good memory reuse and a smaller footprint, and supports WebAssembly directly:

[dependencies]
talc = { version = "4", default-features = false, features = ["lock_api"] }
// src/lib.rs
#[cfg(target_arch = "wasm32")]
#[global_allocator]
static ALLOC: talc::TalckWasm = unsafe { talc::TalckWasm::new_global() };

TalckWasm grows the heap with memory.grow on demand, like the default. For the smallest possible size, lol_alloc offers a free-list allocator in well under a kilobyte, at the cost of a linear search for free blocks:

[dependencies]
lol_alloc = "0.4"
#[cfg(target_arch = "wasm32")]
#[global_allocator]
static ALLOC: lol_alloc::AssumeSingleThreaded<lol_alloc::FreeListAllocator> =
    unsafe { lol_alloc::AssumeSingleThreaded::new(lol_alloc::FreeListAllocator::new()) };

AssumeSingleThreaded is exactly what it says: it skips locking because a single-threaded module cannot race. It is unsound in a threaded build. If you use SharedArrayBuffer threads, stay with an allocator that locks.

Step 3 — consider a bump allocator for one-shot modules

Some modules are called once and thrown away: a function that parses an input and returns a result, run in a fresh instance per call. For those, an allocator that never frees is both the smallest and the fastest. It hands out memory by advancing a pointer and grows memory when it runs out. The design and its pitfalls are covered in implementing a bump allocator in Wasm.

The condition for using one is strict: the module’s lifetime must be short enough that never reusing memory cannot exhaust it. A long-lived instance with a bump allocator leaks every allocation, and will eventually fail when memory reaches its maximum. Instantiating per call, as described in instantiating one module many times, makes that safe.

Choosing an allocator for a Rust Wasm module A decision tree. If the allocator is a small share of the module, keep dlmalloc. If size matters and the module is long-lived, choose talc. If the module is tiny and allocations are few, a free-list allocator. If each instance handles one call and is discarded, a bump allocator. Is the allocator more than ~15% of the module? no Keep dlmalloc saving would not be noticed yes, long-lived instance talc small and reuses memory well yes, few allocations lol_alloc free list tiny, slower search yes, one call per instance Bump allocator never frees; re-instantiate

Step 4 — measure size, speed and peak memory together

An allocator swap changes three numbers, and all three matter:

// bench.mjs — run the same workload against each build
for (const variant of ["dlmalloc", "talc", "lol_alloc"]) {
  const bytes = await readFile(`builds/${variant}/app_bg.wasm`);
  const { instance } = await WebAssembly.instantiate(bytes, imports);
  const t0 = performance.now();
  for (let i = 0; i < 2000; i++) instance.exports.process_document(docPtr, docLen);
  const ms = (performance.now() - t0) / 2000;
  console.log(variant, {
    sizeKB: (bytes.length / 1024).toFixed(1),
    perCallMs: ms.toFixed(3),
    peakMB: (instance.exports.memory.buffer.byteLength / 2 ** 20).toFixed(1),
  });
}

Peak memory is the one people forget. An allocator that fragments badly grows linear memory further for the same workload, and since WebAssembly memory never shrinks — see why Wasm memory never shrinks — that growth is permanent for the life of the instance.

Per-call time for an allocation-heavy workload by allocator A document-processing function that makes thousands of small allocations, called 2,000 times, with three allocators. The smallest allocator is slowest because finding a free block takes a linear search. microseconds per call (lower is better) dlmalloc 412 µs talc 398 µs lol_alloc free list 1,270 µs Code size for the same three builds was 10.2 KB, 3.6 KB and 0.9 KB of allocator code respectively; peak memory was 18, 19 and 31 MB.

Step 5 — reduce allocation instead

The biggest win is often not a smaller allocator but fewer allocations. Reusing a buffer across calls, preallocating with Vec::with_capacity, using arena allocation for data with a shared lifetime, or passing slices into JavaScript-owned memory all reduce the allocator’s work — which makes even the small, slow allocators perform well, and makes the choice between them matter less. A module that does no dynamic allocation at all can drop the allocator entirely with #![no_std] and fixed-size buffers, which is how the smallest useful modules are built.

Thinking about allocation as a design decision

It helps to step back from the allocator itself and ask what the module’s memory looks like over its life. Most WebAssembly modules fall into one of three shapes, and each points to a different choice.

The first is the request handler: a call comes in, the module allocates working data, produces a result, and everything it allocated is garbage when the call returns. Parsers, validators, converters and most server-side functions look like this. For them, the cheapest correct design is often a bump or arena allocator reset after each call, or a fresh instance per call — no fragmentation, no free-list search, almost no allocator code.

The second is the long-lived engine: an editor, a game, a database that lives as long as the page and allocates and frees continuously. Fragmentation and peak memory matter here, because the instance never restarts and its memory never shrinks. A mature reusing allocator — the default, or talc — is the right choice, and the few kilobytes it costs are well spent.

The third is the fixed-buffer kernel: a hash function, a codec inner loop, a DSP routine that works on buffers the caller provides. It may not need dynamic allocation at all. Removing allocation entirely — no_std, buffers passed in from JavaScript, results written into caller-owned memory — removes the allocator from the module and is the smallest design of all.

Knowing which shape your module has usually makes the allocator decision obvious, and it often suggests a better optimization than swapping allocators: changing how memory is used.

Expected output

With talc swapped in for a small validation module:

before: app_bg.wasm 46,812 bytes (17,904 brotli)
after:  app_bg.wasm 40,196 bytes (15,322 brotli)
twiggy: no dlmalloc entries; talc::* totals 3,618 bytes

The module’s tests and benchmarks should pass with no measurable change in speed.

Gotchas

  • Random corruption after switching. The allocator assumed a single thread in a threaded build. Use a locking variant.
  • Memory grows without bound. A bump allocator in a long-lived instance. Re-instantiate per call or use a reusing allocator.
  • No size change at all. The #[global_allocator] was behind a cfg that did not match the target. Check with twiggy that dlmalloc symbols are gone.
  • Slower than expected. The workload makes many small allocations and the free-list search dominates. Measure with your real allocation pattern, not a synthetic one.

Performance note

On a 47 KB validation module, switching to talc saved 2.6 KB compressed — 14% of the transfer — with no measurable speed difference. On a 900 KB image editor, the same switch saved the same 2.6 KB, under 1% of the transfer, and was not worth the extra dependency. The allocator is a small-module optimization.

Frequently Asked Questions

Is dlmalloc slow? No. It is a solid general-purpose allocator; the reason to replace it is size, not speed.

Can I use different allocators for different builds? Yes. Gate the #[global_allocator] on a cargo feature, and pick the smaller one only for the size-critical build.

Does the allocator affect JavaScript code that allocates in Wasm memory? wasm-bindgen’s glue calls the module’s exported allocation functions, which use whatever global allocator the module has. Swapping the allocator changes that too, transparently.

What about Emscripten? Emscripten offers -sMALLOC=emmalloc (smaller) as an alternative to its default dlmalloc, with the same size-for-speed trade-off, and -sMALLOC=none for modules that never allocate.

← Back to Wasm Optimization Flags & Size Reduction