Replacing the Default Allocator to Save Bytes
This page answers one question: the memory allocator is one of the largest fixed costs in a small Rust WebAssembly module — when is it worth replacing, with what, and what do you give up?
Prerequisites
- [ ] A Rust module built for
wasm32-unknown-unknownwith a release profile already tuned for size. - [ ]
twiggyto confirm how much the allocator costs in your build. - [ ] A benchmark or test that exercises allocation the way your real workload does.
Why the allocator matters for small modules
Any Rust code that uses Vec, String, Box or a collection needs a global allocator, and on wasm32-unknown-unknown
the standard library provides a port of dlmalloc. It is a good general-purpose allocator: it reuses freed memory
efficiently, handles a wide range of allocation sizes, and is reasonably fast. It is also around 10 KB of code before
compression, roughly 4 KB after.
For a module that is hundreds of kilobytes, that is noise. For a small module — a validator, a hash function, a parser for one format, a module embedded inline in a page — the allocator can be a fifth of the total. That is when replacing it is worth considering, and the trade-off is always the same: a smaller allocator is simpler, and simpler means it either reuses memory less well or does less work to find a good block.
A note on wee_alloc, which older guides recommend: it is unmaintained, has known memory leaks with some allocation
patterns, and should not be used in new code.
Step 1 — measure what the allocator costs you
wasm-bindgen target/wasm32-unknown-unknown/release/app.wasm --out-dir pkg --target web --keep-debug
twiggy top -n 40 pkg/app_bg.wasm | grep -iE 'dlmalloc|alloc'
4108 ┊ 9.12% ┊ dlmalloc::dlmalloc::Dlmalloc::malloc
2215 ┊ 4.92% ┊ dlmalloc::dlmalloc::Dlmalloc::free
1382 ┊ 3.07% ┊ dlmalloc::dlmalloc::Dlmalloc::dispose_chunk
894 ┊ 1.99% ┊ dlmalloc::dlmalloc::Dlmalloc::memalign
Add up the allocator’s entries. If they are a small fraction of the module — under 5% — stop here: the saving will not be noticed and the risks are not worth it. If they are 15% or more, keep going.
Step 2 — try a smaller general-purpose allocator
talc is a modern allocator with good memory reuse and a smaller footprint, and supports WebAssembly directly:
[dependencies]
talc = { version = "4", default-features = false, features = ["lock_api"] }
// src/lib.rs
#[cfg(target_arch = "wasm32")]
#[global_allocator]
static ALLOC: talc::TalckWasm = unsafe { talc::TalckWasm::new_global() };
TalckWasm grows the heap with memory.grow on demand, like the default. For the smallest possible size, lol_alloc
offers a free-list allocator in well under a kilobyte, at the cost of a linear search for free blocks:
[dependencies]
lol_alloc = "0.4"
#[cfg(target_arch = "wasm32")]
#[global_allocator]
static ALLOC: lol_alloc::AssumeSingleThreaded<lol_alloc::FreeListAllocator> =
unsafe { lol_alloc::AssumeSingleThreaded::new(lol_alloc::FreeListAllocator::new()) };
AssumeSingleThreaded is exactly what it says: it skips locking because a single-threaded module cannot race. It is
unsound in a threaded build. If you use SharedArrayBuffer threads,
stay with an allocator that locks.
Step 3 — consider a bump allocator for one-shot modules
Some modules are called once and thrown away: a function that parses an input and returns a result, run in a fresh instance per call. For those, an allocator that never frees is both the smallest and the fastest. It hands out memory by advancing a pointer and grows memory when it runs out. The design and its pitfalls are covered in implementing a bump allocator in Wasm.
The condition for using one is strict: the module’s lifetime must be short enough that never reusing memory cannot exhaust it. A long-lived instance with a bump allocator leaks every allocation, and will eventually fail when memory reaches its maximum. Instantiating per call, as described in instantiating one module many times, makes that safe.
Step 4 — measure size, speed and peak memory together
An allocator swap changes three numbers, and all three matter:
// bench.mjs — run the same workload against each build
for (const variant of ["dlmalloc", "talc", "lol_alloc"]) {
const bytes = await readFile(`builds/${variant}/app_bg.wasm`);
const { instance } = await WebAssembly.instantiate(bytes, imports);
const t0 = performance.now();
for (let i = 0; i < 2000; i++) instance.exports.process_document(docPtr, docLen);
const ms = (performance.now() - t0) / 2000;
console.log(variant, {
sizeKB: (bytes.length / 1024).toFixed(1),
perCallMs: ms.toFixed(3),
peakMB: (instance.exports.memory.buffer.byteLength / 2 ** 20).toFixed(1),
});
}
Peak memory is the one people forget. An allocator that fragments badly grows linear memory further for the same workload, and since WebAssembly memory never shrinks — see why Wasm memory never shrinks — that growth is permanent for the life of the instance.
Step 5 — reduce allocation instead
The biggest win is often not a smaller allocator but fewer allocations. Reusing a buffer across calls, preallocating with
Vec::with_capacity, using arena allocation for data with a shared lifetime, or passing slices into JavaScript-owned memory
all reduce the allocator’s work — which makes even the small, slow allocators perform well, and makes the choice between
them matter less. A module that does no dynamic allocation at all can drop the allocator entirely with #![no_std] and
fixed-size buffers, which is how the smallest useful modules are built.
Thinking about allocation as a design decision
It helps to step back from the allocator itself and ask what the module’s memory looks like over its life. Most WebAssembly modules fall into one of three shapes, and each points to a different choice.
The first is the request handler: a call comes in, the module allocates working data, produces a result, and everything it allocated is garbage when the call returns. Parsers, validators, converters and most server-side functions look like this. For them, the cheapest correct design is often a bump or arena allocator reset after each call, or a fresh instance per call — no fragmentation, no free-list search, almost no allocator code.
The second is the long-lived engine: an editor, a game, a database that lives as long as the page and allocates and
frees continuously. Fragmentation and peak memory matter here, because the instance never restarts and its memory never
shrinks. A mature reusing allocator — the default, or talc — is the right choice, and the few kilobytes it costs are well
spent.
The third is the fixed-buffer kernel: a hash function, a codec inner loop, a DSP routine that works on buffers the caller
provides. It may not need dynamic allocation at all. Removing allocation entirely — no_std, buffers passed in from
JavaScript, results written into caller-owned memory — removes the allocator from the module and is the smallest design
of all.
Knowing which shape your module has usually makes the allocator decision obvious, and it often suggests a better optimization than swapping allocators: changing how memory is used.
Expected output
With talc swapped in for a small validation module:
before: app_bg.wasm 46,812 bytes (17,904 brotli)
after: app_bg.wasm 40,196 bytes (15,322 brotli)
twiggy: no dlmalloc entries; talc::* totals 3,618 bytes
The module’s tests and benchmarks should pass with no measurable change in speed.
Gotchas
- Random corruption after switching. The allocator assumed a single thread in a threaded build. Use a locking variant.
- Memory grows without bound. A bump allocator in a long-lived instance. Re-instantiate per call or use a reusing allocator.
- No size change at all. The
#[global_allocator]was behind acfgthat did not match the target. Check with twiggy thatdlmallocsymbols are gone. - Slower than expected. The workload makes many small allocations and the free-list search dominates. Measure with your real allocation pattern, not a synthetic one.
Performance note
On a 47 KB validation module, switching to talc saved 2.6 KB compressed — 14% of the transfer — with no measurable speed
difference. On a 900 KB image editor, the same switch saved the same 2.6 KB, under 1% of the transfer, and was not worth the
extra dependency. The allocator is a small-module optimization.
Frequently Asked Questions
Is dlmalloc slow? No. It is a solid general-purpose allocator; the reason to replace it is size, not speed.
Can I use different allocators for different builds?
Yes. Gate the #[global_allocator] on a cargo feature, and pick the smaller one only for the size-critical build.
Does the allocator affect JavaScript code that allocates in Wasm memory? wasm-bindgen’s glue calls the module’s exported allocation functions, which use whatever global allocator the module has. Swapping the allocator changes that too, transparently.
What about Emscripten?
Emscripten offers -sMALLOC=emmalloc (smaller) as an alternative to its default dlmalloc, with the same size-for-speed
trade-off, and -sMALLOC=none for modules that never allocate.
Related
- Removing panic and formatting bloat from Rust Wasm — the other large fixed cost.
- Implementing a free-list allocator in Wasm — how a small reusing allocator works.
- Measuring allocator fragmentation in Wasm — the peak-memory side of the trade.
- Analyzing Wasm size with twiggy — confirming the saving.
← Back to Wasm Optimization Flags & Size Reduction