Choosing an Allocator for Rust Wasm
This page answers one task: a Rust WebAssembly module uses the default allocator, and you have heard that switching allocators can save size or time — so you want to know what the options are, what each trades off, and how to decide for your module rather than by reputation.
Prerequisites
- [ ] A Rust crate built for
wasm32-unknown-unknown(or WASI). - [ ] A benchmark that exercises realistic allocation patterns.
- [ ] A size measurement for the release module (compressed).
What the allocator does in a Wasm module
WebAssembly provides linear memory and memory.grow; it does not provide malloc. Every Rust module that uses Box, Vec, String or HashMap
links a global allocator that manages the heap inside linear memory: it hands out blocks, takes them back, and asks for more pages when the heap is
exhausted. Because the allocator is compiled into the module, its code size counts against your download, its speed counts in every allocation, and its
fragmentation behaviour decides how much memory a long-running module ends up holding — memory that, in WebAssembly, is never returned to the system.
Rust’s standard library uses a port of dlmalloc on wasm32-unknown-unknown: a mature general-purpose allocator, reasonably fast, around 10 KB of code.
Alternatives make different trade-offs. The right choice depends on how your module allocates: a handful of large buffers, millions of small objects,
short-lived per-call allocations, or long-lived data that grows and shrinks for hours.
Step 1 — measure the baseline
Before switching, measure what the allocator costs now: compressed size with and without allocation-heavy code is hard to isolate, so use twiggy to
see the allocator’s functions (dlmalloc:: symbols) and their retained size, and benchmark your real workload. Record allocation counts per operation
with a counting wrapper (see
counting allocations with a wrapping allocator)
so you know whether the module is allocation-heavy at all. A module that allocates a few buffers per call gains nothing measurable from a faster
allocator.
Step 2 — try talc
talc is a modern allocator with a Wasm-specific configuration that is both small and fast, and it is a good first alternative:
[dependencies]
talc = { version = "4", default-features = false, features = ["lock_api"] }
#[global_allocator]
static ALLOC: talc::TalckWasm = unsafe { talc::TalckWasm::new_global() };
TalckWasm claims memory with memory.grow as needed and is intended for single-threaded Wasm. Rebuild, measure size and run the benchmark. In many
modules it is both smaller and faster than the default; confirm on yours.
Step 3 — consider tiny allocators only for tiny modules
For modules where every byte counts and allocation is minimal — a small codec, a hashing function — lol_alloc offers minimal allocators, including a
bump allocator that never frees and a simple free-list allocator:
use lol_alloc::{AssumeSingleThreaded, FreeListAllocator};
#[global_allocator]
static ALLOC: AssumeSingleThreaded<FreeListAllocator> =
unsafe { AssumeSingleThreaded::new(FreeListAllocator::new()) };
A never-freeing allocator is only safe for modules instantiated per task and discarded, where total allocation is bounded. Simple free lists reuse memory
poorly under mixed sizes and can fragment badly in long-running modules. Avoid wee_alloc: it is unmaintained and has known leaks; older guides that
recommend it predate better options.
Step 4 — test fragmentation over long runs
Size and speed are measured in minutes; fragmentation shows up over hours. Run a long benchmark that replays realistic sequences of allocation and
deallocation — open and close documents, process many varied files — and track memory.buffer.byteLength and live bytes over time. An allocator that
looks fine in short benchmarks may grow memory steadily under mixed sizes. Because Wasm memory never shrinks, growth here is permanent until the
instance is recreated. See
measuring allocator fragmentation in Wasm.
Step 5 — decide and document
Pick the allocator that meets your size budget with acceptable speed and stable long-run memory. Typical outcomes: keep dlmalloc for large applications
where its maturity matters and 10 KB is irrelevant; switch to talc when size and speed both matter; use a tiny allocator for small, short-lived modules.
Record the measurement in a comment next to the #[global_allocator] declaration, with the date and versions, so the next person knows why.
Threads change the picture
With the atomics target feature and shared memory, the allocator must be thread-safe. The default dlmalloc uses a global lock on threaded Wasm builds,
which serialises allocation across workers and can become a bottleneck for allocation-heavy parallel code. Allocators with per-thread caches reduce
contention; Emscripten offers mimalloc for C and C++ threaded builds for this reason. In Rust, a common pattern is to reduce allocation in parallel
sections — per-thread arenas or preallocated buffers — rather than to rely on the allocator to scale.
Allocation patterns matter more than allocators
Swapping allocators rarely changes performance as much as changing how code allocates. Reusing buffers across calls, preallocating with
with_capacity, using arenas for per-call temporaries, and storing many small objects in a Vec by index typically save more time than any allocator
switch, and they help with every allocator. Treat the allocator choice as tuning after the allocation pattern is reasonable, not as a substitute for it.
How allocators claim memory from the engine
Allocators differ in how they ask the engine for memory, and that affects both startup and peak memory. When the heap runs out, an allocator calls
memory.grow for more pages; some grow by exactly what the current request needs, others by a larger chunk to amortise the cost of growing. Growing in
large steps reduces the number of memory.grow calls — which can be relatively expensive, especially when the engine must commit new pages — but raises
peak memory slightly. Some allocators also start by claiming everything between __heap_base and the initial memory size, which is free at
instantiation, before growing. If your module’s memory profile matters — for instance on mobile, where every megabyte of committed memory counts — look
at the allocator’s growth policy, and consider setting a larger initial memory for workloads with a known working set, so the allocator never needs to
grow during normal operation.
Interaction with the JavaScript side
Allocations do not only come from Rust code. wasm-bindgen’s glue allocates in linear memory for every string or slice passed into the module, through
exported functions that call the global allocator; Emscripten’s glue does the same with malloc. An application that passes many small strings across
the boundary therefore stresses the allocator with short-lived allocations of varied sizes, which is exactly the pattern where allocator quality shows.
When comparing allocators, include realistic boundary traffic in the benchmark — not only the Rust-internal workload — or the measurement will miss a
large part of what the allocator does in production.
Expected output
A table for your module comparing dlmalloc, talc and one tiny allocator on compressed size, a realistic benchmark and memory after a four-hour replay;
a decision (for example, talc: 7 KB smaller, 12% faster, flat memory) recorded beside the global allocator; and a CI job that reruns the long replay
weekly.
Gotchas
- Choosing by reputation. Results depend on allocation patterns. Measure your workload.
- Using
wee_alloc. Unmaintained and leaky. Pick a maintained allocator. - Never-freeing allocators in long-lived modules. Memory grows without bound.
- Short benchmarks only. Fragmentation appears over hours.
- Ignoring threads. A global lock serialises parallel allocation.
- Benchmarks without boundary traffic. Glue allocations are part of the load. Include realistic JavaScript calls.
Performance note
On an allocation-heavy parser module, talc reduced compressed size by 6 KB and parse time by 14% compared with dlmalloc; a simple free-list allocator
saved 9 KB but parse time rose 35% and memory after a long replay was 2.3× higher.
Frequently Asked Questions
Does the allocator affect wasm-bindgen glue? No — glue calls exported allocation functions that use whichever global allocator you set.
Can I use different allocators in different modules of one app? Yes — each module has its own heap and allocator.
Is std::alloc::System different from the default on Wasm?
On wasm32-unknown-unknown, System is the dlmalloc port.
What about WASI targets?
WASI targets use wasi-libc’s malloc (dlmalloc-based); the same trade-offs apply.
Does boundary traffic affect the allocator choice? Yes — glue allocates for every string and slice passed in, so include realistic JavaScript calls in allocator benchmarks.
Should I set a larger initial memory instead of tuning the allocator? For a known working set, often yes — it removes growth during normal operation regardless of allocator, at the cost of reserving memory at startup.
How often should the choice be revisited? When allocation patterns change substantially or allocator crates release major versions; rerun the same benchmarks and long replay.
Related
- Replacing the default allocator to save bytes — the size angle.
- Benchmarking allocation-heavy Wasm code — measuring allocator cost.
- Implementing a free-list allocator in Wasm — how allocators work.
- Pooling fixed-size objects in Wasm — avoiding allocator churn.
← Back to Linear Memory Management & Allocators