Using an Arena Allocator for Per-Frame Data
This page answers one task: a WebAssembly module allocates many short-lived objects every frame — a game loop, a renderer, a request handler —
and the cost of individual malloc/free calls, and the fragmentation they cause, should go away.
Prerequisites
- [ ] A workload with a clear unit of work: a frame, a request, a document parse.
- [ ] Temporary data that does not outlive that unit.
- [ ] Rust with the
bumpalocrate, or C with the ability to add a small allocator.
The pattern: allocate freely, free all at once
Much of the memory a frame allocates dies at the end of that frame: visibility lists, sorted draw calls, intermediate geometry, formatted strings, scratch buffers. A general-purpose allocator treats each of those objects individually — searching a free list to allocate, updating it again to free — and the mix of lifetimes leaves holes that fragment the heap over time.
An arena (also called a bump or region allocator) exploits the shared lifetime. It owns one large block of memory and a pointer. Allocation advances the pointer by the requested size; there is no individual free. At the end of the frame, the pointer resets to the start, freeing every object at once. Allocation becomes a handful of instructions, freeing becomes one, and fragmentation cannot happen because the arena’s memory is reused in the same order every frame.
Step 1 — create an arena in Rust with bumpalo
bumpalo is the standard Rust arena. It works on wasm32 without changes:
use bumpalo::Bump;
use std::cell::RefCell;
thread_local! {
static FRAME: RefCell<Bump> = RefCell::new(Bump::with_capacity(256 * 1024));
}
#[wasm_bindgen]
pub fn render_frame(t: f64) {
FRAME.with(|arena| {
let mut arena = arena.borrow_mut();
{
let a: &Bump = &arena;
let visible = bumpalo::collections::Vec::from_iter_in(scene().visible(t), a);
let mut draws = bumpalo::collections::Vec::with_capacity_in(visible.len(), a);
for obj in &visible { draws.push(obj.draw_call()); }
draws.sort_unstable_by_key(|d| d.material);
submit(&draws);
} // all arena borrows end here
arena.reset(); // free everything allocated this frame
});
}
Bump::with_capacity reserves an initial chunk so steady-state frames never touch the global allocator. If a frame needs more, bumpalo allocates
another chunk; reset() keeps the largest chunk and frees the rest, so the arena settles at the size of the biggest frame.
Step 2 — let the borrow checker enforce the lifetime
The danger with arenas is using an object after the reset. In Rust, collections allocated in a Bump borrow it, so arena.reset() — which needs
&mut Bump — cannot be called while any of them exist. The inner block above ends every borrow before the reset; moving arena.reset() inside it
fails to compile. That is the main reason to prefer bumpalo over a hand-rolled arena in Rust: use-after-reset becomes a compile error instead of
memory corruption.
What the borrow checker cannot catch is data copied out of the arena by pointer — for example, an address passed to JavaScript. Never hand an arena pointer to JavaScript unless JavaScript is done with it before the reset.
Step 3 — write a simple arena in C
In C the arena is a few lines over a buffer allocated once:
typedef struct { unsigned char *base; size_t cap, used; } Arena;
static Arena frame;
void arena_init(size_t cap) { frame.base = malloc(cap); frame.cap = cap; frame.used = 0; }
void *arena_alloc(size_t n) {
size_t off = (frame.used + 7) & ~(size_t)7; /* 8-byte alignment */
if (off + n > frame.cap) return NULL; /* caller falls back or fails */
frame.used = off + n;
return frame.base + off;
}
void arena_reset(void) { frame.used = 0; }
A fixed-capacity arena makes peak memory explicit: if a frame needs more than the budget, arena_alloc returns NULL and the code can fall back to
malloc for the overflow or degrade gracefully. Pick the capacity from measurement — log the high-water mark of frame.used during development.
Step 4 — reset at the right moment
Reset when the unit of work is completely finished, including anything that reads its output. For a renderer that means after the draw calls are submitted; for a request handler, after the response bytes have been copied out. A common bug is resetting at the start of the next frame instead — which works, but holds the previous frame’s memory across the gap — or resetting before JavaScript has read a result that lives in the arena. If JavaScript reads frame output, copy it out first or move the reset to the next call.
Step 5 — measure the change
Compare frame times and heap behaviour before and after. Count allocations per frame through the global allocator — they should drop to near zero
for the converted paths, which you can confirm with
counting allocations with a wrapping allocator.
Watch memory.buffer.byteLength over a long session: with an arena, it should rise to a plateau in the first seconds and stay there, rather than creeping
upward as fragmentation accumulates.
Choosing what goes in the arena
Not everything allocated during a frame belongs in the arena. The test is lifetime: does the object die before the reset? Scratch vectors, sort buffers,
temporary strings for formatting, intermediate meshes and per-frame command lists all do. Objects that survive — a texture created on first use, an
entity spawned this frame, a cache entry — must use the general allocator, or a second, longer-lived arena reset on a slower cadence such as a level
load. Mixing lifetimes in one arena is the root of most arena bugs. A useful discipline is to make arena-allocated types visibly different: in Rust they
carry the 'bump lifetime, so they cannot be stored in long-lived structures by accident; in C, prefix arena functions and comment the lifetime at every
call site. Several arenas with different reset points — per frame, per level, per session — give most of the benefit of garbage collection with none of
its pauses, and that pattern is common in game engines for exactly that reason.
Arenas for request handlers on the server
The same pattern fits server-side WebAssembly. A request handler running in wasmtime, Spin or an edge runtime parses headers, builds intermediate
structures and formats a response, and almost all of it is garbage once the response is written. Allocating it from a per-request arena, reset when the
handler returns, keeps the guest’s memory flat across millions of requests — important because, as with browsers, a guest’s linear memory never
shrinks, and fragmentation in a long-lived instance turns into steadily rising memory per instance. Many serverless platforms sidestep the issue by
creating a fresh instance for every request, in which case the instance itself is the arena: everything is freed when it is discarded. Where instances
are reused for throughput, an explicit arena gives the same guarantee without the cost of re-instantiation. In Rust, a Bump created at the top of the
handler and passed by reference to the parsing and formatting code is enough; the response bytes are copied out to the host before the handler returns
and the arena is dropped or reset.
Expected output
Allocations through the global allocator during a steady-state frame drop from about 1,800 to 12; frame time falls by a few percent; and linear memory reaches a plateau within two seconds and stays flat for a ten-minute session.
Gotchas
- Using arena data after reset. In C this silently reads the next frame’s data. In Rust, keep borrows scoped so reset cannot compile early.
- Long-lived data in the arena. It disappears at reset. Allocate it from the general allocator.
- Handing arena pointers to JavaScript. JavaScript may read after reset. Copy results out first.
- Arena too small. Overflow falls back to a new chunk (bumpalo) or
NULL(fixed C arena). Size from the measured high-water mark. - Destructors not running.
bumpalodoes not runDropfor values allocated in it by default. Keep arena types free of resources that need dropping.
Performance note
In a particle renderer with 20,000 particles, moving per-frame scratch data from the global allocator to a bumpalo arena cut allocation calls from
1,800 to 12 per frame and frame time from 6.9 ms to 5.8 ms in Chrome. Memory reached its plateau after the first second instead of growing by 14 MB over
ten minutes.
Frequently Asked Questions
Is an arena the same as a bump allocator? Yes, in practice — an arena is a bump allocator with a reset, often with the ability to add chunks when it fills up.
Can several threads share one arena? Not safely without locking. Give each thread its own arena, which also avoids contention.
Does bumpalo increase binary size much? A few kilobytes. It is far smaller than the allocation work it removes from hot paths.
What about Emscripten?
Use a C arena as above, or C++ std::pmr::monotonic_buffer_resource, which provides the same pattern for standard containers.
Related
- Implementing a bump allocator in Wasm — the mechanism underneath.
- Implementing a free-list allocator in Wasm — for data with individual lifetimes.
- Measuring allocator fragmentation in Wasm — the problem arenas avoid.
- Porting a C game loop to Emscripten — where per-frame arenas fit.
← Back to Linear Memory Management & Allocators