Using an Arena Allocator for Per-Frame Data

This page answers one task: a WebAssembly module allocates many short-lived objects every frame — a game loop, a renderer, a request handler — and the cost of individual malloc/free calls, and the fragmentation they cause, should go away.

Prerequisites

  • [ ] A workload with a clear unit of work: a frame, a request, a document parse.
  • [ ] Temporary data that does not outlive that unit.
  • [ ] Rust with the bumpalo crate, or C with the ability to add a small allocator.

The pattern: allocate freely, free all at once

Much of the memory a frame allocates dies at the end of that frame: visibility lists, sorted draw calls, intermediate geometry, formatted strings, scratch buffers. A general-purpose allocator treats each of those objects individually — searching a free list to allocate, updating it again to free — and the mix of lifetimes leaves holes that fragment the heap over time.

An arena (also called a bump or region allocator) exploits the shared lifetime. It owns one large block of memory and a pointer. Allocation advances the pointer by the requested size; there is no individual free. At the end of the frame, the pointer resets to the start, freeing every object at once. Allocation becomes a handful of instructions, freeing becomes one, and fragmentation cannot happen because the arena’s memory is reused in the same order every frame.

Arena memory across three frames During each frame the arena pointer advances as temporary objects are allocated. At the end of the frame it resets to the start in one step, and the next frame reuses the same memory from the beginning. 0 ms frame 1: pointer at 0 9 ms 120 KB allocated 16.7 ms reset (end of frame 1) 26 ms 96 KB allocated 33.4 ms reset (end of frame 2) 43 ms 140 KB allocated

Step 1 — create an arena in Rust with bumpalo

bumpalo is the standard Rust arena. It works on wasm32 without changes:

use bumpalo::Bump;
use std::cell::RefCell;

thread_local! {
    static FRAME: RefCell<Bump> = RefCell::new(Bump::with_capacity(256 * 1024));
}

#[wasm_bindgen]
pub fn render_frame(t: f64) {
    FRAME.with(|arena| {
        let mut arena = arena.borrow_mut();
        {
            let a: &Bump = &arena;
            let visible = bumpalo::collections::Vec::from_iter_in(scene().visible(t), a);
            let mut draws = bumpalo::collections::Vec::with_capacity_in(visible.len(), a);
            for obj in &visible { draws.push(obj.draw_call()); }
            draws.sort_unstable_by_key(|d| d.material);
            submit(&draws);
        }                                   // all arena borrows end here
        arena.reset();                      // free everything allocated this frame
    });
}

Bump::with_capacity reserves an initial chunk so steady-state frames never touch the global allocator. If a frame needs more, bumpalo allocates another chunk; reset() keeps the largest chunk and frees the rest, so the arena settles at the size of the biggest frame.

Step 2 — let the borrow checker enforce the lifetime

The danger with arenas is using an object after the reset. In Rust, collections allocated in a Bump borrow it, so arena.reset() — which needs &mut Bump — cannot be called while any of them exist. The inner block above ends every borrow before the reset; moving arena.reset() inside it fails to compile. That is the main reason to prefer bumpalo over a hand-rolled arena in Rust: use-after-reset becomes a compile error instead of memory corruption.

What the borrow checker cannot catch is data copied out of the arena by pointer — for example, an address passed to JavaScript. Never hand an arena pointer to JavaScript unless JavaScript is done with it before the reset.

Step 3 — write a simple arena in C

In C the arena is a few lines over a buffer allocated once:

typedef struct { unsigned char *base; size_t cap, used; } Arena;

static Arena frame;

void arena_init(size_t cap) { frame.base = malloc(cap); frame.cap = cap; frame.used = 0; }

void *arena_alloc(size_t n) {
    size_t off = (frame.used + 7) & ~(size_t)7;          /* 8-byte alignment */
    if (off + n > frame.cap) return NULL;                 /* caller falls back or fails */
    frame.used = off + n;
    return frame.base + off;
}

void arena_reset(void) { frame.used = 0; }

A fixed-capacity arena makes peak memory explicit: if a frame needs more than the budget, arena_alloc returns NULL and the code can fall back to malloc for the overflow or degrade gracefully. Pick the capacity from measurement — log the high-water mark of frame.used during development.

Step 4 — reset at the right moment

Reset when the unit of work is completely finished, including anything that reads its output. For a renderer that means after the draw calls are submitted; for a request handler, after the response bytes have been copied out. A common bug is resetting at the start of the next frame instead — which works, but holds the previous frame’s memory across the gap — or resetting before JavaScript has read a result that lives in the arena. If JavaScript reads frame output, copy it out first or move the reset to the next call.

General-purpose allocator versus arena for frame data With a general allocator every temporary object is allocated and freed individually and the heap fragments over time. With an arena, allocation is a pointer bump, freeing is a single reset, and memory use stays flat at the size of the largest frame. malloc / free per object search free list on each allocation one free call per object fragmentation grows over time fine for long-lived data arena per frame pointer bump on each allocation one reset per frame memory flat at the peak frame size best for temporary data

Step 5 — measure the change

Compare frame times and heap behaviour before and after. Count allocations per frame through the global allocator — they should drop to near zero for the converted paths, which you can confirm with counting allocations with a wrapping allocator. Watch memory.buffer.byteLength over a long session: with an arena, it should rise to a plateau in the first seconds and stay there, rather than creeping upward as fragmentation accumulates.

Choosing what goes in the arena

Not everything allocated during a frame belongs in the arena. The test is lifetime: does the object die before the reset? Scratch vectors, sort buffers, temporary strings for formatting, intermediate meshes and per-frame command lists all do. Objects that survive — a texture created on first use, an entity spawned this frame, a cache entry — must use the general allocator, or a second, longer-lived arena reset on a slower cadence such as a level load. Mixing lifetimes in one arena is the root of most arena bugs. A useful discipline is to make arena-allocated types visibly different: in Rust they carry the 'bump lifetime, so they cannot be stored in long-lived structures by accident; in C, prefix arena functions and comment the lifetime at every call site. Several arenas with different reset points — per frame, per level, per session — give most of the benefit of garbage collection with none of its pauses, and that pattern is common in game engines for exactly that reason.

Arenas for request handlers on the server

The same pattern fits server-side WebAssembly. A request handler running in wasmtime, Spin or an edge runtime parses headers, builds intermediate structures and formats a response, and almost all of it is garbage once the response is written. Allocating it from a per-request arena, reset when the handler returns, keeps the guest’s memory flat across millions of requests — important because, as with browsers, a guest’s linear memory never shrinks, and fragmentation in a long-lived instance turns into steadily rising memory per instance. Many serverless platforms sidestep the issue by creating a fresh instance for every request, in which case the instance itself is the arena: everything is freed when it is discarded. Where instances are reused for throughput, an explicit arena gives the same guarantee without the cost of re-instantiation. In Rust, a Bump created at the top of the handler and passed by reference to the parsing and formatting code is enough; the response bytes are copied out to the host before the handler returns and the arena is dropped or reset.

Expected output

Allocations through the global allocator during a steady-state frame drop from about 1,800 to 12; frame time falls by a few percent; and linear memory reaches a plateau within two seconds and stays flat for a ten-minute session.

Gotchas

  • Using arena data after reset. In C this silently reads the next frame’s data. In Rust, keep borrows scoped so reset cannot compile early.
  • Long-lived data in the arena. It disappears at reset. Allocate it from the general allocator.
  • Handing arena pointers to JavaScript. JavaScript may read after reset. Copy results out first.
  • Arena too small. Overflow falls back to a new chunk (bumpalo) or NULL (fixed C arena). Size from the measured high-water mark.
  • Destructors not running. bumpalo does not run Drop for values allocated in it by default. Keep arena types free of resources that need dropping.

Performance note

In a particle renderer with 20,000 particles, moving per-frame scratch data from the global allocator to a bumpalo arena cut allocation calls from 1,800 to 12 per frame and frame time from 6.9 ms to 5.8 ms in Chrome. Memory reached its plateau after the first second instead of growing by 14 MB over ten minutes.

Frame time with and without a per-frame arena Average time to build and submit a frame for 20,000 particles, allocating scratch data from the global allocator and from a per-frame bumpalo arena. ms per frame global allocator 6.9 ms per-frame arena 5.8 ms

Frequently Asked Questions

Is an arena the same as a bump allocator? Yes, in practice — an arena is a bump allocator with a reset, often with the ability to add chunks when it fills up.

Can several threads share one arena? Not safely without locking. Give each thread its own arena, which also avoids contention.

Does bumpalo increase binary size much? A few kilobytes. It is far smaller than the allocation work it removes from hot paths.

What about Emscripten? Use a C arena as above, or C++ std::pmr::monotonic_buffer_resource, which provides the same pattern for standard containers.

← Back to Linear Memory Management & Allocators