Why Wasm Memory Never Shrinks
This page answers one task: a WebAssembly module’s memory stays at its peak long after the large operation that caused the peak has finished, and you need to understand why — and how to keep that peak, and therefore the page’s memory footprint, under control.
Prerequisites
- [ ] A module whose memory grows during some operations: decoding a large file, building an index, processing a batch.
- [ ] A way to observe memory size, such as
memory.buffer.byteLengthor the browser’s memory tools.
Growth is one-way by design
WebAssembly gives linear memory exactly one way to change size: memory.grow, which adds whole 64 KiB pages at the end. There is no memory.shrink,
and no way for a module to hand pages back to the engine. Once a module’s memory has grown to 800 MB to decode a huge image, it stays 800 MB until the
instance is discarded — even if the allocator inside the module has freed every byte.
The reasons are practical. Shrinking would invalidate every pointer into the removed region, and because memory is a flat array of bytes, the engine has
no idea which values in memory are pointers — so it cannot check, and the module would have to guarantee nothing lives in the tail. Engines also rely on
memory only growing: bounds checks are often implemented with large reserved address ranges and guard pages, and JavaScript code holds ArrayBuffers
that would need an operation weirder than detachment to become smaller. A proposal to allow discarding page contents (memory.discard) has been
discussed to let engines release physical memory while keeping the address range, but it is not part of the standard.
What “freed” means inside the module
When the module’s allocator frees a block, that memory returns to the allocator’s free list and can be reused for later allocations within the same instance. So the memory is not wasted for the module: the next big decode reuses the same 800 MB without growing. What it cannot do is return to the browser for other tabs, other pages or other parts of the same page. For a single-purpose page that repeatedly does the same large job, that is fine. For a long-lived app where the large job is rare, it means the app carries its worst-case footprint for the rest of the session.
On native platforms, allocators release large free regions back to the OS with munmap or madvise. That escape hatch does not exist for linear
memory, so peak memory matters much more in WebAssembly than it does natively.
Pattern 1 — recycle the instance after big jobs
The only way to release memory is to drop the instance and its memory. If the large job is occasional, run it in a separate instance — or a worker — and discard that instance when the job finishes:
async function decodeLarge(bytes) {
const { instance } = await WebAssembly.instantiate(compiledModule, imports()); // fresh, small memory
try {
return copyResultOut(instance, bytes); // copy the result into JS before dropping
} finally {
// no references kept: instance and its 800 MB memory become garbage
}
}
Instantiating from a cached WebAssembly.Module takes milliseconds, so the cost is small. The essential detail is copying the result out into JavaScript
memory (or transferring it) before dropping the instance, and holding no views, exports or closures that keep the instance alive. Running the job in a
worker and terminating the worker afterwards achieves the same thing even more decisively. Instantiation from a compiled module is covered in
instantiating one module many times.
Pattern 2 — lower the peak with streaming
The cheapest memory to give back is memory never taken. Algorithms that process input in fixed-size chunks — decode an image in strips, parse a file as a stream, compress block by block — have a peak proportional to the chunk size, not the input size. For many formats this is the most effective change available: a 200 MB CSV import that loads the whole file needs over 400 MB of linear memory once parsed structures are included; a streaming parser that handles 1 MB at a time stays under 20 MB. Feeding data in chunks is described in streaming file uploads into Wasm memory.
Pattern 3 — keep big data outside linear memory
Data that JavaScript can hold does not need to live in linear memory. Large results can be produced in chunks and copied out to JavaScript as they are
ready, so linear memory only ever holds the current chunk. JavaScript ArrayBuffers are garbage-collected normally, and their memory is returned when
they are dropped. For an image editor, for example, the full-resolution layers can live in JavaScript or GPU memory, with the module operating on tiles
copied in and out.
Pattern 4 — budget the peak and refuse what exceeds it
When neither recycling nor streaming applies, at least make the peak a decision. Set a maximum memory at link time, compute the memory a job will need from its input header before starting, and refuse jobs that exceed the budget with a clear message, as in handling out-of-memory in Wasm. A refused job is a much better user experience than a tab the operating system kills for using too much memory.
Pattern 5 — avoid accidental peaks
Some peaks are not inherent to the job but come from how it is coded. Common causes are holding the input and the output in memory at the same time when one could be released first; growing a vector by repeated doubling, which briefly holds the old and new buffers (up to three times the final size); copying input into linear memory when the module could read it in chunks; and keeping caches unbounded. Measuring the peak per operation, with the techniques in tracking linear memory growth over time, usually reveals one or two operations responsible for most of it. Pre-sizing buffers with the exact capacity and releasing inputs early can often halve the peak without changing the algorithm.
Servers and long-lived workers
On the server the same rule applies to each instance. Runtimes such as wasmtime typically handle it by instance pooling: an instance serves a request (or a bounded number of requests) and is then reset or replaced, which returns its memory to the pool. Long-lived instances that serve unbounded traffic accumulate the peak of the largest request they ever handled, so per-request memory limits and periodic recycling are standard practice. In the browser, long-lived workers that run a Wasm module for an entire session — a background indexer, a sync engine — benefit from the same approach: restart the worker when its memory exceeds a threshold during an idle moment. The restart costs tens of milliseconds and resets memory to its starting size.
Expected output
After switching occasional large decodes to a temporary instance, the page’s memory returns from 840 MB to 60 MB within a few seconds of each decode finishing (after garbage collection), and the browser’s task manager shows the tab’s footprint following the work instead of staying at the peak.
Gotchas
- Expecting
freeto lower memory size. It only returns blocks to the allocator. The size never decreases. - Holding a reference to a discarded instance. One cached view or export keeps all its memory alive. Drop every reference.
- Doubling growth for large buffers. Briefly needs up to three times the data. Reserve exact capacity.
- Shared memory. Its maximum is reserved up front, and it cannot be released until all threads drop it.
Performance note
On a mid-range laptop, decoding a 120-megapixel image in the long-lived instance left the tab at 910 MB for the rest of the session. Decoding it in a temporary instance created from a cached module cost an extra 4 ms and the tab returned to 70 MB after the next garbage collection.
Frequently Asked Questions
Will WebAssembly get a shrink instruction? Proposals have discussed discarding page contents to release physical memory, but there is no standard way to shrink linear memory today.
Does the browser release untouched pages? Reserved-but-untouched pages usually cost no physical memory. Pages that were written stay committed while the memory exists.
Does garbage collection help? Only once the whole instance and memory are unreachable. GC cannot reclaim part of a live linear memory.
Does WasmGC change this? Objects allocated with WasmGC live on the engine’s garbage-collected heap, not in linear memory, so they are reclaimed normally.
Related
- Growing memory safely from JavaScript — the growth side.
- Using an arena allocator for per-frame data — keeping temporary memory flat.
- Measuring allocator fragmentation in Wasm — when the peak keeps rising.
- Monitoring Wasm memory in production — seeing peaks in the field.
← Back to Linear Memory Management & Allocators