Why Wasm Memory Never Shrinks

This page answers one task: a WebAssembly module’s memory stays at its peak long after the large operation that caused the peak has finished, and you need to understand why — and how to keep that peak, and therefore the page’s memory footprint, under control.

Prerequisites

  • [ ] A module whose memory grows during some operations: decoding a large file, building an index, processing a batch.
  • [ ] A way to observe memory size, such as memory.buffer.byteLength or the browser’s memory tools.

Growth is one-way by design

WebAssembly gives linear memory exactly one way to change size: memory.grow, which adds whole 64 KiB pages at the end. There is no memory.shrink, and no way for a module to hand pages back to the engine. Once a module’s memory has grown to 800 MB to decode a huge image, it stays 800 MB until the instance is discarded — even if the allocator inside the module has freed every byte.

The reasons are practical. Shrinking would invalidate every pointer into the removed region, and because memory is a flat array of bytes, the engine has no idea which values in memory are pointers — so it cannot check, and the module would have to guarantee nothing lives in the tail. Engines also rely on memory only growing: bounds checks are often implemented with large reserved address ranges and guard pages, and JavaScript code holds ArrayBuffers that would need an operation weirder than detachment to become smaller. A proposal to allow discarding page contents (memory.discard) has been discussed to let engines release physical memory while keeping the address range, but it is not part of the standard.

Memory size of one instance across a session Memory starts at 16 MB, grows to 820 MB while a large image is decoded, and stays at 820 MB after the decode finishes and all its buffers are freed, because linear memory cannot shrink. It only returns to a small size when the instance is replaced. 0 s start: 16 MB 8 s decode begins 14 s peak: 820 MB 20 s decode done, buffers freed — still 820 MB 45 s new instance: 16 MB

What “freed” means inside the module

When the module’s allocator frees a block, that memory returns to the allocator’s free list and can be reused for later allocations within the same instance. So the memory is not wasted for the module: the next big decode reuses the same 800 MB without growing. What it cannot do is return to the browser for other tabs, other pages or other parts of the same page. For a single-purpose page that repeatedly does the same large job, that is fine. For a long-lived app where the large job is rare, it means the app carries its worst-case footprint for the rest of the session.

On native platforms, allocators release large free regions back to the OS with munmap or madvise. That escape hatch does not exist for linear memory, so peak memory matters much more in WebAssembly than it does natively.

Pattern 1 — recycle the instance after big jobs

The only way to release memory is to drop the instance and its memory. If the large job is occasional, run it in a separate instance — or a worker — and discard that instance when the job finishes:

async function decodeLarge(bytes) {
  const { instance } = await WebAssembly.instantiate(compiledModule, imports());   // fresh, small memory
  try {
    return copyResultOut(instance, bytes);           // copy the result into JS before dropping
  } finally {
    // no references kept: instance and its 800 MB memory become garbage
  }
}

Instantiating from a cached WebAssembly.Module takes milliseconds, so the cost is small. The essential detail is copying the result out into JavaScript memory (or transferring it) before dropping the instance, and holding no views, exports or closures that keep the instance alive. Running the job in a worker and terminating the worker afterwards achieves the same thing even more decisively. Instantiation from a compiled module is covered in instantiating one module many times.

Keeping one instance versus recycling for big jobs A single long-lived instance keeps its peak memory for the whole session after one large job. Running large jobs in a temporary instance or worker returns memory when the job finishes, at the cost of a few milliseconds of instantiation. one long-lived instance big job grows memory to the peak peak kept for the rest of the session no instantiation cost fine if big jobs are frequent temporary instance per big job fresh memory for the job memory released when dropped a few ms to instantiate best for occasional big jobs

Pattern 2 — lower the peak with streaming

The cheapest memory to give back is memory never taken. Algorithms that process input in fixed-size chunks — decode an image in strips, parse a file as a stream, compress block by block — have a peak proportional to the chunk size, not the input size. For many formats this is the most effective change available: a 200 MB CSV import that loads the whole file needs over 400 MB of linear memory once parsed structures are included; a streaming parser that handles 1 MB at a time stays under 20 MB. Feeding data in chunks is described in streaming file uploads into Wasm memory.

Pattern 3 — keep big data outside linear memory

Data that JavaScript can hold does not need to live in linear memory. Large results can be produced in chunks and copied out to JavaScript as they are ready, so linear memory only ever holds the current chunk. JavaScript ArrayBuffers are garbage-collected normally, and their memory is returned when they are dropped. For an image editor, for example, the full-resolution layers can live in JavaScript or GPU memory, with the module operating on tiles copied in and out.

Pattern 4 — budget the peak and refuse what exceeds it

When neither recycling nor streaming applies, at least make the peak a decision. Set a maximum memory at link time, compute the memory a job will need from its input header before starting, and refuse jobs that exceed the budget with a clear message, as in handling out-of-memory in Wasm. A refused job is a much better user experience than a tab the operating system kills for using too much memory.

Pattern 5 — avoid accidental peaks

Some peaks are not inherent to the job but come from how it is coded. Common causes are holding the input and the output in memory at the same time when one could be released first; growing a vector by repeated doubling, which briefly holds the old and new buffers (up to three times the final size); copying input into linear memory when the module could read it in chunks; and keeping caches unbounded. Measuring the peak per operation, with the techniques in tracking linear memory growth over time, usually reveals one or two operations responsible for most of it. Pre-sizing buffers with the exact capacity and releasing inputs early can often halve the peak without changing the algorithm.

Servers and long-lived workers

On the server the same rule applies to each instance. Runtimes such as wasmtime typically handle it by instance pooling: an instance serves a request (or a bounded number of requests) and is then reset or replaced, which returns its memory to the pool. Long-lived instances that serve unbounded traffic accumulate the peak of the largest request they ever handled, so per-request memory limits and periodic recycling are standard practice. In the browser, long-lived workers that run a Wasm module for an entire session — a background indexer, a sync engine — benefit from the same approach: restart the worker when its memory exceeds a threshold during an idle moment. The restart costs tens of milliseconds and resets memory to its starting size.

Expected output

After switching occasional large decodes to a temporary instance, the page’s memory returns from 840 MB to 60 MB within a few seconds of each decode finishing (after garbage collection), and the browser’s task manager shows the tab’s footprint following the work instead of staying at the peak.

Gotchas

  • Expecting free to lower memory size. It only returns blocks to the allocator. The size never decreases.
  • Holding a reference to a discarded instance. One cached view or export keeps all its memory alive. Drop every reference.
  • Doubling growth for large buffers. Briefly needs up to three times the data. Reserve exact capacity.
  • Shared memory. Its maximum is reserved up front, and it cannot be released until all threads drop it.

Performance note

On a mid-range laptop, decoding a 120-megapixel image in the long-lived instance left the tab at 910 MB for the rest of the session. Decoding it in a temporary instance created from a cached module cost an extra 4 ms and the tab returned to 70 MB after the next garbage collection.

Tab memory ten seconds after a large decode Memory footprint of the tab after decoding a 120-megapixel image, decoding in the long-lived instance, in a temporary instance, and in a worker that was terminated afterwards. MB tab memory after the job long-lived instance 910 MB temporary instance 70 MB worker terminated 64 MB

Frequently Asked Questions

Will WebAssembly get a shrink instruction? Proposals have discussed discarding page contents to release physical memory, but there is no standard way to shrink linear memory today.

Does the browser release untouched pages? Reserved-but-untouched pages usually cost no physical memory. Pages that were written stay committed while the memory exists.

Does garbage collection help? Only once the whole instance and memory are unreachable. GC cannot reclaim part of a live linear memory.

Does WasmGC change this? Objects allocated with WasmGC live on the engine’s garbage-collected heap, not in linear memory, so they are reclaimed normally.

← Back to Linear Memory Management & Allocators