Using Bulk Memory Operations

This page answers one question: what do WebAssembly’s bulk memory instructions do, when do compilers use them, and how much do they speed up copying and initialising memory compared with loops?

Prerequisites

  • [ ] WABT for hand-written examples and wasm-objdump.
  • [ ] A Rust, C or C++ toolchain targeting wasm32 (bulk memory is enabled by default in recent releases).
  • [ ] A browser or runtime from the last few years — every current engine supports the proposal.

What the MVP was missing

The first version of WebAssembly could only move memory one load and one store at a time. Copying a buffer meant a loop of i64.load and i64.store; zeroing memory meant a loop of stores; initialising memory from a data segment happened only once, at instantiation, for the whole segment. Compilers implemented memcpy and memset as WebAssembly functions full of such loops, and those functions were among the most frequently called code in many modules.

The bulk memory proposal, now standard, added instructions that move whole ranges at once. memory.copy copies a range of bytes, correctly handling overlap like memmove. memory.fill sets a range to one byte value. memory.init copies from a passive data segment into memory at any time, and data.drop releases a segment no longer needed. Equivalent table instructions exist for function tables. Engines implement these with optimized native routines — the same ones the browser uses for its own copies — so a single instruction runs at close to memory bandwidth.

Bulk memory instructions and what they replace The instructions added by the bulk memory proposal, the stack operands each takes, and the MVP-era code each replaces. instruction operands replaces memory.copy dst, src, len memcpy / memmove loops memory.fill dst, value, len memset loops memory.init dst, offset, len (+ segment) copying data written at startup data.drop (segment) keeping unused data alive table.copy / table.init dst, src, len table setup by hand

Step 1 — use the instructions directly in WAT

(module
  (memory (export "memory") 1)

  ;; copy len bytes from src to dst (overlap-safe, like memmove)
  (func (export "copy") (param $dst i32) (param $src i32) (param $len i32)
    (memory.copy (local.get $dst) (local.get $src) (local.get $len)))

  ;; zero len bytes at dst
  (func (export "zero") (param $dst i32) (param $len i32)
    (memory.fill (local.get $dst) (i32.const 0) (local.get $len))))
wat2wasm --enable-bulk-memory bulk.wat -o bulk.wasm      # flag only needed on old wabt releases

Both instructions trap if any part of either range is outside memory, before writing anything — there is no partial copy to clean up.

Step 2 — check that your compiler emits them

Recent clang and rustc enable the bulk-memory feature by default for wasm32 targets, and LLVM then lowers memcpy, memmove and memset — including the implicit ones from struct copies and array initialisation — to the instructions directly:

cargo build --release --target wasm32-unknown-unknown
wasm2wat target/wasm32-unknown-unknown/release/app.wasm | grep -cE 'memory\.(copy|fill)'

A count of zero on a module that clearly copies memory means the feature is off — often because an older toolchain or an explicit -C target-feature=-bulk-memory was used for compatibility. Turn it on explicitly where needed:

RUSTFLAGS="-C target-feature=+bulk-memory" cargo build --release --target wasm32-unknown-unknown
clang --target=wasm32 -O2 -mbulk-memory copy.c -c -o copy.o

For Emscripten, -mbulk-memory (or the default in recent versions) has the same effect.

Step 3 — initialise data lazily with passive segments

Normal (active) data segments are copied into memory at instantiation, and their bytes stay in the module afterwards. Passive segments are not copied automatically; the module copies them when it wants, as many times as it wants, and can drop them when done:

(module
  (memory (export "memory") 1)
  (data $greeting "Hello from a passive segment")       ;; passive: no offset given

  (func (export "load_greeting") (param $dst i32) (result i32)
    (memory.init $greeting (local.get $dst) (i32.const 0) (i32.const 28))
    (i32.const 28))

  (func (export "release")
    (data.drop $greeting)))                               ;; further memory.init traps

Passive segments are how threaded modules initialise shared memory exactly once — the first thread runs memory.init, the others skip it — which wasm-ld --shared-memory generates automatically, as noted in importing memory from the host with --import-memory. They are also useful for large lookup tables needed only in some code paths: copy them in on first use instead of at every instantiation.

Active versus passive data segments over a module's life An active segment is written into memory during instantiation. A passive segment stays in the module until code calls memory.init to copy it, possibly more than once, and data.drop releases it. engine module code linear memory instantiate: write active segments memory.init $table → copy on first use memory.init $table again (e.g. reset) data.drop $table → segment released

Step 4 — use them from JavaScript-facing code

When JavaScript needs to move data within the module’s memory — compacting a buffer, shifting a ring buffer’s contents — call an exported copy rather than doing it with typed arrays. TypedArray.prototype.copyWithin on the memory’s buffer is also fast and equivalent; the export simply keeps the operation inside the module, where it composes with the module’s own bookkeeping:

const mem = new Uint8Array(instance.exports.memory.buffer);
mem.copyWithin(0, 4096, 4096 + 65536);          // JS-side equivalent of memory.copy(0, 4096, 65536)
instance.exports.copy(0, 4096, 65536);          // same result, from inside the module

For filling large regions — clearing a framebuffer each frame — memory.fill from inside the module and TypedArray.prototype.fill from JavaScript are both implemented natively and run at similar speed.

Step 5 — measure where it matters

The gain is largest where code copies or clears a lot of memory: image and audio buffers, serialisation, allocators that zero memory, emulators that move framebuffers. For small copies — a few dozen bytes — the difference is negligible, and the optimizer may inline a few loads and stores rather than call the instruction. The benchmark in benchmarking memory bandwidth in Wasm shows memory.copy reaching close to the machine’s memory bandwidth for large ranges.

Why size improves as well as speed

Bulk memory changes module size as well as speed, which is easy to miss. Without it, every module carries its own memcpy, memmove and memset — loop-based functions of a few hundred bytes each, often in several variants for different alignments. With it, those functions become single instructions at each call site, or tiny wrappers, and dead code elimination removes the loop implementations. Modules also shrink when struct copies and array initialisations stop expanding into inline load-store sequences. For small modules the saving is proportionally significant — and freestanding builds, such as those in building a Wasm module without libc, can drop their hand-written memcpy entirely.

When to keep the MVP form

There are still occasional reasons to build without bulk memory. Some embedded and blockchain virtual machines implement only the MVP instruction set and reject modules that use newer opcodes; Rust’s wasm32v1-none target exists precisely for such hosts. Some tooling in a pipeline — an older static analyser, a custom instrumentation pass — may not understand the new instructions. In those cases, disable the feature for that build only, keep the default elsewhere, and treat the MVP build as a compatibility artifact with its own tests rather than as the main product. For browsers, server runtimes and edge platforms there is no remaining reason to disable it.

Expected output

wasm2wat bulk.wasm | grep -E 'memory\.(copy|fill|init)|data\.drop'
    memory.copy
    memory.fill

Calling copy and zero from JavaScript moves and clears the expected ranges, and an out-of-bounds call throws RuntimeError: memory access out of bounds without modifying memory.

Gotchas

  • unknown opcode when instantiating. The engine is very old or the module targets an embedder that does not support bulk memory. Build without it for that target only.
  • Expecting partial writes. Out-of-bounds bulk operations trap before writing anything; code that relied on MVP loops writing up to the fault behaves differently.
  • memory.init after data.drop traps. Dropped segments cannot be reused. Drop only when certain.
  • Tests that check intermediate state. A test that watched a buffer being filled byte by byte sees it change all at once. Assert on the final state.
  • Optimizer inlines small copies. Seeing loads and stores for a 16-byte struct copy is expected and fast.

Performance note

Copying a 64 MB buffer with an MVP-style word loop took 9.8 ms in Chrome on a laptop; memory.copy took 3.1 ms. Zeroing the same buffer took 7.2 ms with a store loop and 2.4 ms with memory.fill. Enabling bulk memory also removed 1.9 KB of memcpy and memset implementations from a 140 KB module.

Copying and clearing 64 MB, loops versus bulk instructions Time to copy and to zero a 64 MB buffer in Chrome using MVP-style 64-bit load and store loops and using memory.copy and memory.fill. ms for a 64 MB operation copy, i64 loop 9.8 ms memory.copy 3.1 ms zero, i64 store loop 7.2 ms memory.fill 2.4 ms

Frequently Asked Questions

Is memory.copy safe for overlapping ranges? Yes. It behaves like memmove, copying as if through a temporary buffer.

Do I need to do anything to benefit? Usually not — recent toolchains enable the feature and lower copies automatically. Check the output once to be sure.

Is memory.fill limited to zero? No — it fills with any byte value, which makes it useful for initialising buffers to a sentinel pattern while debugging, such as 0xAA canaries around allocations.

What about tables? table.copy and table.init are the table equivalents, used by dynamic linking and runtimes that manage function tables; see initialising tables with element segments.

Does bulk memory work with shared memory? Yes. On shared memory, bulk operations are not atomic as a whole, so synchronise threads around them as you would around any copy.

← Back to Post-MVP Wasm Proposals in Practice