Using Bulk Memory Operations
This page answers one question: what do WebAssembly’s bulk memory instructions do, when do compilers use them, and how much do they speed up copying and initialising memory compared with loops?
Prerequisites
- [ ] WABT for hand-written examples and
wasm-objdump. - [ ] A Rust, C or C++ toolchain targeting
wasm32(bulk memory is enabled by default in recent releases). - [ ] A browser or runtime from the last few years — every current engine supports the proposal.
What the MVP was missing
The first version of WebAssembly could only move memory one load and one store at a time. Copying a buffer meant a loop of i64.load and
i64.store; zeroing memory meant a loop of stores; initialising memory from a data segment happened only once, at instantiation, for the
whole segment. Compilers implemented memcpy and memset as WebAssembly functions full of such loops, and those functions were among the
most frequently called code in many modules.
The bulk memory proposal, now standard, added instructions that move whole ranges at once. memory.copy copies a range of bytes, correctly
handling overlap like memmove. memory.fill sets a range to one byte value. memory.init copies from a passive data segment into memory
at any time, and data.drop releases a segment no longer needed. Equivalent table instructions exist for function tables. Engines implement
these with optimized native routines — the same ones the browser uses for its own copies — so a single instruction runs at close to memory
bandwidth.
Step 1 — use the instructions directly in WAT
(module
(memory (export "memory") 1)
;; copy len bytes from src to dst (overlap-safe, like memmove)
(func (export "copy") (param $dst i32) (param $src i32) (param $len i32)
(memory.copy (local.get $dst) (local.get $src) (local.get $len)))
;; zero len bytes at dst
(func (export "zero") (param $dst i32) (param $len i32)
(memory.fill (local.get $dst) (i32.const 0) (local.get $len))))
wat2wasm --enable-bulk-memory bulk.wat -o bulk.wasm # flag only needed on old wabt releases
Both instructions trap if any part of either range is outside memory, before writing anything — there is no partial copy to clean up.
Step 2 — check that your compiler emits them
Recent clang and rustc enable the bulk-memory feature by default for wasm32 targets, and LLVM then lowers memcpy, memmove and
memset — including the implicit ones from struct copies and array initialisation — to the instructions directly:
cargo build --release --target wasm32-unknown-unknown
wasm2wat target/wasm32-unknown-unknown/release/app.wasm | grep -cE 'memory\.(copy|fill)'
A count of zero on a module that clearly copies memory means the feature is off — often because an older toolchain or an explicit
-C target-feature=-bulk-memory was used for compatibility. Turn it on explicitly where needed:
RUSTFLAGS="-C target-feature=+bulk-memory" cargo build --release --target wasm32-unknown-unknown
clang --target=wasm32 -O2 -mbulk-memory copy.c -c -o copy.o
For Emscripten, -mbulk-memory (or the default in recent versions) has the same effect.
Step 3 — initialise data lazily with passive segments
Normal (active) data segments are copied into memory at instantiation, and their bytes stay in the module afterwards. Passive segments are not copied automatically; the module copies them when it wants, as many times as it wants, and can drop them when done:
(module
(memory (export "memory") 1)
(data $greeting "Hello from a passive segment") ;; passive: no offset given
(func (export "load_greeting") (param $dst i32) (result i32)
(memory.init $greeting (local.get $dst) (i32.const 0) (i32.const 28))
(i32.const 28))
(func (export "release")
(data.drop $greeting))) ;; further memory.init traps
Passive segments are how threaded modules initialise shared memory exactly once — the first thread runs memory.init, the others skip it —
which wasm-ld --shared-memory generates automatically, as noted in
importing memory from the host with --import-memory.
They are also useful for large lookup tables needed only in some code paths: copy them in on first use instead of at every instantiation.
Step 4 — use them from JavaScript-facing code
When JavaScript needs to move data within the module’s memory — compacting a buffer, shifting a ring buffer’s contents — call an exported
copy rather than doing it with typed arrays. TypedArray.prototype.copyWithin on the memory’s buffer is also fast and equivalent; the
export simply keeps the operation inside the module, where it composes with the module’s own bookkeeping:
const mem = new Uint8Array(instance.exports.memory.buffer);
mem.copyWithin(0, 4096, 4096 + 65536); // JS-side equivalent of memory.copy(0, 4096, 65536)
instance.exports.copy(0, 4096, 65536); // same result, from inside the module
For filling large regions — clearing a framebuffer each frame — memory.fill from inside the module and TypedArray.prototype.fill from
JavaScript are both implemented natively and run at similar speed.
Step 5 — measure where it matters
The gain is largest where code copies or clears a lot of memory: image and audio buffers, serialisation, allocators that zero memory,
emulators that move framebuffers. For small copies — a few dozen bytes — the difference is negligible, and the optimizer may inline a few
loads and stores rather than call the instruction. The benchmark in
benchmarking memory bandwidth in Wasm
shows memory.copy reaching close to the machine’s memory bandwidth for large ranges.
Why size improves as well as speed
Bulk memory changes module size as well as speed, which is easy to miss. Without it, every module carries its own memcpy, memmove and
memset — loop-based functions of a few hundred bytes each, often in several variants for different alignments. With it, those functions
become single instructions at each call site, or tiny wrappers, and dead code elimination removes the loop implementations. Modules also
shrink when struct copies and array initialisations stop expanding into inline load-store sequences. For small modules the saving is
proportionally significant — and freestanding builds, such as those in
building a Wasm module without libc,
can drop their hand-written memcpy entirely.
When to keep the MVP form
There are still occasional reasons to build without bulk memory. Some embedded and blockchain virtual machines implement only the MVP
instruction set and reject modules that use newer opcodes; Rust’s wasm32v1-none target exists precisely for such hosts. Some tooling in
a pipeline — an older static analyser, a custom instrumentation pass — may not understand the new instructions. In those cases, disable the
feature for that build only, keep the default elsewhere, and treat the MVP build as a compatibility artifact with its own tests rather than
as the main product. For browsers, server runtimes and edge platforms there is no remaining reason to disable it.
Expected output
wasm2wat bulk.wasm | grep -E 'memory\.(copy|fill|init)|data\.drop'
memory.copy
memory.fill
Calling copy and zero from JavaScript moves and clears the expected ranges, and an out-of-bounds call throws RuntimeError: memory access out of bounds without modifying memory.
Gotchas
unknown opcodewhen instantiating. The engine is very old or the module targets an embedder that does not support bulk memory. Build without it for that target only.- Expecting partial writes. Out-of-bounds bulk operations trap before writing anything; code that relied on MVP loops writing up to the fault behaves differently.
memory.initafterdata.droptraps. Dropped segments cannot be reused. Drop only when certain.- Tests that check intermediate state. A test that watched a buffer being filled byte by byte sees it change all at once. Assert on the final state.
- Optimizer inlines small copies. Seeing loads and stores for a 16-byte struct copy is expected and fast.
Performance note
Copying a 64 MB buffer with an MVP-style word loop took 9.8 ms in Chrome on a laptop; memory.copy took 3.1 ms. Zeroing the same buffer
took 7.2 ms with a store loop and 2.4 ms with memory.fill. Enabling bulk memory also removed 1.9 KB of memcpy and memset
implementations from a 140 KB module.
Frequently Asked Questions
Is memory.copy safe for overlapping ranges?
Yes. It behaves like memmove, copying as if through a temporary buffer.
Do I need to do anything to benefit? Usually not — recent toolchains enable the feature and lower copies automatically. Check the output once to be sure.
Is memory.fill limited to zero?
No — it fills with any byte value, which makes it useful for initialising buffers to a sentinel pattern while debugging, such as 0xAA
canaries around allocations.
What about tables?
table.copy and table.init are the table equivalents, used by dynamic linking and runtimes that manage function tables; see
initialising tables with element segments.
Does bulk memory work with shared memory? Yes. On shared memory, bulk operations are not atomic as a whole, so synchronise threads around them as you would around any copy.
Related
- Post-MVP Wasm proposals in practice — the family of standardised extensions.
- Reading and writing memory in WAT — the single-access instructions.
- Avoiding copies when passing image buffers — better than fast copies: none.
- Choosing a Rust Wasm target triple — which targets enable bulk memory by default.
← Back to Post-MVP Wasm Proposals in Practice