Benchmarking Memory Bandwidth in Wasm

This page answers one question: how many gigabytes per second can WebAssembly code move through linear memory, and is your kernel limited by that bandwidth or by computation?

Prerequisites

  • [ ] A Rust or C toolchain targeting Wasm, with SIMD available (-C target-feature=+simd128 or -msimd128).
  • [ ] A browser or Node with a warmed-up benchmark loop, as in avoiding JIT warm-up errors in Wasm benchmarks.
  • [ ] Buffers large enough to exceed the CPU caches — 64 MB or more — plus small ones that fit in cache.

Why bandwidth is the ceiling for many kernels

Many of the workloads people move into WebAssembly — image filters, audio processing, compression, parsing — touch every byte of a large buffer once or a few times and do modest arithmetic per byte. Such kernels are memory-bound: the CPU spends most of its time waiting for data to arrive from RAM, and making the arithmetic faster changes little. Knowing the memory bandwidth available to a Wasm module tells you the best time such a kernel can possibly achieve, and therefore how much optimization headroom is left.

WebAssembly adds one consideration to the native picture. Every load and store to linear memory must stay within bounds, and engines enforce that either with guard pages — reserving a large virtual region so out-of-bounds accesses fault in hardware, which costs nothing per access — or, where guard pages are unavailable, with explicit bounds checks that cost a compare and branch per access. On 64-bit desktop browsers guard pages are the norm for 32-bit memories, so scalar access runs at close to native speed. Memory64 modules, and some mobile configurations, pay for explicit checks, and that difference is one of the things a bandwidth benchmark reveals.

What limits a streaming kernel, from fastest to slowest The speed a streaming kernel can reach is bounded first by L1 and L2 cache bandwidth for small buffers, then by last-level cache, and finally by main memory bandwidth for large buffers. Bounds-checking strategy and access width decide how close Wasm code gets to each ceiling. L1 / L2 cache ~100+ GB/s per core — buffers up to a few hundred KB last-level cache ~40-80 GB/s — buffers up to tens of MB main memory (DRAM) ~15-30 GB/s on a laptop — large buffers bounds checks free with guard pages; a compare per access without access width 1-byte scalar loads far slower than 16-byte v128 loads

Step 1 — write the kernels

Measure four access patterns: a byte-wise scalar sum, a word-wise scalar sum, a SIMD sum, and the bulk-memory copy instruction. Each reads or writes every byte of the buffer once.

use core::arch::wasm32::*;

#[no_mangle] pub extern "C" fn sum_u8(p: *const u8, n: usize) -> u64 {
    let s = unsafe { core::slice::from_raw_parts(p, n) };
    s.iter().map(|&b| b as u64).sum()
}

#[no_mangle] pub extern "C" fn sum_u64(p: *const u64, n: usize) -> u64 {
    let s = unsafe { core::slice::from_raw_parts(p, n / 8) };
    s.iter().fold(0u64, |a, &w| a.wrapping_add(w))
}

#[no_mangle] pub extern "C" fn sum_v128(p: *const v128, n: usize) -> u64 {
    let mut acc = u64x2_splat(0);
    for i in 0..n / 16 {
        acc = i64x2_add(acc, unsafe { v128_load(p.add(i)) });
    }
    u64x2_extract_lane::<0>(acc).wrapping_add(u64x2_extract_lane::<1>(acc))
}

#[no_mangle] pub extern "C" fn copy(dst: *mut u8, src: *const u8, n: usize) {
    unsafe { core::ptr::copy_nonoverlapping(src, dst, n) };    // lowers to memory.copy with bulk-memory
}

Build with SIMD and bulk memory enabled — both are default-on in recent Rust releases for wasm32, but stating them keeps the build explicit:

RUSTFLAGS="-C target-feature=+simd128,+bulk-memory" \
  cargo build --release --target wasm32-unknown-unknown

Step 2 — measure GB/s across buffer sizes

Allocate buffers in linear memory once, warm up, then time each kernel and convert to bytes per second:

const sizes = [32 << 10, 1 << 20, 16 << 20, 128 << 20];   // 32 KB … 128 MB
for (const n of sizes) {
  const src = exports.alloc(n), dst = exports.alloc(n);
  new Uint8Array(exports.memory.buffer, src, n).fill(7);
  const gbps = (fn, bytesTouched) => {
    for (let i = 0; i < 5; i++) fn();                       // warm
    const reps = Math.max(3, Math.floor((256 << 20) / n));
    const t0 = performance.now();
    for (let i = 0; i < reps; i++) fn();
    return ((bytesTouched * reps) / ((performance.now() - t0) / 1000) / 1e9).toFixed(1);
  };
  console.log(n >> 10, "KB", {
    u8: gbps(() => exports.sum_u8(src, n), n),
    u64: gbps(() => exports.sum_u64(src, n), n),
    v128: gbps(() => exports.sum_v128(src, n), n),
    copy: gbps(() => exports.copy(dst, src, n), 2 * n),   // read + write
  });
}

Count bytes honestly: a copy reads and writes every byte, so it touches twice the buffer size.

Step 3 — compare with the host

The same kernels compiled natively give the ceiling for this machine. A native memcpy and a native SIMD sum are the reference points; WebAssembly should approach them for large buffers, where both are limited by DRAM, and may fall short for small, cache-resident buffers, where code quality matters more.

Read bandwidth for a 128 MB buffer by access pattern Gigabytes per second achieved summing or copying a 128 MB buffer in Chrome on a laptop, compared with a native SIMD sum on the same machine. Wide loads and memory.copy approach the native memory-bound ceiling; byte-wise scalar loops do not. GB/s, 128 MB buffer (higher is better) Wasm scalar u8 sum 3.9 GB/s Wasm scalar u64 sum 14.2 GB/s Wasm v128 sum 18.6 GB/s Wasm memory.copy (r+w) 21.3 GB/s native SIMD sum (reference) 19.8 GB/s

For large buffers the SIMD sum and memory.copy reach the memory-bound ceiling, within a few percent of native. The byte-wise loop is compute-bound — one load, one add and one loop iteration per byte — and runs at a fraction of available bandwidth. That gap, not WebAssembly itself, is what a slow streaming kernel usually suffers from.

Step 4 — classify your kernel

Measure your real kernel’s throughput in the same units — bytes of input processed per second — and place it against these reference lines. If it is close to the memory-bound ceiling for its buffer size, it is memory-bound: further compute optimization will not help, and the remaining options are touching fewer bytes (smaller types, fewer passes over the data, fusing passes) or keeping data in cache (processing in tiles). If it is well below the ceiling, it is compute-bound, and SIMD, better algorithms and fewer branches will pay off — see autovectorizing loops for Wasm SIMD.

Reading a kernel's bandwidth against the ceiling If a kernel's throughput is near the memory-bound ceiling for its buffer size, reduce bytes touched or improve cache locality. If it is far below, the kernel is compute-bound and benefits from SIMD, fewer branches and better algorithms. Kernel throughput compared with the measured ceiling within ~20% of ceiling Memory-bound touch fewer bytes, fuse passes, tile far below ceiling Compute-bound SIMD, fewer branches, better algorithm

Step 5 — mind the JavaScript side of the copy

Bandwidth inside the module is only part of the cost. Data often arrives from JavaScript — a decoded image, a file, a network response — and is copied into linear memory before the kernel runs, then copied out afterwards. Uint8Array.prototype.set on a large buffer runs at memory bandwidth too, so a kernel that processes 100 MB at 18 GB/s but needs a 100 MB copy in and out spends two thirds of its time copying. Measure those copies with the same harness, and avoid them where possible with the patterns in zero-copy data transfer patterns.

Interpreting results across devices

Bandwidth varies more across devices than compute does. A desktop with dual-channel DDR5 reaches 50 GB/s or more; a laptop on battery may halve its memory clock; a mid-range phone may manage 8–12 GB/s. A memory-bound kernel’s speed tracks those numbers directly, so running the bandwidth benchmark on a target phone tells you more about a filter’s real-world speed than any amount of profiling on a desktop. It is also why reducing bytes touched — processing 8-bit data instead of 32-bit floats where the quality allows — is often the most effective optimization on mobile.

Expected output

32 KB   { u8: '4.1', u64: '29.8', v128: '61.2', copy: '74.5' }
1024 KB { u8: '4.0', u64: '24.6', v128: '42.0', copy: '48.3' }
16 MB   { u8: '3.9', u64: '16.1', v128: '22.4', copy: '25.0' }
128 MB  { u8: '3.9', u64: '14.2', v128: '18.6', copy: '21.3' }

Gotchas

  • The compiler vectorized the “scalar” loop. LLVM may autovectorize sum_u8 when SIMD is enabled. Check the WAT for v128 instructions, or build that kernel without SIMD to get a true scalar reference.
  • Memory growth during the benchmark. Allocating the buffers inside the timed loop measures growth and allocation. Allocate first.
  • Timer resolution on small buffers. A 32 KB sum takes microseconds; time many repetitions and divide.
  • Huge pages and allocation patterns. The host may back a large fresh allocation with huge pages, which changes bandwidth slightly between runs. Run each size several times and report the median.
  • Comparing against an unwarmed native build. Native code also has caches to warm. Run both the same way.

Performance note

On a phone the same benchmark reached 9.4 GB/s for the SIMD sum on a 128 MB buffer and 2.1 GB/s for the byte-wise loop. A memory-bound image filter measured at 8.7 GB/s on that phone was already at 93% of the ceiling — a clear sign that further work belonged in reducing passes over the image, not in the arithmetic.

Frequently Asked Questions

Is memory.copy always faster than a loop? For more than a few dozen bytes, yes in practice: engines implement it with optimized native copy routines. For tiny copies, a few scalar moves can win because the instruction has fixed overhead.

Does Memory64 change bandwidth? It can, because 64-bit memories usually need explicit bounds checks. Measure the same kernels with a wasm64 build if you plan to use large memories; see Memory64 and large heaps.

Why is the 32 KB copy faster than the theoretical DRAM speed? Because the whole buffer fits in the CPU’s caches after the first pass, so later repetitions never touch main memory. That is the L1/L2 ceiling from the diagram above, and it is why small and large buffers must be measured separately.

Do threads increase bandwidth? Up to a point. One core rarely saturates a desktop’s memory bus, so two to four threads streaming different regions can raise aggregate throughput, until the bus is saturated.

How do I see cache effects in a profile? Browser profilers do not expose cache misses. On Linux, run the module in wasmtime under perf stat to see cache-miss counters; see profiling Wasm hot paths with perf.

← Back to Wasm Performance Benchmarking