Aligning Data in Linear Memory

This page answers one task: data structures in WebAssembly memory are read from both the module and JavaScript, and you need a layout where every field is properly aligned — so that typed arrays can view it, accesses are fast, and the two sides agree on every offset.

Prerequisites

What alignment means in WebAssembly

A value is aligned when its address is a multiple of its size: a 4-byte i32 at an address divisible by 4, an 8-byte f64 at a multiple of 8. WebAssembly itself is lenient about this. Every load and store instruction carries an alignment hint — i32.load align=4 — but the hint is a promise to the engine, not a requirement: a misaligned access still produces the correct result. The specification says only that a wrong hint must not change behaviour. On x86 and modern ARM, misaligned accesses usually cost little or nothing extra, so a misaligned field rarely shows up in profiles.

Two things make alignment matter anyway. First, JavaScript typed arrays require it: new Float64Array(buffer, offset, n) throws a RangeError if offset is not a multiple of 8. If the module places an f64 array at an odd address, JavaScript cannot view it without copying. Second, atomic operations require it: an atomic access to a misaligned address traps in WebAssembly, and Atomics on a misaligned typed array cannot be constructed at all. And on some ARM devices, misaligned SIMD loads are measurably slower. So aligned layouts are the default worth keeping.

The same struct with natural and packed layout With natural alignment, a one-byte flag is followed by three bytes of padding so the next four-byte field starts at offset 4, and the eight-byte field starts at offset 8. Packed, the fields sit at offsets 0, 1 and 5, which typed arrays cannot view. struct { u8 flag; u32 id; f64 value } — natural layout, 16 bytes flag pad id (offset 4) value (offset 8) 0 4 8 16

Step 1 — use natural alignment in the module

Compilers already align data naturally. A Rust #[repr(C)] struct and a C struct get padding inserted so each field is aligned to its size, and the struct’s own alignment is that of its largest field:

#[repr(C)]
pub struct Particle {
    pub flags: u8,       // offset 0
    // 3 bytes padding
    pub id: u32,         // offset 4
    pub x: f64,          // offset 8
    pub y: f64,          // offset 16
}                         // size 24, align 8

Avoid #[repr(packed)] and __attribute__((packed)) for data JavaScript will read. Packing removes padding but misaligns fields, which breaks typed-array views and makes taking a reference to a packed field undefined behaviour in Rust. Plain Rust structs without #[repr(C)] have unspecified field order, so any struct shared with JavaScript must be #[repr(C)].

Step 2 — order fields from largest to smallest

Padding wastes memory. Ordering fields by decreasing alignment removes most of it:

#[repr(C)]
pub struct ParticleCompact {
    pub x: f64,          // 0
    pub y: f64,          // 8
    pub id: u32,         // 16
    pub flags: u8,       // 20
    // 3 bytes tail padding
}                         // size 24 → but with more u8 fields, they fill the tail for free

For a single struct the saving is small, but for arrays of millions of elements, a struct that shrinks from 32 to 24 bytes saves a quarter of the memory and memory bandwidth. Check sizes with std::mem::size_of or static_assert(sizeof(T) == 24) so a change that adds padding is caught at compile time.

Step 3 — mirror the layout in JavaScript

JavaScript has no structs, so it reads fields through typed arrays or a DataView at fixed offsets. Write those offsets down once, next to an assertion on the Rust side:

const _: () = assert!(std::mem::size_of::<Particle>() == 24);
const _: () = assert!(std::mem::offset_of!(Particle, x) == 8);
const PARTICLE = { size: 24, flags: 0, id: 4, x: 8, y: 16 };

function readParticle(memory, ptr) {
  const dv = new DataView(memory.buffer);
  return {
    flags: dv.getUint8(ptr + PARTICLE.flags),
    id: dv.getUint32(ptr + PARTICLE.id, true),          // little-endian, like Wasm
    x: dv.getFloat64(ptr + PARTICLE.x, true),
    y: dv.getFloat64(ptr + PARTICLE.y, true),
  };
}

DataView tolerates any alignment and is convenient for single records. For bulk access, typed arrays are faster — and that is where alignment becomes mandatory.

Step 4 — keep array base addresses aligned for typed arrays

To view all x coordinates of an array of particles as a Float64Array, the array’s base must be 8-byte aligned and the stride must be a multiple of 8. Rust’s Vec<Particle> guarantees the base alignment because the allocator honours the type’s alignment. Then:

// ptr and len come from the module: a Vec<Particle> exported as (ptr, len)
const f64 = new Float64Array(memory.buffer, ptr, (len * PARTICLE.size) / 8);
const xs = (i) => f64[i * 3 + 1];             // 24-byte stride = 3 f64 slots; x is slot 1

If the base were not aligned — data placed by hand at an arbitrary offset, or a byte buffer whose contents are reinterpreted — the constructor throws RangeError: start offset of Float64Array should be a multiple of 8. Fix the producer, not the consumer: allocate with the right alignment, using Layout::from_size_align(n, 8) in Rust or aligned_alloc(8, n) in C.

Choosing how JavaScript reads a field If the data is a single record or its alignment is unknown, use a DataView, which accepts any offset. If it is a bulk array whose base and stride are multiples of the element size, use a typed array for speed. If the base is misaligned, fix the allocation in the module rather than copying. record in memory struct at ptr single record? DataView at offsets bulk array? check base and stride aligned Float64Array view misaligned fix the allocation

Step 5 — be strict for atomics and SIMD

For shared memory, every atomic field must be naturally aligned, or the access traps. Place counters and locks in #[repr(C)] structs with AtomicU32/AtomicI64 fields, or align them explicitly with #[repr(align(8))]. For SIMD, align buffers processed with v128.load to 16 bytes when possible — not required for correctness, but it avoids split loads on some hardware. #[repr(align(16))] on a wrapper struct, or an allocation with 16-byte alignment, achieves it.

Endianness, the other half of layout agreement

Alignment decides where fields go; endianness decides how their bytes are ordered. WebAssembly is always little-endian, on every platform. JavaScript typed arrays use the platform’s native byte order, which on every device that runs browsers today is also little-endian, so typed-array views over Wasm memory read values correctly. DataView methods, however, default to big-endian unless the last argument is true — a frequent source of garbage values when code switches from a typed array to a DataView. Always pass true for little-endian when reading Wasm memory through a DataView. If data also travels over the network or to disk in a big-endian format, convert it at that boundary, not in the shared memory layout. Keeping the in-memory layout little-endian and naturally aligned means both sides can read it with their fastest access methods.

Structure of arrays as an alternative layout

When JavaScript reads one field across many records — every x, every id — an array of structs forces a strided view, and padding inside each struct is carried along in every cache line. A structure of arrays stores each field in its own contiguous array instead: one Vec<f64> for all x values, another for all y, a Vec<u32> for ids, a Vec<u8> for flags. Each array is naturally aligned for its element type with no padding at all, JavaScript views each one with a plain typed array at stride 1, and SIMD code in the module can process four or eight values per instruction without gathering them. The module exports each array’s pointer and length, and JavaScript builds one view per field. The trade-off is that reading one whole record touches several arrays, so the layout suits column-wise processing — physics updates, rendering, statistics — better than record-at-a-time access. Many high-throughput modules use both: structure of arrays for hot numeric data, #[repr(C)] structs for small configuration records.

Expected output

size_of::<Particle>() is 24 and the assertions compile; JavaScript constructs a Float64Array over 100,000 particles without a RangeError, and reading x through the typed array agrees with the value the module reports for every particle.

Gotchas

  • #[repr(packed)] on shared structs. Fields become misaligned and typed arrays cannot view them. Use natural alignment.
  • Missing #[repr(C)]. Rust may reorder fields. Shared structs need a defined layout.
  • DataView without the little-endian flag. It defaults to big-endian. Pass true.
  • Typed-array offset not a multiple of the element size. The constructor throws. Align the allocation.
  • Misaligned atomics. They trap. Align atomic fields to their size.

Performance note

Reading the x field of 1 million particles through a Float64Array took 1.1 ms in Chrome, against 4.6 ms through DataView.getFloat64 and 9.8 ms when a misaligned layout forced copying the data to an aligned buffer first. Inside the module, misaligned scalar loads showed no measurable difference on x86 and about 3% on an older ARM phone.

Reading one field from a million records in JavaScript Time to sum the x field of one million 24-byte records from JavaScript using a typed-array view, a DataView, and a copy into an aligned buffer forced by a misaligned layout. ms to read 1M fields Float64Array view (aligned) 1.1 ms DataView.getFloat64 4.6 ms copy to aligned buffer first 9.8 ms

Frequently Asked Questions

Does WebAssembly trap on misaligned loads? No, ordinary loads and stores work at any address. Only atomic accesses require alignment.

What does the alignment hint do then? It lets the engine choose faster instructions when it trusts the hint. A wrong hint is still correct, possibly slower.

How does wasm-bindgen handle alignment? Its generated glue reads values at aligned offsets and allocates buffers with the right alignment for the element type.

Is C++ alignas supported? Yes — clang honours alignas for Wasm targets, and the allocator returns suitably aligned memory up to 16 bytes by default.

← Back to Linear Memory Management & Allocators