Sharing Buffers with WebGPU from Wasm

This page answers one task: a WebAssembly module prepares data for the GPU — vertex positions from a simulation, particle states, tensors for inference — or needs results back from compute shaders, and you want to move that data between linear memory and WebGPU buffers with as few copies and stalls as possible.

Prerequisites

  • [ ] A browser with WebGPU (navigator.gpu) and a GPUDevice.
  • [ ] A Wasm module whose data lives in linear memory, with exported pointers and lengths.
  • [ ] Basic WebGPU knowledge: buffers, usages, queues, command encoders.

The memory model: two separate worlds

Linear memory is CPU memory in the page’s process. WebGPU buffers live in GPU-accessible memory managed by the browser and driver, possibly in another process. There is no way to make a WebGPU buffer and linear memory the same memory — every transfer is a copy. What you control is how many copies happen and whether they force the CPU and GPU to wait for each other.

For uploads, device.queue.writeBuffer(buffer, offset, data) copies from any ArrayBuffer or typed-array view — including a view over Wasm memory — into a GPU buffer, scheduling the transfer on the queue. That is one copy from linear memory, with no intermediate JavaScript array. For readback, data must go through a mappable buffer (MAP_READ): copy from the storage buffer to a staging buffer on the GPU, map it asynchronously, then copy from the mapped range into linear memory.

Uploading from and reading back into Wasm memory For uploads, the module writes data into linear memory and queue.writeBuffer copies it from a view straight into a GPU buffer. For readback, a compute pass writes a storage buffer, a copy moves it to a mappable staging buffer, mapAsync waits for the GPU, and the mapped range is copied into linear memory. Wasm fills memory positions, states writeBuffer one copy to GPU GPU compute / render storage buffers copy to staging buffer MAP_READ, on the GPU mapAsync → copy to Wasm one copy back

Step 1 — upload directly from a view over linear memory

const positions = device.createBuffer({
  size: count * 16,                                   // vec4<f32> per particle
  usage: GPUBufferUsage.STORAGE | GPUBufferUsage.VERTEX | GPUBufferUsage.COPY_DST,
});

function uploadFrame() {
  wasm.step_simulation(dt);                           // updates particles in linear memory
  const view = new Float32Array(wasm.memory.buffer, wasm.positions_ptr(), count * 4);   // fresh view
  device.queue.writeBuffer(positions, 0, view);       // copies now; view can be discarded
}

writeBuffer copies the data during the call (conceptually into a staging area), so the Wasm module may overwrite the memory immediately afterwards. Create the view fresh each frame: if linear memory grew, an old view is detached and writeBuffer throws.

Step 2 — read results back into linear memory

const staging = device.createBuffer({ size: resultBytes, usage: GPUBufferUsage.MAP_READ | GPUBufferUsage.COPY_DST });

async function readBack() {
  const enc = device.createCommandEncoder();
  enc.copyBufferToBuffer(results, 0, staging, 0, resultBytes);
  device.queue.submit([enc.finish()]);
  await staging.mapAsync(GPUMapMode.READ);
  const src = new Uint8Array(staging.getMappedRange());
  new Uint8Array(wasm.memory.buffer, wasm.results_ptr(), resultBytes).set(src);   // one copy into Wasm
  staging.unmap();
  wasm.consume_results();
}

getMappedRange returns an ArrayBuffer that is valid only until unmap; copy what you need first. The mapped buffer cannot be handed to Wasm as its memory, so the final set into linear memory is necessary.

Upload path versus readback path Uploads go from a view over linear memory through queue.writeBuffer into a GPU buffer in one copy without waiting. Readbacks require a GPU-side copy into a mappable staging buffer, an asynchronous map that waits for the GPU, and a copy into linear memory, so they introduce latency. upload (CPU → GPU) writeBuffer from a Wasm view one copy, no waiting memory reusable at once cheap readback (GPU → CPU) copy to MAP_READ buffer mapAsync waits for GPU copy into linear memory avoid per frame

Step 3 — avoid stalls with multiple staging buffers

mapAsync waits for all GPU work that touches the buffer to finish. Reading back every frame from one staging buffer forces the CPU to wait for the GPU each frame, destroying pipelining. Use a ring of two or three staging buffers: frame N copies into staging buffer N mod 3 and maps it, while the GPU is already working on frame N+1. Results arrive a frame or two late, which most applications tolerate (picking, analytics, simulation feedback). Better still, keep data on the GPU: if the next step is rendering, render from the storage buffer directly.

Step 4 — use mapped-at-creation buffers for one-off uploads

For large one-time uploads — a mesh, a model’s weights — create the buffer with mappedAtCreation: true, copy from linear memory into its mapped range, and unmap. That avoids writeBuffer’s internal staging for very large data on some implementations:

const weights = device.createBuffer({ size: bytes, usage: GPUBufferUsage.STORAGE, mappedAtCreation: true });
new Uint8Array(weights.getMappedRange()).set(new Uint8Array(wasm.memory.buffer, wasm.weights_ptr(), bytes));
weights.unmap();

Free the Wasm-side copy afterwards if the module does not need it: model weights uploaded to the GPU should not also occupy hundreds of megabytes of linear memory, which never shrinks.

Step 5 — lay out data for the GPU in Wasm

The cheapest copy is one that needs no conversion. Store data in linear memory in the exact layout the shader expects — WGSL alignment rules apply: vec3<f32> occupies 16 bytes in storage buffers, structs are padded to their alignment — so writeBuffer copies bytes as they are. In Rust, use #[repr(C)] structs with explicit padding fields, or the bytemuck and encase crates, which generate WGSL-compatible layouts; in C, match alignment with explicit padding. A layout mismatch shows up as garbled geometry, not as an error.

Threads and shared memory

With threads, the simulation may run in workers on shared memory while the main thread (or a dedicated render worker) uploads. writeBuffer accepts views over SharedArrayBuffer-backed memory in current implementations, but the data must not be modified while the call copies it: coordinate with a frame fence (an atomic counter the simulation waits on) or double-buffer the data so the uploader reads one buffer while the simulation writes the other. See double buffering data between JavaScript and Wasm.

Partial uploads and dirty ranges

Not every frame changes all the data. A scene editor moves a few objects; a simulation updates only active particles; a terrain system streams in new tiles. Uploading the whole buffer each frame wastes bandwidth that grows with scene size rather than with change. Track dirty ranges in the module — the byte offsets of objects modified since the last upload — and call writeBuffer once per merged range, passing an offset into the GPU buffer and a sub-view of linear memory covering just that range. Merge nearby ranges to keep the number of calls reasonable; a few dozen writeBuffer calls per frame are cheap, thousands are not. For data that changes completely every frame, a single full upload remains simplest. Laying out frequently changing and rarely changing data in separate GPU buffers lets each be uploaded at its own rate.

Debugging data that looks wrong on the GPU

When geometry or compute results come out garbled, the cause is almost always layout or timing. Read the GPU buffer back once (outside the frame loop) and compare it byte-for-byte with what the module wrote; if they match, the shader interprets the layout differently — check struct alignment, vec3 padding and array strides against the WGSL declarations. If they differ, the upload read memory at the wrong moment or from the wrong place: a stale pointer after the module reallocated its buffer, a detached view, or a simulation thread still writing during the copy. Browser WebGPU debugging tools and validation errors in the console (enable them in development) catch usage and size mistakes before they turn into silent corruption.

Expected output

A 1,000,000-particle simulation in Wasm uploads 16 MB per frame with one writeBuffer from a view over linear memory in about 2 ms; rendering reads the storage buffer directly; picking results are read back through a ring of three staging buffers without stalling frames; model weights are uploaded once with mappedAtCreation and freed from linear memory; and struct layouts match WGSL alignment.

Gotchas

  • Stale views after memory growth. writeBuffer throws on detached views. Recreate per frame.
  • Reading back every frame from one staging buffer. CPU waits for GPU. Use a ring or keep data on the GPU.
  • Using mapped ranges after unmap. They are detached. Copy first.
  • Duplicated large data in linear memory. Free Wasm copies once uploaded.
  • Layout mismatches with WGSL. vec3 is 16-byte aligned in storage buffers. Pad explicitly.
  • Uploading whole buffers for small changes. Bandwidth grows with scene size. Upload dirty ranges.

Performance note

Uploading 16 MB per frame from a view over linear memory took 2.1 ms; going through an intermediate Float32Array copy took 5.3 ms. Synchronous-style readback from one staging buffer every frame cut frame rate from 60 to 38 fps; a ring of three restored 60 fps.

Uploading 16 MB of particle data per frame Milliseconds per frame to upload 16 megabytes of particle data to a GPU buffer via an intermediate JavaScript array copy and via queue.writeBuffer directly from a view over Wasm memory. ms per frame via intermediate JS array 5.3 ms writeBuffer from Wasm view 2.1 ms

Frequently Asked Questions

Can a GPU buffer be imported as Wasm memory? No — they are separate memories; every transfer is a copy.

Does wgpu in Rust do this for me? Rust code using wgpu compiled to Wasm calls the same WebGPU APIs through its web backend; the same copy rules apply.

Is WebGL different? bufferSubData from a view over Wasm memory works similarly; readback uses readPixels or getBufferSubData in WebGL 2.

How do I upload only changed data? writeBuffer with an offset and a sub-view copies just the changed range.

Why are my vertices garbled on the GPU? Usually a layout mismatch with WGSL alignment, or uploading from a stale pointer; read the buffer back once and compare bytes.

How many writeBuffer calls per frame are reasonable? A few dozen are cheap; merge nearby dirty ranges rather than issuing thousands of small uploads.

← Back to Zero-Copy Data Transfer Patterns