Reading fetch Responses into Wasm Memory

This page answers one task: a WebAssembly module needs data from the network — a model file, a map tile set, a large dataset — and the bytes should land in linear memory as they arrive, without first being assembled into a separate ArrayBuffer and then copied in.

Prerequisites

  • [ ] A module that can allocate a region for the data (an alloc(len) export) or accept data incrementally.
  • [ ] A server response, ideally with a Content-Length header.

The double-buffer problem

The usual pattern — const buf = await (await fetch(url)).arrayBuffer() and then copy into the module — has two costs. Peak memory is twice the response size, because the full response exists in JavaScript and again in linear memory. And the module cannot start work until the last byte has arrived and been copied, so download and processing happen strictly one after the other.

A response body is a ReadableStream of chunks. Reading it directly lets each chunk be written into linear memory as soon as it arrives, after which the chunk can be garbage-collected. With a known length, the module can allocate the final region up front and the chunks fill it in order; without one, the module can consume chunks incrementally with a streaming decoder. Either way, memory holds roughly one copy of the data, and processing can overlap with the download.

Buffering a response versus streaming it into Wasm arrayBuffer then copy holds the full response twice and starts processing only after the download. Streaming writes chunks into a pre-allocated region of linear memory as they arrive, holding one copy and allowing work to begin early. arrayBuffer() then copy full response in JS memory copied into linear memory work starts after download simple, double peak stream into Wasm chunks written as they arrive one copy, in linear memory work can overlap download large responses

Step 1 — allocate using Content-Length

async function fetchIntoWasm(url, wasm) {
  const res = await fetch(url);
  if (!res.ok) throw new Error(`HTTP ${res.status} for ${url}`);
  const len = Number(res.headers.get("Content-Length"));
  if (!len || res.headers.get("Content-Encoding")) return fetchUnknownLength(res, wasm);   // see step 3

  const ptr = wasm.alloc(len);                 // may grow memory: allocate before creating views
  let offset = 0;
  const reader = res.body.getReader();
  for (;;) {
    const { value, done } = await reader.read();
    if (done) break;
    if (offset + value.length > len) throw new Error("response longer than Content-Length");
    new Uint8Array(wasm.memory.buffer, ptr + offset, value.length).set(value);
    offset += value.length;
  }
  if (offset !== len) throw new Error(`short read: ${offset} of ${len}`);
  return { ptr, len };
}

The allocation happens once, before any view exists, so growth during allocation cannot detach a view in use. Each chunk is written at its offset with a fresh view. Validating the final length catches truncated downloads, which otherwise appear as corrupt data deep inside the module.

Step 2 — mind Content-Encoding

Content-Length describes the bytes on the wire. If the response is compressed — Content-Encoding: gzip or br — the browser decompresses it transparently, and the body stream yields more bytes than Content-Length. Do not size the allocation from it in that case. Options: serve the file uncompressed with an accurate length (fine for already-compressed formats such as images or quantised model weights), send the decompressed size in a custom header such as X-Decoded-Length, or use the unknown-length path below. Cache headers and CDN behaviour for binary assets are covered in setting cache-control headers for Wasm.

Step 3 — handle unknown lengths

When the length is unknown, either grow a buffer in the module as chunks arrive, or feed chunks to a streaming consumer that does not need the whole input:

async function fetchUnknownLength(res, wasm) {
  const reader = res.body.getReader();
  for (;;) {
    const { value, done } = await reader.read();
    if (done) break;
    const ptr = wasm.input_reserve(value.length);    // module grows its own buffer, returns write position
    new Uint8Array(wasm.memory.buffer, ptr, value.length).set(value);
    wasm.input_commit(value.length);
  }
  return wasm.input_finish();
}

The module’s input_reserve grows its internal vector geometrically, as a Vec would; input_commit records the written bytes. For decoders — image codecs, decompressors, parsers — prefer a true streaming interface that consumes each chunk and keeps only its state, which keeps memory flat, as in streaming file uploads into Wasm memory.

Streaming a response body into linear memory JavaScript starts the fetch and allocates a region of the content length in the module. As each chunk of the body arrives, JavaScript writes it at the next offset in that region. When the stream ends, JavaScript checks the length and tells the module the data is ready. network JavaScript Wasm module alloc(Content-Length) → ptr chunk 1 (64 KB) write at ptr + 0 chunk 2 … n load_model(ptr, len)

Step 4 — overlap processing with the download

When the module can work on partial data — parse records as they complete, decode tiles, build an index incrementally — call it after each chunk rather than at the end. Download and processing then proceed together, and the total time approaches whichever is slower instead of their sum. Report progress from bytes received over Content-Length. For data that must be complete before use, such as model weights, the overlap is limited to the copy itself, which is still worthwhile because it removes the separate copy at the end.

Step 5 — run it where the module lives

If the module runs in a worker, do the fetch in the worker too — fetch is available in workers — so the bytes never touch the main thread. Fetching on the main thread and transferring to a worker works but adds a hop. In Node, fetch returns the same stream type, and the code is identical; for local files, fs.createReadStream gives chunks the same way.

Range requests for partial loading

Not every use needs the whole file. Datasets, archives and map tile packages often have an index at a known position — a header at the start, a directory at the end — that tells the module which byte ranges it actually needs. HTTP range requests fetch just those ranges: fetch(url, { headers: { Range: "bytes=1048576-2097151" } }) returns a 206 Partial Content response whose body streams into linear memory exactly as above. A module that reads a SQLite database or a cloud-optimised GeoTIFF this way can answer a query by fetching a few hundred kilobytes out of a multi-gigabyte file. The JavaScript side becomes a small block-fetching layer: the module asks for a range, JavaScript fetches it into a region the module allocated, and the module continues. Batch adjacent ranges into one request, cache fetched blocks, and check that the server returns 206 rather than 200 — a server that ignores the Range header sends the whole file, which the length check will catch.

Caching large responses

Large binary downloads deserve caching beyond the HTTP cache, which browsers may evict for big entries or ignore for range requests. The Cache API stores full responses and returns them as Response objects whose bodies stream exactly like network responses, so the reading code above works unchanged for cached data:

const cache = await caches.open("model-v3");
let res = await cache.match(url);
if (!res) { res = await fetch(url); await cache.put(url, res.clone()); }
const { ptr, len } = await fetchIntoWasmFromResponse(res, wasm);

Version the cache name with the asset so updates do not mix old and new data, and check the Content-Length of cached responses the same way. For very large assets — hundreds of megabytes of model weights — the Origin Private File System offers faster random access and is easier to manage in pieces, at the cost of more code. Either way, the second visit reads from local storage at disk speed instead of network speed, and the streaming copy into linear memory remains the only copy.

Expected output

Loading a 180 MB model streams it into linear memory with peak tab memory about 190 MB instead of about 370 MB; the module receives a pointer and length that match Content-Length; and a truncated download is reported as a short read instead of a crash inside the model loader.

Gotchas

  • Sizing from Content-Length on compressed responses. The decoded body is larger. Use a decoded-length header or the unknown-length path.
  • Creating views before allocating. Allocation may grow memory and detach them. Allocate first.
  • No length validation. Truncated downloads look like corrupt data. Check the final byte count.
  • Missing Content-Length with CORS. The header must be exposed by the server for cross-origin reads.
  • Fetching on the main thread for a worker module. Fetch inside the worker instead.

Performance note

For a 180 MB uncompressed model over a fast local network, arrayBuffer() plus a copy peaked at 372 MB above baseline and took 1.9 s; streaming into a pre-allocated region peaked at 188 MB and took 1.7 s, the saving coming from the removed final copy.

Peak memory loading a 180 MB file into Wasm Increase in tab memory while loading a 180-megabyte file into linear memory, using arrayBuffer plus a copy, and streaming chunks into a pre-allocated region. MB peak memory increase arrayBuffer() + copy 372 MB stream into allocated region 188 MB

Frequently Asked Questions

Can fetch write directly into a buffer I provide? BYOB reads on response bodies can fill a provided buffer, but not Wasm memory itself, because the buffer would be detached. Copy each chunk instead.

What chunk sizes does the stream deliver? Browsers choose them, typically tens to hundreds of kilobytes. The code must handle any size.

Does WebAssembly.instantiateStreaming use this approach? It streams code into the compiler. Data still needs the technique on this page.

How do I cancel a long download? Pass an AbortSignal to fetch, and free the allocated region when the read loop throws.

Is this worth it for small responses? Below a few megabytes, arrayBuffer() plus one copy is simpler and the difference is negligible.

← Back to Zero-Copy Data Transfer Patterns