Reading fetch Responses into Wasm Memory
This page answers one task: a WebAssembly module needs data from the network — a model file, a map tile set, a large dataset — and the bytes should land
in linear memory as they arrive, without first being assembled into a separate ArrayBuffer and then copied in.
Prerequisites
- [ ] A module that can allocate a region for the data (an
alloc(len)export) or accept data incrementally. - [ ] A server response, ideally with a
Content-Lengthheader.
The double-buffer problem
The usual pattern — const buf = await (await fetch(url)).arrayBuffer() and then copy into the module — has two costs. Peak memory is twice the
response size, because the full response exists in JavaScript and again in linear memory. And the module cannot start work until the last byte has
arrived and been copied, so download and processing happen strictly one after the other.
A response body is a ReadableStream of chunks. Reading it directly lets each chunk be written into linear memory as soon as it arrives, after which the
chunk can be garbage-collected. With a known length, the module can allocate the final region up front and the chunks fill it in order; without one, the
module can consume chunks incrementally with a streaming decoder. Either way, memory holds roughly one copy of the data, and processing can overlap with
the download.
Step 1 — allocate using Content-Length
async function fetchIntoWasm(url, wasm) {
const res = await fetch(url);
if (!res.ok) throw new Error(`HTTP ${res.status} for ${url}`);
const len = Number(res.headers.get("Content-Length"));
if (!len || res.headers.get("Content-Encoding")) return fetchUnknownLength(res, wasm); // see step 3
const ptr = wasm.alloc(len); // may grow memory: allocate before creating views
let offset = 0;
const reader = res.body.getReader();
for (;;) {
const { value, done } = await reader.read();
if (done) break;
if (offset + value.length > len) throw new Error("response longer than Content-Length");
new Uint8Array(wasm.memory.buffer, ptr + offset, value.length).set(value);
offset += value.length;
}
if (offset !== len) throw new Error(`short read: ${offset} of ${len}`);
return { ptr, len };
}
The allocation happens once, before any view exists, so growth during allocation cannot detach a view in use. Each chunk is written at its offset with a fresh view. Validating the final length catches truncated downloads, which otherwise appear as corrupt data deep inside the module.
Step 2 — mind Content-Encoding
Content-Length describes the bytes on the wire. If the response is compressed — Content-Encoding: gzip or br — the browser decompresses it
transparently, and the body stream yields more bytes than Content-Length. Do not size the allocation from it in that case. Options: serve the file
uncompressed with an accurate length (fine for already-compressed formats such as images or quantised model weights), send the decompressed size in a
custom header such as X-Decoded-Length, or use the unknown-length path below. Cache headers and CDN behaviour for binary assets are covered in
setting cache-control headers for Wasm.
Step 3 — handle unknown lengths
When the length is unknown, either grow a buffer in the module as chunks arrive, or feed chunks to a streaming consumer that does not need the whole input:
async function fetchUnknownLength(res, wasm) {
const reader = res.body.getReader();
for (;;) {
const { value, done } = await reader.read();
if (done) break;
const ptr = wasm.input_reserve(value.length); // module grows its own buffer, returns write position
new Uint8Array(wasm.memory.buffer, ptr, value.length).set(value);
wasm.input_commit(value.length);
}
return wasm.input_finish();
}
The module’s input_reserve grows its internal vector geometrically, as a Vec would; input_commit records the written bytes. For decoders — image
codecs, decompressors, parsers — prefer a true streaming interface that consumes each chunk and keeps only its state, which keeps memory flat, as in
streaming file uploads into Wasm memory.
Step 4 — overlap processing with the download
When the module can work on partial data — parse records as they complete, decode tiles, build an index incrementally — call it after each chunk rather
than at the end. Download and processing then proceed together, and the total time approaches whichever is slower instead of their sum. Report progress
from bytes received over Content-Length. For data that must be complete before use, such as model weights, the overlap is limited to the copy itself,
which is still worthwhile because it removes the separate copy at the end.
Step 5 — run it where the module lives
If the module runs in a worker, do the fetch in the worker too — fetch is available in workers — so the bytes never touch the main thread. Fetching on
the main thread and transferring to a worker works but adds a hop. In Node, fetch returns the same stream type, and the code is identical; for local
files, fs.createReadStream gives chunks the same way.
Range requests for partial loading
Not every use needs the whole file. Datasets, archives and map tile packages often have an index at a known position — a header at the start, a
directory at the end — that tells the module which byte ranges it actually needs. HTTP range requests fetch just those ranges:
fetch(url, { headers: { Range: "bytes=1048576-2097151" } }) returns a 206 Partial Content response whose body streams into linear memory exactly as
above. A module that reads a SQLite database or a cloud-optimised GeoTIFF this way can answer a query by fetching a few hundred kilobytes out of a
multi-gigabyte file. The JavaScript side becomes a small block-fetching layer: the module asks for a range, JavaScript fetches it into a region the
module allocated, and the module continues. Batch adjacent ranges into one request, cache fetched blocks, and check that the server returns 206 rather
than 200 — a server that ignores the Range header sends the whole file, which the length check will catch.
Caching large responses
Large binary downloads deserve caching beyond the HTTP cache, which browsers may evict for big entries or ignore for range requests. The Cache API
stores full responses and returns them as Response objects whose bodies stream exactly like network responses, so the reading code above works
unchanged for cached data:
const cache = await caches.open("model-v3");
let res = await cache.match(url);
if (!res) { res = await fetch(url); await cache.put(url, res.clone()); }
const { ptr, len } = await fetchIntoWasmFromResponse(res, wasm);
Version the cache name with the asset so updates do not mix old and new data, and check the Content-Length of cached responses the same way. For very
large assets — hundreds of megabytes of model weights — the Origin Private File System offers faster random access and is easier to manage in
pieces, at the cost of more code. Either way, the second visit reads from local storage at disk speed instead of network speed, and the streaming copy
into linear memory remains the only copy.
Expected output
Loading a 180 MB model streams it into linear memory with peak tab memory about 190 MB instead of about 370 MB; the module receives a pointer and length
that match Content-Length; and a truncated download is reported as a short read instead of a crash inside the model loader.
Gotchas
- Sizing from
Content-Lengthon compressed responses. The decoded body is larger. Use a decoded-length header or the unknown-length path. - Creating views before allocating. Allocation may grow memory and detach them. Allocate first.
- No length validation. Truncated downloads look like corrupt data. Check the final byte count.
- Missing
Content-Lengthwith CORS. The header must be exposed by the server for cross-origin reads. - Fetching on the main thread for a worker module. Fetch inside the worker instead.
Performance note
For a 180 MB uncompressed model over a fast local network, arrayBuffer() plus a copy peaked at 372 MB above baseline and took 1.9 s; streaming into a
pre-allocated region peaked at 188 MB and took 1.7 s, the saving coming from the removed final copy.
Frequently Asked Questions
Can fetch write directly into a buffer I provide?
BYOB reads on response bodies can fill a provided buffer, but not Wasm memory itself, because the buffer would be detached. Copy each chunk instead.
What chunk sizes does the stream deliver? Browsers choose them, typically tens to hundreds of kilobytes. The code must handle any size.
Does WebAssembly.instantiateStreaming use this approach?
It streams code into the compiler. Data still needs the technique on this page.
How do I cancel a long download?
Pass an AbortSignal to fetch, and free the allocated region when the read loop throws.
Is this worth it for small responses?
Below a few megabytes, arrayBuffer() plus one copy is simpler and the difference is negligible.
Related
- Streaming data into Wasm with ReadableStream — the stream plumbing.
- Loading large model weights into linear memory — loading large weights.
- Growing memory safely from JavaScript — reserving space before loading.
- Caching compiled Wasm modules in IndexedDB — caching the code side.
← Back to Zero-Copy Data Transfer Patterns