Decoding Strings Directly from Wasm Memory
This page answers one task: a WebAssembly module produces text — log lines, parsed fields, rendered markup, JSON — as UTF-8 bytes in linear memory, and JavaScript needs those bytes as strings quickly, without extra copies or per-call setup.
Prerequisites
- [ ] A module that returns a pointer and a byte length for each string it produces.
- [ ] Familiarity with encoding strings across the Wasm boundary.
Strings are always a conversion
JavaScript strings are sequences of UTF-16 code units managed by the engine. A Wasm module’s strings are bytes in linear memory, almost always UTF-8. No view can make one look like the other: getting a JavaScript string means decoding the bytes into a new string object. “Zero-copy” for strings therefore means something narrower — decode straight from linear memory, without first copying the bytes into another buffer, and without paying setup costs per call.
TextDecoder is the tool. Its decode method accepts a typed array, including a view — a subarray — over Wasm memory, and decodes those bytes
directly. Engines implement it natively with fast paths for ASCII, so it is much faster than decoding in JavaScript or calling String.fromCharCode per
byte for anything but tiny strings.
Step 1 — decode a subarray with a cached decoder
const decoder = new TextDecoder("utf-8", { fatal: false }); // create once, reuse forever
function readString(memory, ptr, len) {
return decoder.decode(new Uint8Array(memory.buffer, ptr, len));
}
Creating a TextDecoder is comparatively expensive; creating one per call is a common hidden cost. A module-level decoder reused for every call removes
it. Creating the Uint8Array view over the existing buffer is cheap and does not copy. fatal: false replaces invalid UTF-8 with U+FFFD; use fatal: true to throw on invalid input when the module is supposed to emit valid UTF-8 and you want to catch bugs.
Step 2 — handle shared memory
TextDecoder.decode rejects views over a SharedArrayBuffer in some browsers (it may change while being decoded). Threaded modules with shared memory
therefore need a copy into a non-shared buffer first:
function readStringShared(memory, ptr, len) {
const bytes = new Uint8Array(memory.buffer, ptr, len);
return decoder.decode(memory.buffer instanceof SharedArrayBuffer ? bytes.slice() : bytes);
}
The slice() copies the bytes into an ordinary ArrayBuffer. For short strings the copy is negligible; for large text, reuse a scratch buffer and copy
into it with set, decoding a subarray of the scratch buffer. wasm-bindgen’s glue applies the same workaround in threaded builds.
Step 3 — add an ASCII fast path for tiny strings
For very short strings — single identifiers, field names, short tokens — the fixed overhead of calling decode can exceed the work. A JavaScript loop is
faster below roughly 16–32 bytes when the text is ASCII:
function readShortAscii(u8, ptr, len) {
let s = "";
for (let i = 0; i < len; i++) {
const c = u8[ptr + i];
if (c > 0x7f) return decoder.decode(u8.subarray(ptr, ptr + len)); // fall back for non-ASCII
s += String.fromCharCode(c);
}
return s;
}
Measure before adopting this; the crossover point differs between engines, and modern TextDecoder implementations have narrowed the gap considerably.
Step 4 — decode many strings in one pass
Calling into the module and decoding once per string is wasteful for results containing thousands of strings — the rows of a query, the tokens of a document. Have the module lay out all strings in one contiguous UTF-8 blob with an offsets table, and decode lazily:
// module returns: blobPtr, blobLen, offsetsPtr (Uint32Array of count+1 offsets), count
function stringTable(memory, blobPtr, offsetsPtr, count) {
const offsets = new Uint32Array(memory.buffer, offsetsPtr, count + 1).slice(); // copy offsets once
const blob = new Uint8Array(memory.buffer, blobPtr, offsets[count]).slice(); // copy blob once
const cache = new Array(count);
return (i) => cache[i] ??= decoder.decode(blob.subarray(offsets[i], offsets[i + 1]));
}
Copying the blob out once makes the table independent of Wasm memory, so it survives later calls and memory growth. Strings are decoded only when accessed and cached afterwards, which pays off when the UI shows a page of results out of thousands. The layout is the flat, column-oriented approach from choosing between JSON and binary serialization.
Step 5 — know the cost of the other direction
Writing strings into Wasm memory has its own fast path: TextEncoder.encodeInto(str, view) encodes directly into a view over linear memory, without an
intermediate Uint8Array. Allocate str.length * 3 bytes for the worst case (each UTF-16 code unit becomes at most 3 UTF-8 bytes), encode, and use the
written count it returns as the length. wasm-bindgen’s passStringToWasm0 uses exactly this, with an ASCII loop first.
Validating text from untrusted modules
When the module is trusted, fatal: false and silent replacement of invalid bytes are fine. When the bytes come from user content passed through the
module — a parser for uploaded documents, a plugin written by a third party — decide deliberately what invalid UTF-8 means. A decoder with fatal: true
throws a TypeError on the first invalid sequence, which turns corrupt input into a clear error at the boundary. Replacement characters can hide data
loss and make round trips lossy: a file that is decoded, edited and re-encoded will not match the original bytes. Lengths deserve the same care: a
malicious or buggy module can return a pointer and length that run past the end of memory, and the typed-array constructor will throw — good — but a
length that is merely large can make JavaScript decode megabytes of unrelated memory into a string. Check returned lengths against a sensible maximum
before decoding, and treat out-of-range values as a module bug to report rather than data to display.
When strings should not cross at all
The fastest string conversion is the one that never happens. Many applications decode strings only to compare them, look them up or count them — work
the module could do itself. Search results can be returned as ids and scores, with titles decoded only for the rows on screen. Log levels or categories
can be numeric codes mapped to strings in JavaScript once. Templated output — HTML rendered by a Wasm templating engine, for instance — is better
returned as one large string than as many fragments, because one large decode costs much less than many small ones. And when the destination is not
JavaScript at all — writing to a file, sending over the network — the bytes can stay bytes: pass the Uint8Array straight to fetch, a
WritableStream or a Blob without ever creating a JavaScript string. Reserve decoding for text a person will read or JavaScript code must
manipulate.
Expected output
Decoding 10,000 strings from a single blob with a cached decoder takes about 1.2 ms in Chrome instead of 9 ms with a decoder created per call; strings from a threaded build decode correctly after the shared-memory copy; and lazily decoded table rows render as soon as they scroll into view.
Gotchas
- Creating a
TextDecoderper call. It is the largest avoidable cost. Cache one. - Decoding views over shared memory. Some browsers throw. Copy to a non-shared buffer first.
- Length in characters instead of bytes. UTF-8 lengths are byte counts. Have the module return byte lengths.
- Keeping views across calls. Memory growth detaches them. Copy blobs out if they must outlive the call.
- Decoding unbounded lengths. A bad length decodes unrelated memory. Check it against a maximum first.
- Assuming ASCII. A fast path must fall back for bytes above 0x7F.
Performance note
For 10,000 strings averaging 24 bytes, decoding with a new TextDecoder per call took 9.1 ms in Chrome, a cached decoder per string 3.4 ms, and one blob with
lazy cached decoding of all strings 1.2 ms. Decoding a single 1 MB string took 0.6 ms with TextDecoder.
Frequently Asked Questions
Can JavaScript strings point into Wasm memory without decoding? No. Engines own string storage. Proposals for JS string builtins let Wasm create and inspect JavaScript strings, but linear-memory bytes still need decoding.
Is UTF-16 output from the module faster to decode?
Decoding UTF-16 with TextDecoder("utf-16le") is fast too, but doubles the size of ASCII text. UTF-8 is usually the better trade.
What about String.fromCharCode.apply on a view?
It works for ASCII but has argument-count limits and is slower than TextDecoder for long strings.
Does TextDecoder handle a BOM?
By default it strips a leading UTF-8 BOM; pass ignoreBOM: true to keep it.
Should the module NUL-terminate strings? Not for JavaScript’s sake. Return explicit byte lengths; scanning for a terminator in JavaScript is slower and fails on embedded NULs.
Related
- Reading Wasm linear memory with typed arrays — the views being decoded.
- Creating views into Wasm memory safely — view lifetime rules.
- Reading the glue code wasm-bindgen generates — how the glue decodes strings.
- Using Wasm GC for managed languages — where strings live outside linear memory.
← Back to Zero-Copy Data Transfer Patterns