Decoding Strings Directly from Wasm Memory

This page answers one task: a WebAssembly module produces text — log lines, parsed fields, rendered markup, JSON — as UTF-8 bytes in linear memory, and JavaScript needs those bytes as strings quickly, without extra copies or per-call setup.

Prerequisites

Strings are always a conversion

JavaScript strings are sequences of UTF-16 code units managed by the engine. A Wasm module’s strings are bytes in linear memory, almost always UTF-8. No view can make one look like the other: getting a JavaScript string means decoding the bytes into a new string object. “Zero-copy” for strings therefore means something narrower — decode straight from linear memory, without first copying the bytes into another buffer, and without paying setup costs per call.

TextDecoder is the tool. Its decode method accepts a typed array, including a view — a subarray — over Wasm memory, and decodes those bytes directly. Engines implement it natively with fast paths for ASCII, so it is much faster than decoding in JavaScript or calling String.fromCharCode per byte for anything but tiny strings.

Decoding a string from linear memory The module returns a pointer and length. JavaScript creates a subarray view over Wasm memory at that range without copying, and passes it to a cached TextDecoder, which decodes the UTF-8 bytes into a new JavaScript string. module returns ptr, len subarray(ptr, ptr+len) view, no copy cached TextDecoder created once decode(view) UTF-8 → UTF-16 JS string new string object

Step 1 — decode a subarray with a cached decoder

const decoder = new TextDecoder("utf-8", { fatal: false });     // create once, reuse forever

function readString(memory, ptr, len) {
  return decoder.decode(new Uint8Array(memory.buffer, ptr, len));
}

Creating a TextDecoder is comparatively expensive; creating one per call is a common hidden cost. A module-level decoder reused for every call removes it. Creating the Uint8Array view over the existing buffer is cheap and does not copy. fatal: false replaces invalid UTF-8 with U+FFFD; use fatal: true to throw on invalid input when the module is supposed to emit valid UTF-8 and you want to catch bugs.

Step 2 — handle shared memory

TextDecoder.decode rejects views over a SharedArrayBuffer in some browsers (it may change while being decoded). Threaded modules with shared memory therefore need a copy into a non-shared buffer first:

function readStringShared(memory, ptr, len) {
  const bytes = new Uint8Array(memory.buffer, ptr, len);
  return decoder.decode(memory.buffer instanceof SharedArrayBuffer ? bytes.slice() : bytes);
}

The slice() copies the bytes into an ordinary ArrayBuffer. For short strings the copy is negligible; for large text, reuse a scratch buffer and copy into it with set, decoding a subarray of the scratch buffer. wasm-bindgen’s glue applies the same workaround in threaded builds.

Step 3 — add an ASCII fast path for tiny strings

For very short strings — single identifiers, field names, short tokens — the fixed overhead of calling decode can exceed the work. A JavaScript loop is faster below roughly 16–32 bytes when the text is ASCII:

function readShortAscii(u8, ptr, len) {
  let s = "";
  for (let i = 0; i < len; i++) {
    const c = u8[ptr + i];
    if (c > 0x7f) return decoder.decode(u8.subarray(ptr, ptr + len));   // fall back for non-ASCII
    s += String.fromCharCode(c);
  }
  return s;
}

Measure before adopting this; the crossover point differs between engines, and modern TextDecoder implementations have narrowed the gap considerably.

Step 4 — decode many strings in one pass

Calling into the module and decoding once per string is wasteful for results containing thousands of strings — the rows of a query, the tokens of a document. Have the module lay out all strings in one contiguous UTF-8 blob with an offsets table, and decode lazily:

// module returns: blobPtr, blobLen, offsetsPtr (Uint32Array of count+1 offsets), count
function stringTable(memory, blobPtr, offsetsPtr, count) {
  const offsets = new Uint32Array(memory.buffer, offsetsPtr, count + 1).slice();   // copy offsets once
  const blob = new Uint8Array(memory.buffer, blobPtr, offsets[count]).slice();      // copy blob once
  const cache = new Array(count);
  return (i) => cache[i] ??= decoder.decode(blob.subarray(offsets[i], offsets[i + 1]));
}

Copying the blob out once makes the table independent of Wasm memory, so it survives later calls and memory growth. Strings are decoded only when accessed and cached afterwards, which pays off when the UI shows a page of results out of thousands. The layout is the flat, column-oriented approach from choosing between JSON and binary serialization.

Many strings as one blob plus an offsets table All strings are stored back to back as UTF-8 in one blob. A separate table of offsets marks where each string starts, with one extra entry for the end. JavaScript decodes string i from the bytes between offsets i and i plus one, only when needed. UTF-8 blob: "alpha" "beta" "gamma" "δέλτα" alpha beta gamma δέλτα (10 B) 0 5 9 14 24

Step 5 — know the cost of the other direction

Writing strings into Wasm memory has its own fast path: TextEncoder.encodeInto(str, view) encodes directly into a view over linear memory, without an intermediate Uint8Array. Allocate str.length * 3 bytes for the worst case (each UTF-16 code unit becomes at most 3 UTF-8 bytes), encode, and use the written count it returns as the length. wasm-bindgen’s passStringToWasm0 uses exactly this, with an ASCII loop first.

Validating text from untrusted modules

When the module is trusted, fatal: false and silent replacement of invalid bytes are fine. When the bytes come from user content passed through the module — a parser for uploaded documents, a plugin written by a third party — decide deliberately what invalid UTF-8 means. A decoder with fatal: true throws a TypeError on the first invalid sequence, which turns corrupt input into a clear error at the boundary. Replacement characters can hide data loss and make round trips lossy: a file that is decoded, edited and re-encoded will not match the original bytes. Lengths deserve the same care: a malicious or buggy module can return a pointer and length that run past the end of memory, and the typed-array constructor will throw — good — but a length that is merely large can make JavaScript decode megabytes of unrelated memory into a string. Check returned lengths against a sensible maximum before decoding, and treat out-of-range values as a module bug to report rather than data to display.

When strings should not cross at all

The fastest string conversion is the one that never happens. Many applications decode strings only to compare them, look them up or count them — work the module could do itself. Search results can be returned as ids and scores, with titles decoded only for the rows on screen. Log levels or categories can be numeric codes mapped to strings in JavaScript once. Templated output — HTML rendered by a Wasm templating engine, for instance — is better returned as one large string than as many fragments, because one large decode costs much less than many small ones. And when the destination is not JavaScript at all — writing to a file, sending over the network — the bytes can stay bytes: pass the Uint8Array straight to fetch, a WritableStream or a Blob without ever creating a JavaScript string. Reserve decoding for text a person will read or JavaScript code must manipulate.

Expected output

Decoding 10,000 strings from a single blob with a cached decoder takes about 1.2 ms in Chrome instead of 9 ms with a decoder created per call; strings from a threaded build decode correctly after the shared-memory copy; and lazily decoded table rows render as soon as they scroll into view.

Gotchas

  • Creating a TextDecoder per call. It is the largest avoidable cost. Cache one.
  • Decoding views over shared memory. Some browsers throw. Copy to a non-shared buffer first.
  • Length in characters instead of bytes. UTF-8 lengths are byte counts. Have the module return byte lengths.
  • Keeping views across calls. Memory growth detaches them. Copy blobs out if they must outlive the call.
  • Decoding unbounded lengths. A bad length decodes unrelated memory. Check it against a maximum first.
  • Assuming ASCII. A fast path must fall back for bytes above 0x7F.

Performance note

For 10,000 strings averaging 24 bytes, decoding with a new TextDecoder per call took 9.1 ms in Chrome, a cached decoder per string 3.4 ms, and one blob with lazy cached decoding of all strings 1.2 ms. Decoding a single 1 MB string took 0.6 ms with TextDecoder.

Decoding 10,000 short strings from Wasm memory Milliseconds to decode ten thousand strings averaging 24 bytes, with a TextDecoder created per call, a cached decoder called per string, and one blob decoded lazily through an offsets table. ms for 10,000 strings new decoder per call 9.1 ms cached decoder per string 3.4 ms blob + offsets, cached decoder 1.2 ms

Frequently Asked Questions

Can JavaScript strings point into Wasm memory without decoding? No. Engines own string storage. Proposals for JS string builtins let Wasm create and inspect JavaScript strings, but linear-memory bytes still need decoding.

Is UTF-16 output from the module faster to decode? Decoding UTF-16 with TextDecoder("utf-16le") is fast too, but doubles the size of ASCII text. UTF-8 is usually the better trade.

What about String.fromCharCode.apply on a view? It works for ASCII but has argument-count limits and is slower than TextDecoder for long strings.

Does TextDecoder handle a BOM? By default it strips a leading UTF-8 BOM; pass ignoreBOM: true to keep it.

Should the module NUL-terminate strings? Not for JavaScript’s sake. Return explicit byte lengths; scanning for a terminator in JavaScript is slower and fails on embedded NULs.

← Back to Zero-Copy Data Transfer Patterns