Choosing Between JSON and Binary Serialization

This page answers one task: a WebAssembly module exchanges a lot of structured data with JavaScript — query results, scene graphs, parsed documents — and the format used to move it has become a measurable cost, so you need to choose deliberately among JSON, binary formats and flat memory layouts.

Prerequisites

  • [ ] A module that already passes structured data, so there is a baseline to measure.
  • [ ] Representative payloads: typical and worst-case sizes from real use.
  • [ ] Familiarity with serializing data with serde-wasm-bindgen.

What the format actually costs

Structured data crosses the boundary in three stages: the module encodes its in-memory structures into some representation, the representation is transferred (a copy into JavaScript, or a view into linear memory), and JavaScript decodes it into objects it can use. Each format moves cost between these stages. JSON is relatively expensive to encode in Wasm but decoded by JSON.parse, one of the fastest parsers in any engine. Binary formats such as MessagePack and bincode are cheap to encode and compact, but are decoded by JavaScript libraries that are slower than native JSON.parse per byte. Flat layouts — fixed offsets in a buffer, read lazily — skip decoding almost entirely, at the cost of a more rigid schema and less convenient access.

So the “fastest format” depends on payload shape and on what JavaScript does with the data. Reading every field of every record favours formats that decode quickly into objects. Reading a few fields of a large result favours formats that can be read lazily without decoding.

Serialization formats for the Wasm boundary JSON is readable and parsed natively but larger and slower to encode. serde-wasm-bindgen builds objects directly with many boundary calls. MessagePack and bincode are compact binary formats decoded by JavaScript libraries. A flat layout is read lazily with typed arrays and needs no decoding. format decode in JS size schema change JSON JSON.parse (native) large flexible serde-wasm-bindgen built directly n/a (objects) flexible MessagePack JS library compact flexible bincode / postcard hand or generated very compact rigid flat layout lazy typed-array reads compact rigid

Step 1 — measure the baseline end to end

Before changing format, time the whole path for realistic payloads: from the call into Wasm until JavaScript has the data in the form it needs. Measure encoding, transfer and decoding separately if you can, because the optimisation depends on which dominates:

const t0 = performance.now();
const json = wasm.query_json(q);                  // Rust: serde_json::to_string
const t1 = performance.now();
const rows = JSON.parse(json);
const t2 = performance.now();
console.log({ encodeAndCopy: t1 - t0, parse: t2 - t1, bytes: json.length });

Repeat with warm-up and many iterations, as in avoiding JIT warm-up errors in Wasm benchmarks. If the total is a small fraction of the operation’s run time, stop here: JSON’s debuggability is worth more than the saving.

Step 2 — try JSON first, done well

JSON through serde_json plus JSON.parse is a strong baseline. Two details make it faster. Return the JSON as bytes and decode with TextDecoder instead of letting the glue build a string character by character, or let wasm-bindgen return a String (which uses TextDecoder internally). And keep the encoded form compact — short field names via #[serde(rename)] for very large arrays, no pretty-printing. For payloads of many thousands of records, JSON frequently beats direct object construction, because JSON.parse builds objects faster than thousands of individual boundary calls.

Step 3 — use a binary format when size or encode cost dominates

When payloads are large and encoding time in Wasm is the bottleneck, or when the data also goes over the network or into storage, a compact binary format helps. MessagePack is self-describing and schema-flexible like JSON, with good JavaScript decoders (@msgpack/msgpack) and the rmp-serde crate on the Rust side:

#[wasm_bindgen]
pub fn query_msgpack(q: &str) -> Vec<u8> {
    rmp_serde::to_vec_named(&run_query(q)).unwrap()      // named: maps with field names
}
import { decode } from "@msgpack/msgpack";
const rows = decode(wasm.query_msgpack(q));

Non-self-describing formats such as bincode or postcard are smaller and faster to encode, but JavaScript needs a decoder that knows the exact schema, usually hand-written or generated. They suit cases where both sides are generated from one schema, or where the data is stored rather than read by JavaScript at all.

Step 4 — use a flat layout for large, read-mostly results

For the largest payloads — a million points, a big table — the fastest option is not to decode at all. Lay the data out in linear memory as columns or fixed-size records and let JavaScript read it through typed-array views, decoding only the fields it touches:

#[wasm_bindgen]
pub struct Points { xs: Vec<f32>, ys: Vec<f32>, ids: Vec<u32> }

#[wasm_bindgen]
impl Points {
    pub fn len(&self) -> usize { self.xs.len() }
    pub fn xs_ptr(&self) -> *const f32 { self.xs.as_ptr() }
    pub fn ys_ptr(&self) -> *const f32 { self.ys.as_ptr() }
    pub fn ids_ptr(&self) -> *const u32 { self.ids.as_ptr() }
}
const n = pts.len();
const xs = new Float32Array(memory.buffer, pts.xs_ptr(), n);    // no decode, no copy

The cost of encoding and decoding disappears; what remains is the discipline of views — valid only until memory grows or the object is freed — described in creating views into Wasm memory safely. Schema libraries such as FlatBuffers and Cap’n Proto formalise this approach for nested data, with generated accessors on both sides.

Choosing a format for boundary data For small or moderate payloads, use serde-wasm-bindgen or JSON for simplicity. For large payloads read in full, compare JSON with MessagePack. For large numeric or read-mostly data, use a flat layout read through typed arrays. How big is the payload, and how is it read? small to moderate serde-wasm-bindgen or JSON — simplest large, every field read JSON or MessagePack — measure both large, numeric or partly read flat layout + typed-array views

Step 5 — weigh evolution and debuggability

Speed is only one axis. Self-describing formats (JSON, MessagePack) tolerate added fields gracefully: old readers ignore what they do not know. Schema-bound formats (bincode, flat layouts) need both sides updated together, which is fine when the module and its JavaScript wrapper ship as one package and risky when data is persisted or shared between versions. JSON can be read in a network panel or a log; binary formats need tooling. A common, sensible split is JSON (or serde-wasm-bindgen) for control-plane data — configuration, small results, errors — and a flat binary layout for the one or two bulk data paths where profiling shows the format matters.

Payload shape matters more than format

The numbers in benchmarks are dominated less by the format than by the shape of the data. Arrays of many small objects are expensive in every format, because each object has to be created in JavaScript, and object creation, not parsing, becomes the bottleneck. Restructuring the payload often beats switching format: returning columns — an array of ids, an array of scores, an array of titles — instead of an array of { id, score, title } objects cuts object creation to a handful of arrays. For numeric columns, typed arrays make it nearly free. Paging large results, so that JavaScript asks for the 200 rows it will display rather than all 50,000, removes most of the cost entirely. And deduplicating repeated strings — category names, units — into a lookup table shrinks payloads in any format. Try these shape changes before reaching for a new serialisation library; they are usually simpler and larger wins, and they keep the format debuggable.

Whatever the format, keep the encoding code in one place on each side. A single encode_results function in Rust and a single decodeResults in JavaScript make it straightforward to swap formats later, to measure each in isolation, and to add a version byte at the start of binary payloads so that a future layout change can be detected rather than misread.

Expected output

For a 50,000-row query result, measurements show JSON at about 23 ms end to end, MessagePack at about 28 ms with a 40% smaller payload, and a flat columnar layout at about 6 ms with lazy reads — and the chosen design uses JSON for metadata and the columnar layout for the rows.

Gotchas

  • Optimising without measuring. Format changes add complexity. Prove the format is the bottleneck first.
  • Assuming binary is always faster. JavaScript binary decoders can be slower than native JSON.parse.
  • Rigid formats for persisted data. bincode files break when the struct changes. Version the format.
  • Holding flat-layout views too long. Growth invalidates them. Re-create after calls that may allocate.
  • 64-bit integers in JSON. They lose precision above 2⁵³. Encode as strings or use a binary format with BigInt.

Performance note

For 50,000 records of five fields in Chrome, serde-wasm-bindgen took 41 ms, JSON 23 ms, MessagePack 28 ms, and a columnar flat layout 6 ms including reading every value once. For 200 records, all approaches were under 0.5 ms and the differences were irrelevant.

Moving 50,000 records from Wasm to JavaScript End-to-end milliseconds to deliver fifty thousand five-field records to JavaScript in usable form, for serde-wasm-bindgen, JSON, MessagePack and a columnar flat layout. ms end to end serde-wasm-bindgen 41 ms MessagePack 28 ms JSON 23 ms columnar flat layout 6 ms

Frequently Asked Questions

Is Protocol Buffers a good choice? It works, with generated code on both sides, and suits data that also travels to servers. For boundary-only data it offers little over MessagePack.

Does compression help? Not for in-process boundary data — compressing and decompressing costs more than copying. It helps for network and storage.

Can I use structuredClone-style transfer instead? Only between JavaScript realms (workers). Wasm data still has to be encoded into JavaScript values first.

What about strings in flat layouts? Store them as offsets into a shared UTF-8 blob and decode on access; see decoding strings directly from Wasm memory.

Should the same format be used for the network and the boundary? It can be convenient — a MessagePack response can be passed straight into the module without re-encoding — but optimise each path for its own costs.

← Back to Passing Complex Types Across the Boundary