Choosing Between JSON and Binary Serialization
This page answers one task: a WebAssembly module exchanges a lot of structured data with JavaScript — query results, scene graphs, parsed documents — and the format used to move it has become a measurable cost, so you need to choose deliberately among JSON, binary formats and flat memory layouts.
Prerequisites
- [ ] A module that already passes structured data, so there is a baseline to measure.
- [ ] Representative payloads: typical and worst-case sizes from real use.
- [ ] Familiarity with serializing data with serde-wasm-bindgen.
What the format actually costs
Structured data crosses the boundary in three stages: the module encodes its in-memory structures into some representation, the representation is
transferred (a copy into JavaScript, or a view into linear memory), and JavaScript decodes it into objects it can use. Each format moves cost between
these stages. JSON is relatively expensive to encode in Wasm but decoded by JSON.parse, one of the fastest parsers in any engine. Binary formats such
as MessagePack and bincode are cheap to encode and compact, but are decoded by JavaScript libraries that are slower than native JSON.parse per byte.
Flat layouts — fixed offsets in a buffer, read lazily — skip decoding almost entirely, at the cost of a more rigid schema and less convenient access.
So the “fastest format” depends on payload shape and on what JavaScript does with the data. Reading every field of every record favours formats that decode quickly into objects. Reading a few fields of a large result favours formats that can be read lazily without decoding.
Step 1 — measure the baseline end to end
Before changing format, time the whole path for realistic payloads: from the call into Wasm until JavaScript has the data in the form it needs. Measure encoding, transfer and decoding separately if you can, because the optimisation depends on which dominates:
const t0 = performance.now();
const json = wasm.query_json(q); // Rust: serde_json::to_string
const t1 = performance.now();
const rows = JSON.parse(json);
const t2 = performance.now();
console.log({ encodeAndCopy: t1 - t0, parse: t2 - t1, bytes: json.length });
Repeat with warm-up and many iterations, as in avoiding JIT warm-up errors in Wasm benchmarks. If the total is a small fraction of the operation’s run time, stop here: JSON’s debuggability is worth more than the saving.
Step 2 — try JSON first, done well
JSON through serde_json plus JSON.parse is a strong baseline. Two details make it faster. Return the JSON as bytes and decode with TextDecoder
instead of letting the glue build a string character by character, or let wasm-bindgen return a String (which uses TextDecoder internally). And keep
the encoded form compact — short field names via #[serde(rename)] for very large arrays, no pretty-printing. For payloads of many thousands of records,
JSON frequently beats direct object construction, because JSON.parse builds objects faster than thousands of individual boundary calls.
Step 3 — use a binary format when size or encode cost dominates
When payloads are large and encoding time in Wasm is the bottleneck, or when the data also goes over the network or into storage, a compact binary format
helps. MessagePack is self-describing and schema-flexible like JSON, with good JavaScript decoders (@msgpack/msgpack) and the rmp-serde crate on the
Rust side:
#[wasm_bindgen]
pub fn query_msgpack(q: &str) -> Vec<u8> {
rmp_serde::to_vec_named(&run_query(q)).unwrap() // named: maps with field names
}
import { decode } from "@msgpack/msgpack";
const rows = decode(wasm.query_msgpack(q));
Non-self-describing formats such as bincode or postcard are smaller and faster to encode, but JavaScript needs a decoder that knows the exact schema, usually hand-written or generated. They suit cases where both sides are generated from one schema, or where the data is stored rather than read by JavaScript at all.
Step 4 — use a flat layout for large, read-mostly results
For the largest payloads — a million points, a big table — the fastest option is not to decode at all. Lay the data out in linear memory as columns or fixed-size records and let JavaScript read it through typed-array views, decoding only the fields it touches:
#[wasm_bindgen]
pub struct Points { xs: Vec<f32>, ys: Vec<f32>, ids: Vec<u32> }
#[wasm_bindgen]
impl Points {
pub fn len(&self) -> usize { self.xs.len() }
pub fn xs_ptr(&self) -> *const f32 { self.xs.as_ptr() }
pub fn ys_ptr(&self) -> *const f32 { self.ys.as_ptr() }
pub fn ids_ptr(&self) -> *const u32 { self.ids.as_ptr() }
}
const n = pts.len();
const xs = new Float32Array(memory.buffer, pts.xs_ptr(), n); // no decode, no copy
The cost of encoding and decoding disappears; what remains is the discipline of views — valid only until memory grows or the object is freed — described in creating views into Wasm memory safely. Schema libraries such as FlatBuffers and Cap’n Proto formalise this approach for nested data, with generated accessors on both sides.
Step 5 — weigh evolution and debuggability
Speed is only one axis. Self-describing formats (JSON, MessagePack) tolerate added fields gracefully: old readers ignore what they do not know. Schema-bound formats (bincode, flat layouts) need both sides updated together, which is fine when the module and its JavaScript wrapper ship as one package and risky when data is persisted or shared between versions. JSON can be read in a network panel or a log; binary formats need tooling. A common, sensible split is JSON (or serde-wasm-bindgen) for control-plane data — configuration, small results, errors — and a flat binary layout for the one or two bulk data paths where profiling shows the format matters.
Payload shape matters more than format
The numbers in benchmarks are dominated less by the format than by the shape of the data. Arrays of many small objects are expensive in every format,
because each object has to be created in JavaScript, and object creation, not parsing, becomes the bottleneck. Restructuring the payload often beats
switching format: returning columns — an array of ids, an array of scores, an array of titles — instead of an array of { id, score, title } objects
cuts object creation to a handful of arrays. For numeric columns, typed arrays make it nearly free. Paging large results, so that JavaScript asks for the
200 rows it will display rather than all 50,000, removes most of the cost entirely. And deduplicating repeated strings — category names, units — into a
lookup table shrinks payloads in any format. Try these shape changes before reaching for a new serialisation library; they are usually simpler and
larger wins, and they keep the format debuggable.
Whatever the format, keep the encoding code in one place on each side. A single encode_results function in Rust and a single decodeResults in
JavaScript make it straightforward to swap formats later, to measure each in isolation, and to add a version byte at the start of binary payloads so that
a future layout change can be detected rather than misread.
Expected output
For a 50,000-row query result, measurements show JSON at about 23 ms end to end, MessagePack at about 28 ms with a 40% smaller payload, and a flat columnar layout at about 6 ms with lazy reads — and the chosen design uses JSON for metadata and the columnar layout for the rows.
Gotchas
- Optimising without measuring. Format changes add complexity. Prove the format is the bottleneck first.
- Assuming binary is always faster. JavaScript binary decoders can be slower than native
JSON.parse. - Rigid formats for persisted data. bincode files break when the struct changes. Version the format.
- Holding flat-layout views too long. Growth invalidates them. Re-create after calls that may allocate.
- 64-bit integers in JSON. They lose precision above 2⁵³. Encode as strings or use a binary format with BigInt.
Performance note
For 50,000 records of five fields in Chrome, serde-wasm-bindgen took 41 ms, JSON 23 ms, MessagePack 28 ms, and a columnar flat layout 6 ms including reading every value once. For 200 records, all approaches were under 0.5 ms and the differences were irrelevant.
Frequently Asked Questions
Is Protocol Buffers a good choice? It works, with generated code on both sides, and suits data that also travels to servers. For boundary-only data it offers little over MessagePack.
Does compression help? Not for in-process boundary data — compressing and decompressing costs more than copying. It helps for network and storage.
Can I use structuredClone-style transfer instead?
Only between JavaScript realms (workers). Wasm data still has to be encoded into JavaScript values first.
What about strings in flat layouts? Store them as offsets into a shared UTF-8 blob and decode on access; see decoding strings directly from Wasm memory.
Should the same format be used for the network and the boundary? It can be convenient — a MessagePack response can be passed straight into the module without re-encoding — but optimise each path for its own costs.
Related
- Passing arrays between JavaScript and Wasm — the numeric building block.
- Encoding strings across the Wasm boundary — the text building block.
- Understanding the canonical ABI — the Component Model’s standard encoding.
- Measuring JS-to-Wasm call overhead — why many small calls cost so much.