Passing Nested Objects Efficiently
This page answers one task: a WebAssembly function takes or returns a deeply nested structure — a document tree, a scene graph, a syntax tree, a nested JSON configuration — and converting it costs more than the work the function does. You want an encoding that moves the structure across the boundary in a few bulk copies instead of thousands of small conversions.
Prerequisites
- [ ] A module that exchanges tree-shaped or graph-shaped data with JavaScript.
- [ ] A benchmark of the current conversion (serde-wasm-bindgen, JSON, or manual
Reflectcalls). - [ ] Knowledge of which fields each side actually reads.
Why nested objects are expensive to convert
Converting a JavaScript object into Rust with serde-wasm-bindgen walks the object: for each property it calls into JavaScript to read the value, checks
its type, converts strings to UTF-8, and recurses. Each of those steps is a boundary call, and a tree of 50,000 nodes with five fields each means hundreds
of thousands of calls. The other direction is similar: creating each JavaScript object and setting each property is a call. JSON is sometimes faster —
JSON.stringify runs natively in the engine, and Rust parses one string — but it still encodes every number as text and every key repeatedly.
The fix is to change the shape of the data at the boundary. Trees can be represented as a few flat arrays: one entry per node, with numeric fields in typed arrays, children referenced by index, and strings in a shared table. That representation crosses the boundary in a handful of copies, and both sides can rebuild whatever native structure they prefer — or work on the flat arrays directly.
Step 1 — measure where time goes
Before redesigning, measure conversion separately from the work:
const t0 = performance.now();
const handle = wasm.load_scene(sceneObject); // conversion + build
const t1 = performance.now();
wasm.layout(handle); // the actual work
const t2 = performance.now();
console.log({ convertMs: t1 - t0, workMs: t2 - t1 });
If conversion is a small fraction, stop here. If it dominates — common for trees above a few thousand nodes — continue.
Step 2 — flatten the tree into parallel arrays
Represent the tree as arrays indexed by node number, in depth-first order so children follow their parents:
function flatten(root) {
const kinds = [], parents = [], xs = [], ys = [], nameIdx = [];
const strings = new Map(); // string → index
const intern = (s) => strings.has(s) ? strings.get(s) : (strings.set(s, strings.size), strings.size - 1);
(function walk(node, parent) {
const id = kinds.length;
kinds.push(KIND[node.type]); parents.push(parent);
xs.push(node.x ?? 0); ys.push(node.y ?? 0); nameIdx.push(intern(node.name ?? ""));
for (const c of node.children ?? []) walk(c, id);
})(root, -1);
return {
kinds: Uint8Array.from(kinds), parents: Int32Array.from(parents),
xs: Float32Array.from(xs), ys: Float32Array.from(ys), names: Uint32Array.from(nameIdx),
stringTable: [...strings.keys()],
};
}
On the Rust side, accept the arrays as slices and build or use them directly:
#[wasm_bindgen]
pub fn load_scene(kinds: &[u8], parents: &[i32], xs: &[f32], ys: &[f32], names: &[u32], strings: Vec<String>) -> SceneHandle {
Scene::from_columns(kinds, parents, xs, ys, names, strings).into()
}
Five typed-array copies and one string array replace a recursive conversion of every node.
Step 3 — encode strings once
Strings are often the most expensive part. A string table — each distinct string once, nodes referring to it by index — removes repetition (element names, CSS class names and identifiers repeat constantly). For very large tables, go further: concatenate all strings into one UTF-8 buffer with an offsets array, so the whole table crosses as two typed arrays, and decode lazily on the side that needs text.
Step 4 — return results in the same shape
Results from the module — computed positions, new nodes, diagnostics — come back the same way: typed arrays indexed by node number, returned as views
over linear memory or copied out once. For a layout engine, returning Float32Arrays of x, y, width and height for every node is one copy each; the
JavaScript side applies them to its own objects in a tight loop. For sparse results (only changed nodes), return an index array plus value arrays.
Step 5 — keep the tree on one side when you can
The biggest saving is not converting at all. If the module owns the document — an editor core, a game world, a layout engine — keep the tree inside the
module and expose operations and queries: insert_node, set_attribute, nodes_in_rect. JavaScript holds node IDs, not copies. The UI reads only what it
renders, through bulk queries. That design moves data across the boundary in proportion to changes, not to the tree’s size.
Binary serialisation as a middle ground
When the structure is complex — many node types with different fields — maintaining parallel arrays by hand becomes error-prone. A binary serialisation
format with code generation for both sides (FlatBuffers, Cap’n Proto, or a schema-driven format such as bincode/postcard on the Rust side with a
matching JavaScript encoder) sends one byte buffer across the boundary. FlatBuffers and Cap’n Proto can be read in place without parsing, so the
receiving side touches only the fields it needs. The trade-off is schema maintenance and code size for the generated readers; see
choosing between JSON and binary serialization.
When serde-wasm-bindgen is the right choice
For small and medium structures — configuration objects, API responses of a few hundred nodes, results shown in a form — serde-wasm-bindgen is
convenient, type-safe and fast enough, and flattening would add complexity for no measurable gain. Reserve the techniques here for structures where
measurement shows conversion dominating, typically above several thousand nodes or when the conversion runs every frame.
Optional and variable fields
Real trees are not uniform: some nodes have a name, some have a list of attributes, some carry text. Parallel arrays handle optional scalar fields with a
sentinel (NaN for a missing float, 0xffffffff for a missing string index) or a separate presence bitmap when sentinels are ambiguous. Variable-length
fields — a node’s attribute list, a polygon’s points — use the same trick as strings: one flat array holding all the values back to back, plus an offsets
array with one entry per node marking where its values start, and one extra entry for the end. Node i’s values occupy values[offsets[i]..offsets[i+1]].
This “offsets plus values” pattern is the columnar representation used by Apache Arrow, and for very large tabular or tree data, Arrow itself — with
readers available for both JavaScript and Rust — can be the encoding, giving a well-specified format instead of a hand-rolled one.
Validating flattened input
Flattened arrays are fast but trust-sensitive: a parent index that points outside the array, a cycle in parent references, or offsets that decrease can crash or hang code that assumes a well-formed tree. When the arrays come from code you control, a debug-build check is enough; when they come from elsewhere — a file, another team’s code, a plugin — validate in the module before building the tree: every parent index is less than the node’s own index (which also rules out cycles in depth-first order), every string index is within the string table, and offsets are non-decreasing and end at the values array’s length. These checks are linear, cheap compared with conversion, and turn malformed input into a clear error instead of a trap.
Expected output
Loading a 50,000-node scene drops from 180 ms with serde-wasm-bindgen to 9 ms with flattened columns and a string table; layout results return as four
Float32Arrays applied in 2 ms; the editor keeps the document inside the module and sends only edits across the boundary; and small configuration
objects still use serde.
Gotchas
- Optimising before measuring. Conversion may be negligible. Time it separately.
- Repeated strings. Encode each distinct string once.
- Index mismatches between sides. Generate the flattening code from one schema, or test round trips.
- Converting the whole tree per change. Keep it on one side and send changes.
- Hand-maintained layouts for complex schemas. Use a schema-driven binary format.
Performance note
For the 50,000-node scene, the flattened encoding moved 2.1 MB in six copies; serde-wasm-bindgen made about 400,000 boundary calls; JSON took 95 ms
including stringify and parsing.
Frequently Asked Questions
Does reference-types support make serde faster? It reduces some glue overhead, but per-property conversion still dominates for large trees.
Can I pass the flattened arrays to a worker too? Yes — typed arrays transfer to workers without copying.
How do I handle cycles in graphs? Flattened encodings handle them naturally with index references; serde does not.
Is JSON ever the best choice? For data already arriving as JSON text from the network, parsing it directly in Rust avoids building JavaScript objects at all.
How do I encode a node’s variable-length list in flat arrays? Put all values back to back in one array and add an offsets array marking where each node’s values start.
Related
- Serializing data with serde-wasm-bindgen — the convenient default.
- Choosing between JSON and binary serialization — format trade-offs.
- Passing maps and sets across the boundary — key-value data.
- Decoding strings directly from Wasm memory — string tables.