Passing Nested Objects Efficiently

This page answers one task: a WebAssembly function takes or returns a deeply nested structure — a document tree, a scene graph, a syntax tree, a nested JSON configuration — and converting it costs more than the work the function does. You want an encoding that moves the structure across the boundary in a few bulk copies instead of thousands of small conversions.

Prerequisites

  • [ ] A module that exchanges tree-shaped or graph-shaped data with JavaScript.
  • [ ] A benchmark of the current conversion (serde-wasm-bindgen, JSON, or manual Reflect calls).
  • [ ] Knowledge of which fields each side actually reads.

Why nested objects are expensive to convert

Converting a JavaScript object into Rust with serde-wasm-bindgen walks the object: for each property it calls into JavaScript to read the value, checks its type, converts strings to UTF-8, and recurses. Each of those steps is a boundary call, and a tree of 50,000 nodes with five fields each means hundreds of thousands of calls. The other direction is similar: creating each JavaScript object and setting each property is a call. JSON is sometimes faster — JSON.stringify runs natively in the engine, and Rust parses one string — but it still encodes every number as text and every key repeatedly.

The fix is to change the shape of the data at the boundary. Trees can be represented as a few flat arrays: one entry per node, with numeric fields in typed arrays, children referenced by index, and strings in a shared table. That representation crosses the boundary in a handful of copies, and both sides can rebuild whatever native structure they prefer — or work on the flat arrays directly.

Converting a tree object by object versus as flat arrays Converting a nested object with serde-wasm-bindgen makes boundary calls for every node and property, so cost grows with node count times fields. Flattening the tree into typed arrays of fields, parent indices and a string table moves it in a few bulk copies regardless of size. object-by-object (serde) calls per node and property strings converted one by one cost ∝ nodes × fields fine for small trees flat arrays + string table few bulk copies one UTF-8 blob for strings cost ∝ bytes large trees

Step 1 — measure where time goes

Before redesigning, measure conversion separately from the work:

const t0 = performance.now();
const handle = wasm.load_scene(sceneObject);     // conversion + build
const t1 = performance.now();
wasm.layout(handle);                             // the actual work
const t2 = performance.now();
console.log({ convertMs: t1 - t0, workMs: t2 - t1 });

If conversion is a small fraction, stop here. If it dominates — common for trees above a few thousand nodes — continue.

Step 2 — flatten the tree into parallel arrays

Represent the tree as arrays indexed by node number, in depth-first order so children follow their parents:

function flatten(root) {
  const kinds = [], parents = [], xs = [], ys = [], nameIdx = [];
  const strings = new Map();                     // string → index
  const intern = (s) => strings.has(s) ? strings.get(s) : (strings.set(s, strings.size), strings.size - 1);
  (function walk(node, parent) {
    const id = kinds.length;
    kinds.push(KIND[node.type]); parents.push(parent);
    xs.push(node.x ?? 0); ys.push(node.y ?? 0); nameIdx.push(intern(node.name ?? ""));
    for (const c of node.children ?? []) walk(c, id);
  })(root, -1);
  return {
    kinds: Uint8Array.from(kinds), parents: Int32Array.from(parents),
    xs: Float32Array.from(xs), ys: Float32Array.from(ys), names: Uint32Array.from(nameIdx),
    stringTable: [...strings.keys()],
  };
}

On the Rust side, accept the arrays as slices and build or use them directly:

#[wasm_bindgen]
pub fn load_scene(kinds: &[u8], parents: &[i32], xs: &[f32], ys: &[f32], names: &[u32], strings: Vec<String>) -> SceneHandle {
    Scene::from_columns(kinds, parents, xs, ys, names, strings).into()
}

Five typed-array copies and one string array replace a recursive conversion of every node.

A tree encoded as parallel arrays Each node has one entry in every column. The parents column stores the index of each node's parent, minus one for the root. Numeric fields are in typed arrays. Names are indices into a string table, so repeated names are stored once. nodes in depth-first order: 0 root, 1–2 children of 0, 3 child of 2 kinds (u8) parents (i32) x, y (f32) names (u32 idx) 4 nodes parents: −1, 0, 0, 2 + string table

Step 3 — encode strings once

Strings are often the most expensive part. A string table — each distinct string once, nodes referring to it by index — removes repetition (element names, CSS class names and identifiers repeat constantly). For very large tables, go further: concatenate all strings into one UTF-8 buffer with an offsets array, so the whole table crosses as two typed arrays, and decode lazily on the side that needs text.

Step 4 — return results in the same shape

Results from the module — computed positions, new nodes, diagnostics — come back the same way: typed arrays indexed by node number, returned as views over linear memory or copied out once. For a layout engine, returning Float32Arrays of x, y, width and height for every node is one copy each; the JavaScript side applies them to its own objects in a tight loop. For sparse results (only changed nodes), return an index array plus value arrays.

Step 5 — keep the tree on one side when you can

The biggest saving is not converting at all. If the module owns the document — an editor core, a game world, a layout engine — keep the tree inside the module and expose operations and queries: insert_node, set_attribute, nodes_in_rect. JavaScript holds node IDs, not copies. The UI reads only what it renders, through bulk queries. That design moves data across the boundary in proportion to changes, not to the tree’s size.

Binary serialisation as a middle ground

When the structure is complex — many node types with different fields — maintaining parallel arrays by hand becomes error-prone. A binary serialisation format with code generation for both sides (FlatBuffers, Cap’n Proto, or a schema-driven format such as bincode/postcard on the Rust side with a matching JavaScript encoder) sends one byte buffer across the boundary. FlatBuffers and Cap’n Proto can be read in place without parsing, so the receiving side touches only the fields it needs. The trade-off is schema maintenance and code size for the generated readers; see choosing between JSON and binary serialization.

When serde-wasm-bindgen is the right choice

For small and medium structures — configuration objects, API responses of a few hundred nodes, results shown in a form — serde-wasm-bindgen is convenient, type-safe and fast enough, and flattening would add complexity for no measurable gain. Reserve the techniques here for structures where measurement shows conversion dominating, typically above several thousand nodes or when the conversion runs every frame.

Optional and variable fields

Real trees are not uniform: some nodes have a name, some have a list of attributes, some carry text. Parallel arrays handle optional scalar fields with a sentinel (NaN for a missing float, 0xffffffff for a missing string index) or a separate presence bitmap when sentinels are ambiguous. Variable-length fields — a node’s attribute list, a polygon’s points — use the same trick as strings: one flat array holding all the values back to back, plus an offsets array with one entry per node marking where its values start, and one extra entry for the end. Node i’s values occupy values[offsets[i]..offsets[i+1]]. This “offsets plus values” pattern is the columnar representation used by Apache Arrow, and for very large tabular or tree data, Arrow itself — with readers available for both JavaScript and Rust — can be the encoding, giving a well-specified format instead of a hand-rolled one.

Validating flattened input

Flattened arrays are fast but trust-sensitive: a parent index that points outside the array, a cycle in parent references, or offsets that decrease can crash or hang code that assumes a well-formed tree. When the arrays come from code you control, a debug-build check is enough; when they come from elsewhere — a file, another team’s code, a plugin — validate in the module before building the tree: every parent index is less than the node’s own index (which also rules out cycles in depth-first order), every string index is within the string table, and offsets are non-decreasing and end at the values array’s length. These checks are linear, cheap compared with conversion, and turn malformed input into a clear error instead of a trap.

Expected output

Loading a 50,000-node scene drops from 180 ms with serde-wasm-bindgen to 9 ms with flattened columns and a string table; layout results return as four Float32Arrays applied in 2 ms; the editor keeps the document inside the module and sends only edits across the boundary; and small configuration objects still use serde.

Gotchas

  • Optimising before measuring. Conversion may be negligible. Time it separately.
  • Repeated strings. Encode each distinct string once.
  • Index mismatches between sides. Generate the flattening code from one schema, or test round trips.
  • Converting the whole tree per change. Keep it on one side and send changes.
  • Hand-maintained layouts for complex schemas. Use a schema-driven binary format.

Performance note

For the 50,000-node scene, the flattened encoding moved 2.1 MB in six copies; serde-wasm-bindgen made about 400,000 boundary calls; JSON took 95 ms including stringify and parsing.

Loading a 50,000-node scene into Wasm Milliseconds to move a 50,000-node scene tree into the module with serde-wasm-bindgen, with JSON stringify plus parsing in Rust, and with flattened typed-array columns and a string table. ms per load serde-wasm-bindgen 180 ms JSON.stringify + serde_json 95 ms flattened columns + string table 9 ms

Frequently Asked Questions

Does reference-types support make serde faster? It reduces some glue overhead, but per-property conversion still dominates for large trees.

Can I pass the flattened arrays to a worker too? Yes — typed arrays transfer to workers without copying.

How do I handle cycles in graphs? Flattened encodings handle them naturally with index references; serde does not.

Is JSON ever the best choice? For data already arriving as JSON text from the network, parsing it directly in Rust avoids building JavaScript objects at all.

How do I encode a node’s variable-length list in flat arrays? Put all values back to back in one array and add an offsets array marking where each node’s values start.

← Back to Passing Complex Types Across the Boundary