Porting a JavaScript Parser to Rust Wasm

This page answers one task: a JavaScript parser — for a log format, a query language, CSV, a binary record format — is the measured bottleneck, and you want to port it to Rust compiled to WebAssembly while callers keep using exactly the same JavaScript API.

Prerequisites

  • [ ] A profile showing the parser dominates the scenario, as in profiling JavaScript to find Wasm candidates.
  • [ ] The existing parser’s test suite, or a corpus of real inputs with known outputs.
  • [ ] Rust, wasm-pack or wasm-bindgen-cli, and the project’s bundler.

The shape of a good port

A naive port mirrors the JavaScript structure: a Tokenizer class exported from Rust, a next() method called once per token from JavaScript, and a result object built field by field through the boundary. That design pays a boundary crossing per token and per field, and often ends up slower than the original. A good port moves the whole loop across the boundary: JavaScript hands over the entire input once, Rust tokenizes and parses everything, and the result comes back in one step — as a compact binary buffer, a single JSON string, or a small number of typed arrays.

The second design principle is that the public API does not change. Callers import parse(text) and receive the same objects, with the same error messages and positions, as before. A thin JavaScript wrapper adapts the Wasm module’s efficient interface to the old API, which also makes it possible to run both implementations side by side during rollout.

Two ways to port a parser A token-by-token port calls into Wasm for every token and builds result objects through the boundary, paying a crossing per token. A whole-input port passes the input once, parses entirely in Rust, and returns a compact result decoded by a thin wrapper that preserves the old API. token-by-token next() called per token fields read one by one crossings dominate often slower than JS whole input, one result input copied in once parse loop entirely in Rust one compact result back the design to use

Step 1 — pin down the existing behaviour

Before writing Rust, freeze what “correct” means. Collect inputs that exercise every branch of the existing parser — valid documents, every error case, edge cases like empty input, trailing newlines, Unicode, very long lines — and record the current outputs and error messages as golden files:

for (const file of await glob("corpus/**/*.txt")) {
  const text = await readFile(file, "utf8");
  let out;
  try { out = { ok: parseJs(text) }; } catch (e) { out = { error: e.message, line: e.line, col: e.column }; }
  await writeFile(file + ".golden.json", JSON.stringify(out, null, 1));
}

These goldens become the port’s acceptance test. Differences between the two implementations are bugs in the port unless you decide otherwise.

Step 2 — write the parser in Rust over bytes

Parse UTF-8 bytes rather than JavaScript strings. Rust’s &str is UTF-8, so the input crosses the boundary as bytes and is validated once:

#[derive(Debug)]
pub struct Record { pub ts: u64, pub level: u8, pub msg_start: u32, pub msg_len: u32 }

pub fn parse(input: &str) -> Result<Vec<Record>, ParseError> {
    let mut out = Vec::with_capacity(input.len() / 80);
    for (lineno, line) in input.lines().enumerate() {
        let (ts, rest) = parse_ts(line).map_err(|c| ParseError::at(lineno + 1, c, "expected timestamp"))?;
        let (level, msg) = parse_level(rest).map_err(|c| ParseError::at(lineno + 1, c, "unknown level"))?;
        let msg_start = (msg.as_ptr() as usize - input.as_ptr() as usize) as u32;
        out.push(Record { ts, level, msg_start, msg_len: msg.len() as u32 });
    }
    Ok(out)
}

Records refer to message text by offset and length into the input rather than copying strings; the JavaScript wrapper decodes only the messages callers actually read. Keep the Rust code free of wasm-bindgen types so it can be tested natively with cargo test against the same golden files.

Step 3 — design a one-call boundary

Export a single function that takes the input and returns a flat result. With wasm-bindgen, returning a Vec<u32> copies one typed array back:

#[wasm_bindgen]
pub fn parse_flat(input: &str) -> Result<Vec<u32>, JsError> {
    let recs = parse(input).map_err(|e| JsError::new(&format!("{}:{}:{}", e.line, e.col, e.msg)))?;
    let mut flat = Vec::with_capacity(recs.len() * 5);
    for r in recs {
        flat.extend([(r.ts >> 32) as u32, r.ts as u32, r.level as u32, r.msg_start, r.msg_len]);
    }
    Ok(flat)
}

The wrapper rebuilds the familiar objects, lazily decoding message text from the original string’s bytes:

import { parse_flat } from "./pkg/logparse.js";
const encoder = new TextEncoder(), decoder = new TextDecoder();

export function parse(text) {
  const bytes = encoder.encode(text);
  const flat = parse_flat(text);                         // one boundary call
  const out = new Array(flat.length / 5);
  for (let i = 0, j = 0; i < flat.length; i += 5, j++) {
    const s = flat[i + 3], n = flat[i + 4];
    out[j] = {
      ts: flat[i] * 2 ** 32 + flat[i + 1],
      level: LEVELS[flat[i + 2]],
      get message() { return decoder.decode(bytes.subarray(s, s + n)); },
    };
  }
  return out;
}

The getter means callers that never read message never pay to decode it. If callers spread or serialise records, replace the getter with eager decoding — measure both.

One parse call through the ported module The wrapper encodes the input to UTF-8 and calls parse_flat once. Rust parses all lines and returns a flat Uint32Array of five fields per record. The wrapper rebuilds the original record objects, decoding message text lazily from offsets into the input bytes. parse(text) same public API parse_flat(text) one crossing, UTF-8 in Rust parse loop no boundary calls flat Uint32Array 5 fields per record rebuild objects lazy message decode

Step 4 — match errors exactly

Error behaviour is part of the API. The original parser threw SyntaxError objects with line and column; the port must too. Encode the position in the Rust error and translate it in the wrapper:

try { return parseInternal(text); }
catch (e) {
  const [line, col, ...msg] = String(e.message).split(":");
  const err = new SyntaxError(msg.join(":")); err.line = +line; err.column = +col;
  throw err;
}

Columns deserve care: JavaScript counted UTF-16 code units, Rust counts bytes. Convert byte offsets to UTF-16 columns for the error line before reporting, or golden tests with non-ASCII input will disagree.

Step 5 — measure every stage

Time the old parser, the Wasm call alone, and the full wrapper including encoding and object rebuilding, on production-sized inputs. The full wrapper is what callers experience. If object rebuilding dominates, the remaining win may lie in letting performance-sensitive callers use the flat result directly through a second, lower-level API, while keeping the compatible one.

Sharing code between the browser and the server

A parser often runs in more than one place: in the browser for previews and validation, on the server for ingestion. A Rust core gives you one implementation for both — compiled to WebAssembly for the browser and Node, natively for a Rust server, or as a Wasm module inside a Node or Deno backend. Keep the core crate free of browser-specific code and put the wasm-bindgen exports in a thin separate crate, as described in using Cargo features for Wasm and native builds. The golden tests then run against every build: cargo test for the native core, and a JavaScript test for the Wasm wrapper. That guarantees the browser and the server agree, which was rarely true when they had separate JavaScript and server-language implementations.

When the port is not enough

Sometimes the port lands and the end-to-end gain is smaller than the profile suggested. The usual reasons are visible in the stage timings: encoding the input to UTF-8 and copying it in takes noticeable time for very large strings; rebuilding millions of JavaScript objects costs as much as the parse; or the data arrived as a string only because an earlier step decoded bytes that could have been passed directly. Each has a fix — take Uint8Array input from fetch or File instead of a string, expose the flat result to callers that can use it, or stream input in chunks. The profile after the port is as important as the one before it, and the porting decision is about the whole pipeline, not just the parse function.

Expected output

All golden tests pass against the Rust parser natively and through the Wasm wrapper; the public parse() API is unchanged; and parsing a 40 MB log file drops from 2.6 s to 0.7 s end to end in Chrome, with the Wasm call itself taking 0.35 s.

Gotchas

  • Token-by-token APIs. Boundary crossings swamp the gain. Move the whole loop into Rust.
  • Byte versus UTF-16 positions. Error columns differ for non-ASCII input. Convert before reporting.
  • Copying every string out. Return offsets and decode lazily or on demand.
  • Changing the error type. Callers catch SyntaxError. Keep the same type and fields.
  • Testing only the Wasm build. Run golden tests natively too, for fast iteration.

Performance note

On a 40 MB log file in Chrome: the original JavaScript parser took 2.6 s; the Wasm parse call took 0.35 s; encoding and copying the input 0.08 s; rebuilding objects with lazy messages 0.27 s — 0.7 s end to end, a 3.7× improvement. A token-by-token port of the same parser took 3.1 s.

Parsing a 40 MB log file in Chrome Seconds to parse a forty-megabyte log file with the original JavaScript parser, a token-by-token Wasm port, and a whole-input Wasm port with a compatible wrapper. seconds end to end original JavaScript 2.6 s token-by-token Wasm port 3.1 s whole-input Wasm port 0.7 s

Frequently Asked Questions

Should the parser return JSON instead of a flat buffer? JSON is simpler and JSON.parse is fast; flat buffers win when there are many small records or numeric fields. Measure both on real inputs.

Can I use a Rust parser library? Yes — nom, logos or winnow work well in Wasm; check their binary size impact.

How do I keep the old parser for comparison? Keep it in the repository behind a flag and run both on the same inputs, as described in the feature-flag guide.

What about streaming input? Expose a stateful parser that accepts chunks and returns completed records, with the same flat format.

Does this approach work for C or Zig? Yes. The boundary design — one call in, one compact result out — matters more than the language.

← Back to Porting JavaScript Hot Paths to Wasm