Porting a JavaScript Parser to Rust Wasm
This page answers one task: a JavaScript parser — for a log format, a query language, CSV, a binary record format — is the measured bottleneck, and you want to port it to Rust compiled to WebAssembly while callers keep using exactly the same JavaScript API.
Prerequisites
- [ ] A profile showing the parser dominates the scenario, as in profiling JavaScript to find Wasm candidates.
- [ ] The existing parser’s test suite, or a corpus of real inputs with known outputs.
- [ ] Rust,
wasm-packorwasm-bindgen-cli, and the project’s bundler.
The shape of a good port
A naive port mirrors the JavaScript structure: a Tokenizer class exported from Rust, a next() method called once per token from JavaScript, and a
result object built field by field through the boundary. That design pays a boundary crossing per token and per field, and often ends up slower than
the original. A good port moves the whole loop across the boundary: JavaScript hands over the entire input once, Rust tokenizes and parses everything,
and the result comes back in one step — as a compact binary buffer, a single JSON string, or a small number of typed arrays.
The second design principle is that the public API does not change. Callers import parse(text) and receive the same objects, with the same error
messages and positions, as before. A thin JavaScript wrapper adapts the Wasm module’s efficient interface to the old API, which also makes it possible to
run both implementations side by side during rollout.
Step 1 — pin down the existing behaviour
Before writing Rust, freeze what “correct” means. Collect inputs that exercise every branch of the existing parser — valid documents, every error case, edge cases like empty input, trailing newlines, Unicode, very long lines — and record the current outputs and error messages as golden files:
for (const file of await glob("corpus/**/*.txt")) {
const text = await readFile(file, "utf8");
let out;
try { out = { ok: parseJs(text) }; } catch (e) { out = { error: e.message, line: e.line, col: e.column }; }
await writeFile(file + ".golden.json", JSON.stringify(out, null, 1));
}
These goldens become the port’s acceptance test. Differences between the two implementations are bugs in the port unless you decide otherwise.
Step 2 — write the parser in Rust over bytes
Parse UTF-8 bytes rather than JavaScript strings. Rust’s &str is UTF-8, so the input crosses the boundary as bytes and is validated once:
#[derive(Debug)]
pub struct Record { pub ts: u64, pub level: u8, pub msg_start: u32, pub msg_len: u32 }
pub fn parse(input: &str) -> Result<Vec<Record>, ParseError> {
let mut out = Vec::with_capacity(input.len() / 80);
for (lineno, line) in input.lines().enumerate() {
let (ts, rest) = parse_ts(line).map_err(|c| ParseError::at(lineno + 1, c, "expected timestamp"))?;
let (level, msg) = parse_level(rest).map_err(|c| ParseError::at(lineno + 1, c, "unknown level"))?;
let msg_start = (msg.as_ptr() as usize - input.as_ptr() as usize) as u32;
out.push(Record { ts, level, msg_start, msg_len: msg.len() as u32 });
}
Ok(out)
}
Records refer to message text by offset and length into the input rather than copying strings; the JavaScript wrapper decodes only the messages callers
actually read. Keep the Rust code free of wasm-bindgen types so it can be tested natively with cargo test against the same golden files.
Step 3 — design a one-call boundary
Export a single function that takes the input and returns a flat result. With wasm-bindgen, returning a Vec<u32> copies one typed array back:
#[wasm_bindgen]
pub fn parse_flat(input: &str) -> Result<Vec<u32>, JsError> {
let recs = parse(input).map_err(|e| JsError::new(&format!("{}:{}:{}", e.line, e.col, e.msg)))?;
let mut flat = Vec::with_capacity(recs.len() * 5);
for r in recs {
flat.extend([(r.ts >> 32) as u32, r.ts as u32, r.level as u32, r.msg_start, r.msg_len]);
}
Ok(flat)
}
The wrapper rebuilds the familiar objects, lazily decoding message text from the original string’s bytes:
import { parse_flat } from "./pkg/logparse.js";
const encoder = new TextEncoder(), decoder = new TextDecoder();
export function parse(text) {
const bytes = encoder.encode(text);
const flat = parse_flat(text); // one boundary call
const out = new Array(flat.length / 5);
for (let i = 0, j = 0; i < flat.length; i += 5, j++) {
const s = flat[i + 3], n = flat[i + 4];
out[j] = {
ts: flat[i] * 2 ** 32 + flat[i + 1],
level: LEVELS[flat[i + 2]],
get message() { return decoder.decode(bytes.subarray(s, s + n)); },
};
}
return out;
}
The getter means callers that never read message never pay to decode it. If callers spread or serialise records, replace the getter with eager decoding
— measure both.
Step 4 — match errors exactly
Error behaviour is part of the API. The original parser threw SyntaxError objects with line and column; the port must too. Encode the position in
the Rust error and translate it in the wrapper:
try { return parseInternal(text); }
catch (e) {
const [line, col, ...msg] = String(e.message).split(":");
const err = new SyntaxError(msg.join(":")); err.line = +line; err.column = +col;
throw err;
}
Columns deserve care: JavaScript counted UTF-16 code units, Rust counts bytes. Convert byte offsets to UTF-16 columns for the error line before reporting, or golden tests with non-ASCII input will disagree.
Step 5 — measure every stage
Time the old parser, the Wasm call alone, and the full wrapper including encoding and object rebuilding, on production-sized inputs. The full wrapper is what callers experience. If object rebuilding dominates, the remaining win may lie in letting performance-sensitive callers use the flat result directly through a second, lower-level API, while keeping the compatible one.
Sharing code between the browser and the server
A parser often runs in more than one place: in the browser for previews and validation, on the server for ingestion. A Rust core gives you one
implementation for both — compiled to WebAssembly for the browser and Node, natively for a Rust server, or as a Wasm module inside a Node or Deno backend.
Keep the core crate free of browser-specific code and put the wasm-bindgen exports in a thin separate crate, as described in
using Cargo features for Wasm and native builds.
The golden tests then run against every build: cargo test for the native core, and a JavaScript test for the Wasm wrapper. That guarantees the browser
and the server agree, which was rarely true when they had separate JavaScript and server-language implementations.
When the port is not enough
Sometimes the port lands and the end-to-end gain is smaller than the profile suggested. The usual reasons are visible in the stage timings: encoding the
input to UTF-8 and copying it in takes noticeable time for very large strings; rebuilding millions of JavaScript objects costs as much as the parse; or the
data arrived as a string only because an earlier step decoded bytes that could have been passed directly. Each has a fix — take Uint8Array input from
fetch or File instead of a string, expose the flat result to callers that can use it, or stream input in chunks. The profile after the port is as
important as the one before it, and the porting decision is about the whole pipeline, not just the parse function.
Expected output
All golden tests pass against the Rust parser natively and through the Wasm wrapper; the public parse() API is unchanged; and parsing a 40 MB log
file drops from 2.6 s to 0.7 s end to end in Chrome, with the Wasm call itself taking 0.35 s.
Gotchas
- Token-by-token APIs. Boundary crossings swamp the gain. Move the whole loop into Rust.
- Byte versus UTF-16 positions. Error columns differ for non-ASCII input. Convert before reporting.
- Copying every string out. Return offsets and decode lazily or on demand.
- Changing the error type. Callers catch
SyntaxError. Keep the same type and fields. - Testing only the Wasm build. Run golden tests natively too, for fast iteration.
Performance note
On a 40 MB log file in Chrome: the original JavaScript parser took 2.6 s; the Wasm parse call took 0.35 s; encoding and copying the input 0.08 s; rebuilding objects with lazy messages 0.27 s — 0.7 s end to end, a 3.7× improvement. A token-by-token port of the same parser took 3.1 s.
Frequently Asked Questions
Should the parser return JSON instead of a flat buffer?
JSON is simpler and JSON.parse is fast; flat buffers win when there are many small records or numeric fields. Measure both on real inputs.
Can I use a Rust parser library?
Yes — nom, logos or winnow work well in Wasm; check their binary size impact.
How do I keep the old parser for comparison? Keep it in the repository behind a flag and run both on the same inputs, as described in the feature-flag guide.
What about streaming input? Expose a stateful parser that accepts chunks and returns completed records, with the same flat format.
Does this approach work for C or Zig? Yes. The boundary design — one call in, one compact result out — matters more than the language.
Related
- Keeping JavaScript and Wasm results identical — the parity tests.
- Shipping a ported module behind a feature flag — rolling it out safely.
- Choosing between JSON and binary serialization — result formats.
- Encoding strings across the Wasm boundary — text input costs.
← Back to Porting JavaScript Hot Paths to Wasm