Measuring JS-to-Wasm Call Overhead
This page answers one question: how much does crossing the JavaScript–WebAssembly boundary cost per call, and at what granularity does that cost stop mattering?
Prerequisites
- [ ] A module with a few trivial exports you can call in a loop: an empty function, one that adds two numbers, one that takes a string.
- [ ] A warmed-up benchmark loop, as in avoiding JIT warm-up errors in Wasm benchmarks.
- [ ] Chrome or Node with a recent V8, and optionally Firefox for comparison.
What a boundary crossing involves
A call from JavaScript into a WebAssembly export is not free, but it is far cheaper than its reputation. For an export that
takes and returns numbers, the engine converts each JavaScript value to the Wasm type — a Number to i32 or f64, a
BigInt to i64 — switches to the Wasm calling convention, runs the function, and converts the result back. Modern engines
compile a specialised wrapper for each export signature and, once the calling JavaScript is optimized, can inline that wrapper
into the caller. The result is a crossing cost measured in nanoseconds.
The cost that dominates in practice is not the crossing itself but what has to happen to data at the boundary. Numbers pass in registers. Strings, arrays and objects do not exist in WebAssembly, so they must be copied into linear memory — encoded, allocated, written — before the call and often read back and freed after it. That work scales with the size of the data and is done by glue code, usually generated by wasm-bindgen or Emscripten. A “slow boundary” is almost always slow marshalling.
Step 1 — measure the baseline: an empty call
Export functions that do nothing, so the measurement is pure overhead:
#[wasm_bindgen] pub fn noop() {}
#[wasm_bindgen] pub fn add(a: i32, b: i32) -> i32 { a + b }
#[wasm_bindgen] pub fn add_f64(a: f64, b: f64) -> f64 { a + b }
#[wasm_bindgen] pub fn str_len(s: &str) -> usize { s.len() }
#[wasm_bindgen] pub fn sum(v: &[f64]) -> f64 { v.iter().sum() }
Time each in a batched loop, so timer resolution does not dominate, and compare against an equivalent JavaScript function so the loop’s own cost is visible:
function nsPerCall(fn, n = 5_000_000) {
for (let i = 0; i < 200_000; i++) fn(i); // warm up
const t0 = performance.now();
let acc = 0;
for (let i = 0; i < n; i++) acc += fn(i) | 0;
const ns = ((performance.now() - t0) * 1e6) / n;
if (acc === 42) console.log(""); // keep acc alive
return ns.toFixed(2);
}
const jsAdd = (a, b) => a + b;
console.table({
"js add": nsPerCall((i) => jsAdd(i, 1)),
"wasm noop": nsPerCall(() => wasm.noop()),
"wasm add": nsPerCall((i) => wasm.add(i, 1)),
"wasm add_f64": nsPerCall((i) => wasm.add_f64(i, 1.5)),
});
Step 2 — measure data-carrying calls
Repeat with strings and arrays of different sizes. This is where the numbers change by orders of magnitude:
const short = "hello";
const long = "x".repeat(10_000);
const small = new Float64Array(8);
const big = new Float64Array(100_000);
console.table({
"str_len(5 chars)": nsPerCall(() => wasm.str_len(short), 1_000_000),
"str_len(10k chars)": nsPerCall(() => wasm.str_len(long), 20_000),
"sum(8 f64)": nsPerCall(() => wasm.sum(small), 1_000_000),
"sum(100k f64)": nsPerCall(() => wasm.sum(big), 2_000),
});
The array case is instructive. Passing a Float64Array to a function taking &[f64] makes the glue allocate in linear
memory and copy the whole array in; for 100,000 elements that copy is most of the call’s time. The fix is not a faster
boundary but no copy — see zero-copy data transfer patterns.
Step 3 — find the break-even granularity
Overhead only matters relative to the work done per call. The useful question is: how much work must a call do before the crossing is a small fraction of its cost? Measure a real kernel with varying amounts of work per call — for example processing an image one pixel per call, one row per call, and the whole image per call:
const w = 1920, h = 1080;
const perPixel = () => { for (let i = 0; i < w * h; i++) wasm.invert_pixel(ptr + i * 4); };
const perRow = () => { for (let y = 0; y < h; y++) wasm.invert_row(ptr + y * w * 4, w); };
const whole = () => wasm.invert_image(ptr, w, h);
Per-pixel calls cost two million crossings per frame; even at 2 ns each, that is 4 ms of pure overhead, larger than the work itself. Per-row calls cost 1,080 crossings — microseconds. A single call costs one. The rule of thumb that falls out of such measurements: once a call does at least a microsecond of work, the crossing is under one percent of its cost.
Step 4 — reduce glue work on chatty boundaries
When an interface genuinely needs many calls — an event handler per keystroke, a callback per item — make each call as cheap
as possible. Pass numbers instead of strings where an enum or an index will do. Keep large data resident in linear memory and
pass offsets, so nothing is copied per call. Prefer &str arguments over String where the function only reads, since the glue
can then free the temporary immediately. For JavaScript objects passed by reference, wasm-bindgen’s --reference-types mode
uses externref and avoids a heap-slab lookup per object. Each of these removes glue work, which is where the time was.
Step 5 — measure in the context that matters
Microbenchmarks like these run the boundary in its best case: a hot loop the optimizer can see entirely. Real applications call from code that may not be optimized, from callbacks with polymorphic arguments, and from frameworks with their own overhead. The crossing in such code can be several times more expensive than in the microbenchmark, though still usually small next to the data-marshalling costs. Use the microbenchmark to understand the shape of the costs, and a profile of the real application — see profiling Wasm with the Chrome Performance panel — to decide whether the boundary is actually where your time goes.
Interpreting the numbers
Two observations from the table generalise well beyond this example. First, the crossing itself is in the same order of magnitude as a JavaScript function call — a few nanoseconds — so the folk wisdom that “calling Wasm is expensive” is about glue and data, not the call. Second, the cost of a data-carrying call is roughly proportional to the bytes moved, which means it can be predicted: if a call copies a megabyte, expect it to cost about what copying a megabyte costs, regardless of how fast the function on the other side is.
That second point gives a quick sanity check for any interface. Multiply the bytes each call moves by the number of calls per second the application makes, and compare it with memory bandwidth. An interface that moves 50 MB per second through copies is spending real time on them; one that moves a few kilobytes is not, however chatty it looks.
Expected output
┌──────────────────────┬────────┐
│ js add │ '0.92' │
│ wasm noop │ '2.08' │
│ wasm add │ '2.37' │
│ wasm add_f64 │ '2.41' │
│ str_len(5 chars) │ '61.20'│
│ str_len(10k chars) │ '3398' │
│ sum(8 f64) │ '73.80'│
│ sum(100k f64) │ '71240'│
└──────────────────────┴────────┘
Gotchas
- The loop is optimized away. A call whose result is never used may be eliminated. Accumulate results and use them.
- Comparing against cold JavaScript. The JavaScript reference must be warmed up as thoroughly as the Wasm side.
- Measuring with DevTools open. The profiler and console instrumentation change call costs significantly. Close DevTools.
- Assuming numbers carry over between engines. Firefox and Safari have different wrapper strategies. Measure the engines you ship to before drawing conclusions about chatty designs.
Performance note
The difference between the per-pixel and whole-image versions of the invert kernel was 9.1 ms against 0.7 ms per frame — almost entirely boundary overhead. No optimization flag or faster engine could have closed that gap; moving the loop into Wasm did. Interface design decides boundary cost far more than the boundary itself.
Frequently Asked Questions
Are calls from Wasm into JavaScript as cheap? Imported JavaScript functions are called through a similar wrapper, and numbers are similarly cheap. Calls that touch the DOM pay for the DOM operation, which dwarfs the crossing.
Does the component model make the boundary faster? Its canonical ABI standardises how strings and lists are copied, which can reduce glue code, but data still has to be copied between separate memories. See understanding the canonical ABI.
Is i64 slower than i32 at the boundary?
Yes — i64 values are converted to and from BigInt, which costs an allocation. Avoid i64 in hot interfaces where 53-bit
precision suffices; see passing 64-bit integers with BigInt.
Should every hot path be a single call? Coarse calls are the right default, but not at the cost of copying large data in and out each time. Combine coarse calls with data that stays resident in linear memory.
Related
- Measuring Wasm vs JavaScript throughput — whole-workload comparisons.
- Encoding strings across the Wasm boundary — the cost behind string arguments.
- Reading Wasm linear memory with typed arrays — avoiding the copy entirely.
- Is Wasm faster than JavaScript for DOM manipulation? — the worst case for chatty boundaries.
← Back to Wasm Performance Benchmarking