Measuring JS-to-Wasm Call Overhead

This page answers one question: how much does crossing the JavaScript–WebAssembly boundary cost per call, and at what granularity does that cost stop mattering?

Prerequisites

  • [ ] A module with a few trivial exports you can call in a loop: an empty function, one that adds two numbers, one that takes a string.
  • [ ] A warmed-up benchmark loop, as in avoiding JIT warm-up errors in Wasm benchmarks.
  • [ ] Chrome or Node with a recent V8, and optionally Firefox for comparison.

What a boundary crossing involves

A call from JavaScript into a WebAssembly export is not free, but it is far cheaper than its reputation. For an export that takes and returns numbers, the engine converts each JavaScript value to the Wasm type — a Number to i32 or f64, a BigInt to i64 — switches to the Wasm calling convention, runs the function, and converts the result back. Modern engines compile a specialised wrapper for each export signature and, once the calling JavaScript is optimized, can inline that wrapper into the caller. The result is a crossing cost measured in nanoseconds.

The cost that dominates in practice is not the crossing itself but what has to happen to data at the boundary. Numbers pass in registers. Strings, arrays and objects do not exist in WebAssembly, so they must be copied into linear memory — encoded, allocated, written — before the call and often read back and freed after it. That work scales with the size of the data and is done by glue code, usually generated by wasm-bindgen or Emscripten. A “slow boundary” is almost always slow marshalling.

Where the time goes in one call carrying a string The layers of work for a single call from JavaScript into a Wasm export that takes a string. The crossing itself is a few nanoseconds; encoding the string into linear memory and the allocator calls cost far more. JS argument checks in glue type checks, length — tens of ns allocate in linear memory __wbindgen_malloc call — tens of ns encode UTF-16 → UTF-8 TextEncoder.encodeInto — scales with length the crossing itself wrapper + call convention — a few ns free after use __wbindgen_free — tens of ns

Step 1 — measure the baseline: an empty call

Export functions that do nothing, so the measurement is pure overhead:

#[wasm_bindgen] pub fn noop() {}
#[wasm_bindgen] pub fn add(a: i32, b: i32) -> i32 { a + b }
#[wasm_bindgen] pub fn add_f64(a: f64, b: f64) -> f64 { a + b }
#[wasm_bindgen] pub fn str_len(s: &str) -> usize { s.len() }
#[wasm_bindgen] pub fn sum(v: &[f64]) -> f64 { v.iter().sum() }

Time each in a batched loop, so timer resolution does not dominate, and compare against an equivalent JavaScript function so the loop’s own cost is visible:

function nsPerCall(fn, n = 5_000_000) {
  for (let i = 0; i < 200_000; i++) fn(i);           // warm up
  const t0 = performance.now();
  let acc = 0;
  for (let i = 0; i < n; i++) acc += fn(i) | 0;
  const ns = ((performance.now() - t0) * 1e6) / n;
  if (acc === 42) console.log("");                     // keep acc alive
  return ns.toFixed(2);
}

const jsAdd = (a, b) => a + b;
console.table({
  "js add": nsPerCall((i) => jsAdd(i, 1)),
  "wasm noop": nsPerCall(() => wasm.noop()),
  "wasm add": nsPerCall((i) => wasm.add(i, 1)),
  "wasm add_f64": nsPerCall((i) => wasm.add_f64(i, 1.5)),
});

Step 2 — measure data-carrying calls

Repeat with strings and arrays of different sizes. This is where the numbers change by orders of magnitude:

const short = "hello";
const long = "x".repeat(10_000);
const small = new Float64Array(8);
const big = new Float64Array(100_000);

console.table({
  "str_len(5 chars)": nsPerCall(() => wasm.str_len(short), 1_000_000),
  "str_len(10k chars)": nsPerCall(() => wasm.str_len(long), 20_000),
  "sum(8 f64)": nsPerCall(() => wasm.sum(small), 1_000_000),
  "sum(100k f64)": nsPerCall(() => wasm.sum(big), 2_000),
});

The array case is instructive. Passing a Float64Array to a function taking &[f64] makes the glue allocate in linear memory and copy the whole array in; for 100,000 elements that copy is most of the call’s time. The fix is not a faster boundary but no copy — see zero-copy data transfer patterns.

Cost per call by argument type Measured nanoseconds per call from optimized JavaScript into Wasm exports in Chrome, after warm-up. Number arguments cost a few nanoseconds; strings and arrays cost what the copy into linear memory costs. nanoseconds per call (Chrome, warmed up) JS function (reference) 0.9 ns Wasm noop 2.1 ns Wasm add(i32, i32) 2.4 ns str_len 5 chars 61 ns sum 8 × f64 74 ns str_len 10k chars 3,400 ns

Step 3 — find the break-even granularity

Overhead only matters relative to the work done per call. The useful question is: how much work must a call do before the crossing is a small fraction of its cost? Measure a real kernel with varying amounts of work per call — for example processing an image one pixel per call, one row per call, and the whole image per call:

const w = 1920, h = 1080;
const perPixel = () => { for (let i = 0; i < w * h; i++) wasm.invert_pixel(ptr + i * 4); };
const perRow = () => { for (let y = 0; y < h; y++) wasm.invert_row(ptr + y * w * 4, w); };
const whole = () => wasm.invert_image(ptr, w, h);

Per-pixel calls cost two million crossings per frame; even at 2 ns each, that is 4 ms of pure overhead, larger than the work itself. Per-row calls cost 1,080 crossings — microseconds. A single call costs one. The rule of thumb that falls out of such measurements: once a call does at least a microsecond of work, the crossing is under one percent of its cost.

Moving the loop across the boundary Calling Wasm once per pixel pays the crossing millions of times per frame. Calling once per row pays it about a thousand times. Calling once per image with a pointer to pixels already in linear memory pays it once, and the loop runs entirely inside Wasm. per pixel 2,073,600 calls / frame per row 1,080 calls / frame per image 1 call / frame The work is identical in all three; only where the loop lives changes.

Step 4 — reduce glue work on chatty boundaries

When an interface genuinely needs many calls — an event handler per keystroke, a callback per item — make each call as cheap as possible. Pass numbers instead of strings where an enum or an index will do. Keep large data resident in linear memory and pass offsets, so nothing is copied per call. Prefer &str arguments over String where the function only reads, since the glue can then free the temporary immediately. For JavaScript objects passed by reference, wasm-bindgen’s --reference-types mode uses externref and avoids a heap-slab lookup per object. Each of these removes glue work, which is where the time was.

Step 5 — measure in the context that matters

Microbenchmarks like these run the boundary in its best case: a hot loop the optimizer can see entirely. Real applications call from code that may not be optimized, from callbacks with polymorphic arguments, and from frameworks with their own overhead. The crossing in such code can be several times more expensive than in the microbenchmark, though still usually small next to the data-marshalling costs. Use the microbenchmark to understand the shape of the costs, and a profile of the real application — see profiling Wasm with the Chrome Performance panel — to decide whether the boundary is actually where your time goes.

Interpreting the numbers

Two observations from the table generalise well beyond this example. First, the crossing itself is in the same order of magnitude as a JavaScript function call — a few nanoseconds — so the folk wisdom that “calling Wasm is expensive” is about glue and data, not the call. Second, the cost of a data-carrying call is roughly proportional to the bytes moved, which means it can be predicted: if a call copies a megabyte, expect it to cost about what copying a megabyte costs, regardless of how fast the function on the other side is.

That second point gives a quick sanity check for any interface. Multiply the bytes each call moves by the number of calls per second the application makes, and compare it with memory bandwidth. An interface that moves 50 MB per second through copies is spending real time on them; one that moves a few kilobytes is not, however chatty it looks.

Expected output

┌──────────────────────┬────────┐
│ js add               │ '0.92' │
│ wasm noop            │ '2.08' │
│ wasm add             │ '2.37' │
│ wasm add_f64         │ '2.41' │
│ str_len(5 chars)     │ '61.20'│
│ str_len(10k chars)   │ '3398' │
│ sum(8 f64)           │ '73.80'│
│ sum(100k f64)        │ '71240'│
└──────────────────────┴────────┘

Gotchas

  • The loop is optimized away. A call whose result is never used may be eliminated. Accumulate results and use them.
  • Comparing against cold JavaScript. The JavaScript reference must be warmed up as thoroughly as the Wasm side.
  • Measuring with DevTools open. The profiler and console instrumentation change call costs significantly. Close DevTools.
  • Assuming numbers carry over between engines. Firefox and Safari have different wrapper strategies. Measure the engines you ship to before drawing conclusions about chatty designs.

Performance note

The difference between the per-pixel and whole-image versions of the invert kernel was 9.1 ms against 0.7 ms per frame — almost entirely boundary overhead. No optimization flag or faster engine could have closed that gap; moving the loop into Wasm did. Interface design decides boundary cost far more than the boundary itself.

Frequently Asked Questions

Are calls from Wasm into JavaScript as cheap? Imported JavaScript functions are called through a similar wrapper, and numbers are similarly cheap. Calls that touch the DOM pay for the DOM operation, which dwarfs the crossing.

Does the component model make the boundary faster? Its canonical ABI standardises how strings and lists are copied, which can reduce glue code, but data still has to be copied between separate memories. See understanding the canonical ABI.

Is i64 slower than i32 at the boundary? Yes — i64 values are converted to and from BigInt, which costs an allocation. Avoid i64 in hot interfaces where 53-bit precision suffices; see passing 64-bit integers with BigInt.

Should every hot path be a single call? Coarse calls are the right default, but not at the cost of copying large data in and out each time. Combine coarse calls with data that stays resident in linear memory.

← Back to Wasm Performance Benchmarking