Benchmarking Wasm on Mobile Devices

This page answers one task: a WebAssembly feature is fast on a laptop, but most users are on phones, and desktop throttling only approximates them — so you want to run benchmarks on real Android and iOS devices and get numbers that are repeatable enough to compare builds.

Prerequisites

  • [ ] At least one mid-range Android phone and one iPhone, ideally matching your users’ devices from analytics.
  • [ ] A USB cable and Chrome on the desktop (for Android) or Safari on a Mac (for iOS).
  • [ ] A benchmark page that runs the workload and reports results, served from your network or a staging host.

Why phones need their own measurements

Mobile CPUs differ from desktop CPUs in more than speed. They mix high-performance and efficiency cores, and the operating system decides which core runs a thread; they throttle aggressively when warm; they have smaller caches and less memory bandwidth; and their browsers impose lower memory limits. WebAssembly engines on mobile also make different tier-up trade-offs to save power. A workload that is 3× faster in Wasm than JavaScript on a desktop might be 2× or 4× on a phone, and absolute times are often 3–8× slower than on a laptop.

The difficulty is repeatability. A phone that ran a benchmark a minute ago is warmer and may be slower; the same phone on battery saver is slower still; background apps and notifications steal time. Without control, run-to-run variation on phones easily reaches 20–30%, which hides real differences between builds. Most of the work in mobile benchmarking is removing that variation.

A repeatable mobile benchmark run Prepare the device by charging, cooling, disabling battery saver and closing apps. Connect over USB with remote debugging. Load the benchmark page, warm up, run several measured iterations with cool-down pauses between them, and send results to a collector automatically. prepare device charged, cool, no saver remote debugging USB + DevTools / Web Inspector warm up workload engine tier-up measured runs with cool-down pauses post results collector endpoint

Step 1 — prepare the device

Before each session: charge the phone above 80% and keep it plugged in, or run on battery consistently — but do not mix; disable low-power and battery saver modes; close other apps; enable “Do not disturb”; and let the phone cool to room temperature, out of its case. Use the same browser version for compared runs, and note it. On Android, developer options include “Stay awake” while charging, which avoids the screen locking mid-run. Keep the screen on and the benchmark page in the foreground: background tabs are heavily throttled.

Step 2 — connect remote debugging

For Android, enable USB debugging, connect, and open chrome://inspect on the desktop to inspect the phone’s Chrome tab — the console, network panel and Performance panel all work against the device. For iOS, enable Web Inspector in Safari’s advanced settings on the phone, connect to a Mac, and open the Develop menu in desktop Safari. Remote DevTools let you see console output and record traces, but recording a trace adds overhead; use traces to understand behaviour and plain timings to compare builds.

Step 3 — structure runs to control heat

Run the workload in short bursts with pauses, rather than in a tight loop for minutes:

async function bench(fn, { warmup = 10, runs = 15, pauseMs = 3000 } = {}) {
  for (let i = 0; i < warmup; i++) fn();
  const times = [];
  for (let i = 0; i < runs; i++) {
    const t0 = performance.now();
    fn();
    times.push(performance.now() - t0);
    await new Promise((r) => setTimeout(r, pauseMs));     // let the SoC cool
  }
  times.sort((a, b) => a - b);
  return { median: times[runs >> 1], min: times[0], max: times[runs - 1] };
}

Report the median, and keep min and max to see the spread. If later runs are consistently slower than earlier ones, the device is throttling; lengthen pauses or shorten bursts until the sequence is flat. Interleave the builds you compare — A, B, A, B — so both see the same thermal conditions.

Throttling in a long benchmark loop on a phone In a continuous loop, per-run time starts at 42 milliseconds, rises as the chip heats, and settles near 61 milliseconds after about two minutes. With three-second pauses between runs, times stay near 42 to 44 milliseconds throughout. 0 s 42 ms (cool) 45 s 47 ms 90 s 55 ms 150 s 61 ms (throttled) 180 s with pauses: 43 ms

Step 4 — collect results automatically

Copying numbers from a phone screen is error-prone. Have the benchmark page post results to a small collector — a local HTTP endpoint or a spreadsheet webhook — with the device model, browser version, build identifier and timestamp:

const result = await bench(() => wasm.process(input));
await fetch("http://192.168.1.20:8787/results", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ build: BUILD_ID, ua: navigator.userAgent, cores: navigator.hardwareConcurrency, ...result }),
});

With results in one place, comparing builds and devices becomes a query rather than a transcription exercise.

Step 5 — compare with desktop and field data

Run the same benchmark page on a desktop with and without CPU throttling and compare with the phones. The ratio tells you how well desktop throttling predicts each device: if a 4× throttled desktop matches your mid-range Android phone within 10–15%, you can use throttled desktop runs for day-to-day work and check phones before releases. Compare both with real-user monitoring data, which shows the distribution across all devices; lab devices cover specific points in that distribution, as described in measuring Wasm performance with real-user monitoring.

Threads, SIMD and memory on mobile

Feature support and behaviour vary more on mobile. Threads need cross-origin isolation and SharedArrayBuffer, available in current mobile Chrome and Safari but with different numbers of usable cores — navigator.hardwareConcurrency reports logical cores, some of which are efficiency cores that run Wasm threads far slower. A thread pool sized to all cores may run slower than one sized to the performance cores. SIMD is supported in current engines; measure its benefit on device, since narrower or slower vector units change the gain. Memory limits are lower, especially on iOS, where large memories can cause the tab to be killed. Benchmark the realistic largest input on the smallest supported device, not only the typical input.

Device labs and cloud devices

Owning a few phones covers the essentials. For wider coverage, cloud device services run tests on real devices remotely, and some offer performance profiling. They are useful for checking that results generalise across chip vendors, but device state — temperature, background activity — is less controlled than on a phone on your desk, so use them for breadth and your own devices for precise comparisons. Keep a record of each device’s model, operating system version and browser version with every result, because browser updates on phones arrive automatically and can change performance.

Measuring startup as well as throughput

Throughput benchmarks run after the module is warm, but on phones the first seconds often matter more. Measure the cold path separately: clear the cache, load the page, and record download, compile and instantiate times plus the first call, using performance.mark around each phase. Mobile engines compile more lazily and tier up later than desktop engines to save power, so the first calls can be several times slower than steady state, and code caching behaves differently across repeat visits. Report cold and warm numbers side by side; a build that improves steady-state throughput by 10% but adds 200 ms of compile time on a mid-range phone may be a net loss for users who use the feature once per visit.

Writing up mobile results

Mobile numbers are easy to misread without context. Every result should carry the device model, operating system and browser version, the power state, whether threads and SIMD were enabled, the input size and the build identifier, plus the spread, not only the median. When presenting a comparison, show each device separately rather than averaging across devices — an average of a flagship and a low-end phone describes neither — and state which device represents the audience the decision is for. Keeping this format consistent makes results from different weeks and different people comparable, which is what turns occasional measurements into a performance history.

Expected output

On a mid-range Android phone, the Wasm filter runs at a median of 43 ms (spread ±3 ms) with pauses, against 9 ms on the laptop and 38 ms on the laptop with 4× throttling; on an iPhone, 21 ms; results for both builds sit in the collector with device and browser details; and a thread pool sized to four workers beats one sized to eight on the Android phone.

Gotchas

  • Long continuous loops. Thermal throttling distorts later runs. Pause between runs.
  • Mixing power states. Battery saver and charging change performance. Keep them constant.
  • Background tabs. Timers and execution are throttled. Keep the page in front.
  • Profiling overhead in comparisons. Traces slow the run. Compare plain timings.
  • Sizing thread pools by hardwareConcurrency. Efficiency cores are slow. Measure pool sizes.

Performance note

On the same mid-range Android phone, run-to-run spread fell from ±24% with continuous loops and no device preparation to ±4% with preparation, pauses and interleaving — small enough to detect a 10% regression reliably.

Run-to-run spread on a mid-range Android phone Percentage spread of benchmark timings on the same phone with no preparation and a continuous loop, compared with a prepared device, pauses between runs and interleaved builds. run-to-run spread (±%) unprepared, continuous loop 24 % prepared, pauses, interleaved 4 %

Frequently Asked Questions

Is desktop CPU throttling good enough? For day-to-day work, often yes once calibrated against real devices; check phones before releases.

Should I benchmark in a WebView? If your app runs in one, yes — WebViews can lag behind the browser’s engine version.

How many devices do I need? One mid-range Android, one iPhone, and one low-end Android cover most decisions.

Can I automate runs on phones? Yes — remote debugging protocols and cloud device services support scripted runs.

Why is the first call so much slower on phones? Mobile engines compile lazily and tier up later to save power; measure cold and warm paths separately.

← Back to Wasm Performance Benchmarking