Benchmarking Wasm on Mobile Devices
This page answers one task: a WebAssembly feature is fast on a laptop, but most users are on phones, and desktop throttling only approximates them — so you want to run benchmarks on real Android and iOS devices and get numbers that are repeatable enough to compare builds.
Prerequisites
- [ ] At least one mid-range Android phone and one iPhone, ideally matching your users’ devices from analytics.
- [ ] A USB cable and Chrome on the desktop (for Android) or Safari on a Mac (for iOS).
- [ ] A benchmark page that runs the workload and reports results, served from your network or a staging host.
Why phones need their own measurements
Mobile CPUs differ from desktop CPUs in more than speed. They mix high-performance and efficiency cores, and the operating system decides which core runs a thread; they throttle aggressively when warm; they have smaller caches and less memory bandwidth; and their browsers impose lower memory limits. WebAssembly engines on mobile also make different tier-up trade-offs to save power. A workload that is 3× faster in Wasm than JavaScript on a desktop might be 2× or 4× on a phone, and absolute times are often 3–8× slower than on a laptop.
The difficulty is repeatability. A phone that ran a benchmark a minute ago is warmer and may be slower; the same phone on battery saver is slower still; background apps and notifications steal time. Without control, run-to-run variation on phones easily reaches 20–30%, which hides real differences between builds. Most of the work in mobile benchmarking is removing that variation.
Step 1 — prepare the device
Before each session: charge the phone above 80% and keep it plugged in, or run on battery consistently — but do not mix; disable low-power and battery saver modes; close other apps; enable “Do not disturb”; and let the phone cool to room temperature, out of its case. Use the same browser version for compared runs, and note it. On Android, developer options include “Stay awake” while charging, which avoids the screen locking mid-run. Keep the screen on and the benchmark page in the foreground: background tabs are heavily throttled.
Step 2 — connect remote debugging
For Android, enable USB debugging, connect, and open chrome://inspect on the desktop to inspect the phone’s Chrome tab — the console, network panel and
Performance panel all work against the device. For iOS, enable Web Inspector in Safari’s advanced settings on the phone, connect to a Mac, and open the
Develop menu in desktop Safari. Remote DevTools let you see console output and record traces, but recording a trace adds overhead; use traces to
understand behaviour and plain timings to compare builds.
Step 3 — structure runs to control heat
Run the workload in short bursts with pauses, rather than in a tight loop for minutes:
async function bench(fn, { warmup = 10, runs = 15, pauseMs = 3000 } = {}) {
for (let i = 0; i < warmup; i++) fn();
const times = [];
for (let i = 0; i < runs; i++) {
const t0 = performance.now();
fn();
times.push(performance.now() - t0);
await new Promise((r) => setTimeout(r, pauseMs)); // let the SoC cool
}
times.sort((a, b) => a - b);
return { median: times[runs >> 1], min: times[0], max: times[runs - 1] };
}
Report the median, and keep min and max to see the spread. If later runs are consistently slower than earlier ones, the device is throttling; lengthen pauses or shorten bursts until the sequence is flat. Interleave the builds you compare — A, B, A, B — so both see the same thermal conditions.
Step 4 — collect results automatically
Copying numbers from a phone screen is error-prone. Have the benchmark page post results to a small collector — a local HTTP endpoint or a spreadsheet webhook — with the device model, browser version, build identifier and timestamp:
const result = await bench(() => wasm.process(input));
await fetch("http://192.168.1.20:8787/results", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ build: BUILD_ID, ua: navigator.userAgent, cores: navigator.hardwareConcurrency, ...result }),
});
With results in one place, comparing builds and devices becomes a query rather than a transcription exercise.
Step 5 — compare with desktop and field data
Run the same benchmark page on a desktop with and without CPU throttling and compare with the phones. The ratio tells you how well desktop throttling predicts each device: if a 4× throttled desktop matches your mid-range Android phone within 10–15%, you can use throttled desktop runs for day-to-day work and check phones before releases. Compare both with real-user monitoring data, which shows the distribution across all devices; lab devices cover specific points in that distribution, as described in measuring Wasm performance with real-user monitoring.
Threads, SIMD and memory on mobile
Feature support and behaviour vary more on mobile. Threads need cross-origin isolation and SharedArrayBuffer, available in current mobile Chrome and
Safari but with different numbers of usable cores — navigator.hardwareConcurrency reports logical cores, some of which are efficiency cores that run
Wasm threads far slower. A thread pool sized to all cores may run slower than one sized to the performance cores. SIMD is supported in current engines;
measure its benefit on device, since narrower or slower vector units change the gain. Memory limits are lower, especially on iOS, where large memories can
cause the tab to be killed. Benchmark the realistic largest input on the smallest supported device, not only the typical input.
Device labs and cloud devices
Owning a few phones covers the essentials. For wider coverage, cloud device services run tests on real devices remotely, and some offer performance profiling. They are useful for checking that results generalise across chip vendors, but device state — temperature, background activity — is less controlled than on a phone on your desk, so use them for breadth and your own devices for precise comparisons. Keep a record of each device’s model, operating system version and browser version with every result, because browser updates on phones arrive automatically and can change performance.
Measuring startup as well as throughput
Throughput benchmarks run after the module is warm, but on phones the first seconds often matter more. Measure the cold path separately: clear the
cache, load the page, and record download, compile and instantiate times plus the first call, using performance.mark around each phase. Mobile engines
compile more lazily and tier up later than desktop engines to save power, so the first calls can be several times slower than steady state, and code
caching behaves differently across repeat visits. Report cold and warm numbers side by side; a build that improves steady-state throughput by 10% but adds
200 ms of compile time on a mid-range phone may be a net loss for users who use the feature once per visit.
Writing up mobile results
Mobile numbers are easy to misread without context. Every result should carry the device model, operating system and browser version, the power state, whether threads and SIMD were enabled, the input size and the build identifier, plus the spread, not only the median. When presenting a comparison, show each device separately rather than averaging across devices — an average of a flagship and a low-end phone describes neither — and state which device represents the audience the decision is for. Keeping this format consistent makes results from different weeks and different people comparable, which is what turns occasional measurements into a performance history.
Expected output
On a mid-range Android phone, the Wasm filter runs at a median of 43 ms (spread ±3 ms) with pauses, against 9 ms on the laptop and 38 ms on the laptop with 4× throttling; on an iPhone, 21 ms; results for both builds sit in the collector with device and browser details; and a thread pool sized to four workers beats one sized to eight on the Android phone.
Gotchas
- Long continuous loops. Thermal throttling distorts later runs. Pause between runs.
- Mixing power states. Battery saver and charging change performance. Keep them constant.
- Background tabs. Timers and execution are throttled. Keep the page in front.
- Profiling overhead in comparisons. Traces slow the run. Compare plain timings.
- Sizing thread pools by
hardwareConcurrency. Efficiency cores are slow. Measure pool sizes.
Performance note
On the same mid-range Android phone, run-to-run spread fell from ±24% with continuous loops and no device preparation to ±4% with preparation, pauses and interleaving — small enough to detect a 10% regression reliably.
Frequently Asked Questions
Is desktop CPU throttling good enough? For day-to-day work, often yes once calibrated against real devices; check phones before releases.
Should I benchmark in a WebView? If your app runs in one, yes — WebViews can lag behind the browser’s engine version.
How many devices do I need? One mid-range Android, one iPhone, and one low-end Android cover most decisions.
Can I automate runs on phones? Yes — remote debugging protocols and cloud device services support scripted runs.
Why is the first call so much slower on phones? Mobile engines compile lazily and tier up later to save power; measure cold and warm paths separately.
Related
- Comparing Wasm performance across browsers fairly — engine comparisons.
- Avoiding JIT warm-up errors in Wasm benchmarks — warm-up rules.
- Simulating slow networks for Wasm loading — loading on phones.
- Building a reproducible Wasm benchmark harness — the harness.
← Back to Wasm Performance Benchmarking