Testing Wasm Performance Budgets in CI

This page answers one task: a WebAssembly module exists because it is fast, but performance regressions slip in unnoticed — a dependency upgrade, a changed compiler flag, an innocent refactor — and you want CI to fail when a hot function becomes meaningfully slower, without failing randomly on noise.

Prerequisites

  • [ ] A benchmark harness for the module’s hot functions (Criterion for Rust, or a JavaScript harness).
  • [ ] CI that can run benchmarks, ideally on consistent runners.
  • [ ] A small set of representative inputs checked into the repository.

Why performance tests are harder than correctness tests

A correctness test has a fixed answer. A performance measurement varies from run to run, and on shared CI runners it varies a lot: noisy neighbours, CPU frequency scaling, different hardware across runs, and JIT tiering in the engine all move timings by 5–30%. A naive budget — “fail if decode takes over 12 ms” — either sits so far above normal that it catches nothing, or so close that it fails every other run. Teams then disable it, and regressions return.

Three techniques make budgets reliable. Measure something stable where possible: instruction counts in a deterministic runtime vary by less than 1% regardless of machine load. Compare relatively: run the main branch and the change on the same runner in the same job, and fail on the ratio rather than an absolute number. And set thresholds from observed noise, not intuition, so a failure means a real regression.

Wall-clock versus instruction-count budgets Wall-clock timings reflect what users feel but vary 5 to 30 percent on shared CI runners and need relative comparison and generous thresholds. Instruction counts in a deterministic runtime vary under 1 percent, allowing tight budgets, but miss effects such as cache behaviour and engine tiering. wall-clock time what users experience 5–30% noise on shared runners needs relative comparison realistic but noisy instruction count deterministic, under 1% noise tight thresholds possible misses cache and JIT effects stable proxy

Step 1 — choose what to budget

Budget the operations users wait for, not everything: the codec’s decode of a typical file, the parser on a large document, the physics step at the expected entity count, module instantiation. Each budget needs a fixed input, a clear operation boundary and an owner who decides what a regression means. Start with three to five; a dashboard of fifty microbenchmarks produces alerts nobody reads.

Step 2 — measure instruction counts for stable budgets

Running the module in a runtime with deterministic fuel or instruction counting gives a number that does not depend on runner load. Wasmtime’s fuel mechanism counts executed operations approximately; cachegrind or iai-callgrind counts native instructions when benchmarking the native build of the same Rust code:

// benches/iai.rs — instruction counts with iai-callgrind (native build of the core)
use iai_callgrind::{library_benchmark, library_benchmark_group, main};

#[library_benchmark]
fn decode_typical() -> usize {
    let input = include_bytes!("../fixtures/typical.bin");
    codec::decode(std::hint::black_box(input)).unwrap().len()
}
library_benchmark_group!(name = codec; benchmarks = decode_typical);
main!(library_benchmark_groups = codec);
// Fuel-based counting of the actual Wasm module under Wasmtime
let mut config = wasmtime::Config::new();
config.consume_fuel(true);
let engine = wasmtime::Engine::new(&config)?;
let mut store = wasmtime::Store::new(&engine, ());
store.set_fuel(u64::MAX)?;
// ... instantiate and call decode ...
let used = u64::MAX - store.get_fuel()?;
println!("decode fuel: {used}");

Fuel counts on the real .wasm catch regressions caused by Wasm-specific code generation, such as a lost SIMD path or a changed wasm-opt level, that native counts miss.

Step 3 — compare wall-clock time relative to the base branch

For time in a real engine, run both versions in the same job. Build the base branch’s module and the pull request’s module, then run each benchmark interleaved — A, B, A, B — many times in the same headless browser or Node process, and compare medians:

// bench/compare.mjs — interleaved runs of base and head modules in one process
import { load } from "./harness.mjs";
const base = await load("base/codec_bg.wasm"), head = await load("head/codec_bg.wasm");
const input = await readFixture("typical.bin");
for (let i = 0; i < 50; i++) { base.decode(input); head.decode(input); }   // warm-up both
const tb = [], th = [];
for (let i = 0; i < 200; i++) {
  tb.push(time(() => base.decode(input)));
  th.push(time(() => head.decode(input)));
}
const ratio = median(th) / median(tb);
console.log(`decode ratio head/base: ${ratio.toFixed(3)}`);
if (ratio > 1.10) process.exit(1);

Interleaving exposes both versions to the same machine conditions, so the ratio is far more stable than either absolute time. Warm-up matters because engines tier up hot functions; see avoiding JIT warm-up errors in Wasm benchmarks.

A relative performance check in a pull request CI builds the module from the base branch and from the pull request, loads both in one process, warms both up, runs them interleaved many times on the same input, computes the ratio of medians, and fails the check if the ratio exceeds the threshold derived from measured noise. build base + head same toolchain load both one process warm up both tier-up done interleaved runs A B A B … ratio of medians fail above threshold

Step 4 — set thresholds from measured noise

Before enforcing a budget, run the comparison of the main branch against itself twenty or thirty times on the CI runners and record the ratios. The spread of those ratios is your noise floor. Set the failure threshold comfortably above it — if self-comparisons range from 0.97 to 1.04, fail above 1.10 and warn above 1.06. For instruction counts, a threshold of 1–2% is usually safe. Revisit thresholds when runners change.

Step 5 — report clearly and allow intentional regressions

A failing budget should say which benchmark regressed, by how much, and against which base commit, in the pull request itself — a comment or a check summary. Some regressions are intentional: a correctness fix or a security hardening that costs time. Provide an explicit escape hatch — a label such as perf-regression-accepted or an updated baseline file committed with a justification — so the budget records the decision instead of being disabled.

Budgets beyond speed

The same machinery covers other numbers that regress silently. Module size, compressed size, peak linear memory for a representative workload, and instantiation time are cheap to measure and stable. Size budgets are deterministic and belong next to performance budgets; see catching size regressions in CI. Peak memory is deterministic for a given input and allocator, and a jump usually means a leak or a changed algorithm.

Running on dedicated hardware

When wall-clock accuracy matters — a codec whose performance is the product — shared runners are the wrong place for absolute measurements. A small dedicated machine or self-hosted runner, with CPU frequency scaling disabled and nothing else scheduled, reduces noise to a few percent and makes absolute budgets possible. Use it for nightly runs that track the trend over time, and keep the relative check on shared runners for pull requests. Results over time belong in a dashboard, as described in tracking benchmark results in CI.

Investigating a failed budget

A budget failure is the start of an investigation, not the end. First reproduce locally with the same comparison script: if the ratio is far above the threshold locally too, the regression is real. Then narrow it down. Compare the two modules’ sizes and function counts with twiggy diff or wasm-objdump — a function that doubled in size or a missing SIMD path often explains everything. Check the build flags in both builds, since many regressions come from a changed profile, a dropped -C target-feature=+simd128, or wasm-opt running at a different level after a tooling update. If the code itself changed, profile both versions with the browser’s profiler or perf on the native build, and compare where time moved. Record the cause in the pull request; over time those notes show which kinds of change tend to cost performance, which is useful when reviewing.

Keeping benchmark inputs representative

Budgets are only as good as their inputs. A decode benchmark on a tiny synthetic file measures setup overhead; one on a pathological file measures a path users rarely hit. Choose inputs from real usage — anonymised samples of typical sizes and contents — and keep a small set covering the common case and the expensive-but-realistic case. When the product’s typical input changes, update the fixtures and reset the baseline deliberately in a separate commit, so a change in what is measured is never confused with a change in performance.

Expected output

Each pull request runs instruction-count benchmarks for four operations (threshold 2%) and an interleaved wall-clock comparison in headless Chromium (threshold 10%); self-comparison noise on the main branch stays within ±4%; a change that dropped SIMD from the build failed both checks with a 1.8× ratio; and an accepted regression is recorded in the baseline file with its justification.

Gotchas

  • Absolute timings on shared runners. They flake. Compare relatively in one job.
  • No warm-up. Tier-up skews the first runs. Warm up both versions.
  • Thresholds from intuition. Measure self-comparison noise first.
  • Too many benchmarks. Alerts get ignored. Budget the operations users feel.
  • No escape hatch. Teams disable the check. Allow recorded, justified regressions.

Performance note

On GitHub-hosted runners, absolute wall-clock times for the decode benchmark varied by ±18% across runs, while the interleaved head/base ratio varied by ±4% and fuel counts by under 0.1%.

Run-to-run variation of three measurements on shared runners Percentage spread across 30 runs of the same commit for absolute wall-clock time, the ratio of interleaved head and base timings, and Wasmtime fuel counts. run-to-run spread (±%) absolute wall-clock time 18 % interleaved ratio 4 % fuel count 0.1 %

Frequently Asked Questions

Are instruction counts enough on their own? No — they miss cache effects, memory bandwidth and engine tiering. Pair them with a relative wall-clock check.

Which engine should the wall-clock check use? The one most of your users run; add others for nightly runs.

How many iterations are needed? Enough that the self-comparison spread is stable — usually a few hundred short runs or a few dozen long ones.

Should budgets block merges? Instruction-count budgets can; wall-clock checks often start as warnings until thresholds are trusted.

What should I check first when a budget fails? Reproduce the ratio locally, then compare module sizes, function sizes and build flags between the base and head builds before profiling.

← Back to Testing & Verifying Wasm Builds