Tracking Benchmark Results in CI
This page answers one task: catch performance regressions in a WebAssembly module automatically, the way tests catch correctness regressions — by running benchmarks in CI, comparing with a baseline, and flagging slowdowns that are real.
Prerequisites
- [ ] A benchmark harness that prints machine-readable results, such as the one from building a reproducible Wasm benchmark harness.
- [ ] A CI system where you can store artifacts or push to a branch — GitHub Actions in the examples.
- [ ] Realistic expectations about CI machines, which are noisy.
The problem with timing on shared runners
Benchmarks in CI face a measurement problem that tests do not. A test either passes or fails; a benchmark produces a number, and on a shared CI runner that number varies from run to run by 5–15% for reasons unrelated to your code: other workloads on the same host, different CPU models in the runner pool, frequency scaling, and cache state. A naive setup that fails the build when a benchmark is 5% slower than last time fails constantly, and teams learn to ignore it within a week.
Three things make CI benchmarking useful despite the noise. Compare against a baseline measured in the same job, on the same machine, rather than against a stored number from another day. Use enough samples and a statistic that resists outliers. And set thresholds based on the measured noise of each benchmark, flagging only changes larger than that. With those in place, CI reliably catches the regressions that matter — a 20% slowdown from an accidentally disabled optimization flag — while ignoring the noise.
Step 1 — make the harness emit JSON
The harness should print one record per benchmark with the statistics you will compare:
// bench/run.mjs — prints JSON lines
import { loadModule, bench } from "./harness.mjs";
const m = await loadModule(process.argv[2]); // path to a built pkg/
for (const [name, fn] of Object.entries(m.benchmarks())) {
const r = await bench(fn, { warmupMs: 500, samples: 60 });
console.log(JSON.stringify({ name, median_ns: r.median * 1e6, p90_ns: r.p90 * 1e6, mad_ns: r.mad * 1e6 }));
}
MAD — the median absolute deviation — is a robust measure of spread, and it is what the threshold will use. Run in Node or a headless browser; Node is faster to start and good enough for tracking relative changes, while a browser run adds coverage of the engine your users have.
Step 2 — build base and head in the same job
# .github/workflows/bench.yml
on: pull_request
jobs:
bench:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with: { fetch-depth: 0 }
- uses: Swatinem/rust-cache@v2
- name: build base and head
run: |
git worktree add ../base ${{ github.event.pull_request.base.sha }}
(cd ../base && wasm-pack build --release --target nodejs --out-dir ../pkg-base)
wasm-pack build --release --target nodejs --out-dir pkg-head
- name: run interleaved
run: |
for round in 1 2 3; do
node bench/run.mjs pkg-base >> base.jsonl
node bench/run.mjs pkg-head >> head.jsonl
done
- run: node bench/compare.mjs base.jsonl head.jsonl > summary.md
- uses: marocchino/sticky-pull-request-comment@v2
with: { path: summary.md }
Building both in the same job is the key decision. Whatever the runner’s speed today, it affects both builds equally, and the comparison measures the change rather than the machine. Interleaving rounds spreads time-dependent noise evenly between them.
Step 3 — compare with a noise-aware threshold
For each benchmark, combine the rounds, take the median of medians, and compare. Flag a change only when it exceeds both a minimum percentage and a multiple of the observed spread:
// bench/compare.mjs
const load = (f) => readFileSync(f, "utf8").trim().split("\n").map(JSON.parse);
const group = (rows) => rows.reduce((m, r) => ((m[r.name] ??= []).push(r), m), {});
const med = (xs) => xs.sort((a, b) => a - b)[xs.length >> 1];
const base = group(load(process.argv[2])), head = group(load(process.argv[3]));
console.log("| benchmark | base | head | change | |\n|---|---:|---:|---:|---|");
for (const name of Object.keys(base)) {
const b = med(base[name].map((r) => r.median_ns)), h = med(head[name].map((r) => r.median_ns));
const noise = 3 * med(base[name].map((r) => r.mad_ns)) / b; // relative
const change = (h - b) / b;
const flag = Math.abs(change) < Math.max(0.05, noise) ? "" : change > 0 ? "🔴 slower" : "🟢 faster";
console.log(`| ${name} | ${(b / 1e3).toFixed(1)} µs | ${(h / 1e3).toFixed(1)} µs | ${(change * 100).toFixed(1)}% | ${flag} |`);
}
The threshold max(5%, 3 × relative MAD) means a noisy benchmark needs a larger change to be flagged than a stable one, which
is exactly what keeps false alarms down.
Step 4 — keep a history on the main branch
Pull-request comparisons catch sudden regressions. Slow drift — one percent per change over a quarter — needs history. On each push to the main branch, append results with the commit hash to a data file on a dedicated branch, and plot it:
history:
if: github.ref == 'refs/heads/main'
steps:
- run: node bench/run.mjs pkg-head | jq -c --arg sha "$GITHUB_SHA" '. + {sha: $sha, t: now}' >> results.jsonl
- uses: benchmark-action/github-action-benchmark@v1
with:
tool: customSmallerIsBetter
output-file-path: bench-history.json
auto-push: true
gh-pages-branch: bench-data
History on shared runners is noisier than within-job comparisons, so read it for trends over weeks rather than for individual points. Size history belongs on the same chart; catching size regressions in CI covers that side.
Step 5 — decide what fails the build
Most teams find it best to report performance changes on every pull request and fail only on large, unambiguous regressions — say, more than 20% slower on a benchmark with low noise. Everything between is information for the reviewer, who knows whether a 7% slowdown in exchange for a correctness fix is acceptable. When something does fail, the comment should name the benchmark and link the base and head numbers, so the author can reproduce it locally with the same harness.
Interpreting a flagged result
A flag is a prompt to look, not a verdict. Before reverting anything, re-run the job: a regression that disappears on a second run was noise that slipped past the threshold. If it persists, reproduce it locally on a quiet machine, where noise is lower and the size of the change is clearer. Then look at what changed in the module, not just the source: a dependency bump, a profile flag, or a new feature gate can change the compiled code more than the lines in the diff suggest. Comparing the two builds with twiggy — or the hot function’s WAT — usually explains a real regression within minutes.
Keep the benchmark suite itself under review as the code evolves. A benchmark that no longer exercises a hot path in the product — because the product changed — still consumes CI time and adds noise to every comparison. Prune benchmarks that nobody would act on, and add one whenever a performance bug is fixed, so the fix is protected.
Expected output
The pull-request comment:
| benchmark | base | head | change | |
|-------------|---------:|---------:|-------:|-----------|
| parse_small | 41.2 µs | 42.5 µs | +3.1% | |
| blur_hd | 1312 µs | 1409 µs | +7.4% | |
| diff_large | 883 µs | 1084 µs | +22.8% | 🔴 slower |
| encode_png | 5120 µs | 4557 µs | −11.0% | 🟢 faster |
Over time, the comments become a record of the module’s performance history in the places where decisions were made — the pull requests themselves — which is more useful than a dashboard nobody opens.
Gotchas
- Every run flags something. Thresholds are tighter than the noise. Use per-benchmark noise in the threshold and more rounds.
- Base build is cached, head is not. Build both with the same cache state, or the comparison includes compile-time effects.
- Benchmarks take too long for every pull request. Run a fast subset on pull requests and the full suite nightly.
- Different runner CPUs between base and head. Avoided by building and running both in one job — never compare across jobs.
Performance note
With three interleaved rounds of 60 samples each, the noise floor on standard GitHub-hosted runners was 2–4% for most benchmarks and about 9% for the one that allocated heavily. That made 5% a workable default threshold and explained why the allocation-heavy benchmark needed a looser one.
Frequently Asked Questions
Should benchmarks run on dedicated hardware? If performance is central to the product, a dedicated, quiet runner reduces noise dramatically and allows tighter thresholds. For most projects, same-job comparison on shared runners is good enough.
Can I count instructions instead of timing? Instruction counts under a deterministic runtime are far less noisy and catch many regressions, though they miss cache and memory effects. Some teams track both.
Browser or Node? Node for every pull request — it is fast and stable. A browser run nightly adds confidence that the engine your users run behaves the same.
How many benchmarks should run per pull request? Few enough that the job finishes in a few minutes — a handful of representative hot paths. A suite of fifty microbenchmarks on every pull request mostly produces noise and waiting; run the full set nightly instead.
What about compile time and startup? Track them as separate benchmarks: module size, compile time and first-call time regress independently of steady-state speed.
Related
- Avoiding JIT warm-up errors in Wasm benchmarks — getting trustworthy samples.
- Caching Rust Wasm builds in GitHub Actions — keeping two builds per job affordable.
- Measuring Wasm vs JavaScript throughput — benchmarks worth tracking.
- Measuring Wasm performance with real-user monitoring — the field counterpart to CI numbers.
← Back to Wasm Performance Benchmarking