Tracking Benchmark Results in CI

This page answers one task: catch performance regressions in a WebAssembly module automatically, the way tests catch correctness regressions — by running benchmarks in CI, comparing with a baseline, and flagging slowdowns that are real.

Prerequisites

  • [ ] A benchmark harness that prints machine-readable results, such as the one from building a reproducible Wasm benchmark harness.
  • [ ] A CI system where you can store artifacts or push to a branch — GitHub Actions in the examples.
  • [ ] Realistic expectations about CI machines, which are noisy.

The problem with timing on shared runners

Benchmarks in CI face a measurement problem that tests do not. A test either passes or fails; a benchmark produces a number, and on a shared CI runner that number varies from run to run by 5–15% for reasons unrelated to your code: other workloads on the same host, different CPU models in the runner pool, frequency scaling, and cache state. A naive setup that fails the build when a benchmark is 5% slower than last time fails constantly, and teams learn to ignore it within a week.

Three things make CI benchmarking useful despite the noise. Compare against a baseline measured in the same job, on the same machine, rather than against a stored number from another day. Use enough samples and a statistic that resists outliers. And set thresholds based on the measured noise of each benchmark, flagging only changes larger than that. With those in place, CI reliably catches the regressions that matter — a 20% slowdown from an accidentally disabled optimization flag — while ignoring the noise.

Benchmarking a pull request against its base in one job The job checks out and builds both the base commit and the pull request head on the same runner, runs the benchmarks for both in interleaved rounds, compares medians against each benchmark's noise threshold, and posts a summary to the pull request. build base + head same runner, same flags interleaved rounds A B A B … compare medians per-benchmark threshold PR comment table + flags

Step 1 — make the harness emit JSON

The harness should print one record per benchmark with the statistics you will compare:

// bench/run.mjs — prints JSON lines
import { loadModule, bench } from "./harness.mjs";

const m = await loadModule(process.argv[2]);          // path to a built pkg/
for (const [name, fn] of Object.entries(m.benchmarks())) {
  const r = await bench(fn, { warmupMs: 500, samples: 60 });
  console.log(JSON.stringify({ name, median_ns: r.median * 1e6, p90_ns: r.p90 * 1e6, mad_ns: r.mad * 1e6 }));
}

MAD — the median absolute deviation — is a robust measure of spread, and it is what the threshold will use. Run in Node or a headless browser; Node is faster to start and good enough for tracking relative changes, while a browser run adds coverage of the engine your users have.

Step 2 — build base and head in the same job

# .github/workflows/bench.yml
on: pull_request
jobs:
  bench:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with: { fetch-depth: 0 }
      - uses: Swatinem/rust-cache@v2
      - name: build base and head
        run: |
          git worktree add ../base ${{ github.event.pull_request.base.sha }}
          (cd ../base && wasm-pack build --release --target nodejs --out-dir ../pkg-base)
          wasm-pack build --release --target nodejs --out-dir pkg-head
      - name: run interleaved
        run: |
          for round in 1 2 3; do
            node bench/run.mjs pkg-base >> base.jsonl
            node bench/run.mjs pkg-head >> head.jsonl
          done
      - run: node bench/compare.mjs base.jsonl head.jsonl > summary.md
      - uses: marocchino/sticky-pull-request-comment@v2
        with: { path: summary.md }

Building both in the same job is the key decision. Whatever the runner’s speed today, it affects both builds equally, and the comparison measures the change rather than the machine. Interleaving rounds spreads time-dependent noise evenly between them.

Step 3 — compare with a noise-aware threshold

For each benchmark, combine the rounds, take the median of medians, and compare. Flag a change only when it exceeds both a minimum percentage and a multiple of the observed spread:

// bench/compare.mjs
const load = (f) => readFileSync(f, "utf8").trim().split("\n").map(JSON.parse);
const group = (rows) => rows.reduce((m, r) => ((m[r.name] ??= []).push(r), m), {});
const med = (xs) => xs.sort((a, b) => a - b)[xs.length >> 1];

const base = group(load(process.argv[2])), head = group(load(process.argv[3]));
console.log("| benchmark | base | head | change | |\n|---|---:|---:|---:|---|");
for (const name of Object.keys(base)) {
  const b = med(base[name].map((r) => r.median_ns)), h = med(head[name].map((r) => r.median_ns));
  const noise = 3 * med(base[name].map((r) => r.mad_ns)) / b;   // relative
  const change = (h - b) / b;
  const flag = Math.abs(change) < Math.max(0.05, noise) ? "" : change > 0 ? "🔴 slower" : "🟢 faster";
  console.log(`| ${name} | ${(b / 1e3).toFixed(1)} µs | ${(h / 1e3).toFixed(1)} µs | ${(change * 100).toFixed(1)}% | ${flag} |`);
}

The threshold max(5%, 3 × relative MAD) means a noisy benchmark needs a larger change to be flagged than a stable one, which is exactly what keeps false alarms down.

How the threshold treats different changes Example benchmark results showing the change between base and head, the benchmark's own noise level, and whether the comparison flags it. Only changes larger than both five percent and three times the noise are reported. benchmark change noise (3×MAD) result parse_small +3.1% ±2.0% not flagged (< 5%) blur_hd +7.4% ±9.5% not flagged (noisy) diff_large +22.8% ±3.2% flagged slower encode_png −11.0% ±2.4% flagged faster

Step 4 — keep a history on the main branch

Pull-request comparisons catch sudden regressions. Slow drift — one percent per change over a quarter — needs history. On each push to the main branch, append results with the commit hash to a data file on a dedicated branch, and plot it:

  history:
    if: github.ref == 'refs/heads/main'
    steps:
      - run: node bench/run.mjs pkg-head | jq -c --arg sha "$GITHUB_SHA" '. + {sha: $sha, t: now}' >> results.jsonl
      - uses: benchmark-action/github-action-benchmark@v1
        with:
          tool: customSmallerIsBetter
          output-file-path: bench-history.json
          auto-push: true
          gh-pages-branch: bench-data

History on shared runners is noisier than within-job comparisons, so read it for trends over weeks rather than for individual points. Size history belongs on the same chart; catching size regressions in CI covers that side.

Step 5 — decide what fails the build

Most teams find it best to report performance changes on every pull request and fail only on large, unambiguous regressions — say, more than 20% slower on a benchmark with low noise. Everything between is information for the reviewer, who knows whether a 7% slowdown in exchange for a correctness fix is acceptable. When something does fail, the comment should name the benchmark and link the base and head numbers, so the author can reproduce it locally with the same harness.

Interpreting a flagged result

A flag is a prompt to look, not a verdict. Before reverting anything, re-run the job: a regression that disappears on a second run was noise that slipped past the threshold. If it persists, reproduce it locally on a quiet machine, where noise is lower and the size of the change is clearer. Then look at what changed in the module, not just the source: a dependency bump, a profile flag, or a new feature gate can change the compiled code more than the lines in the diff suggest. Comparing the two builds with twiggy — or the hot function’s WAT — usually explains a real regression within minutes.

Keep the benchmark suite itself under review as the code evolves. A benchmark that no longer exercises a hot path in the product — because the product changed — still consumes CI time and adds noise to every comparison. Prune benchmarks that nobody would act on, and add one whenever a performance bug is fixed, so the fix is protected.

Expected output

The pull-request comment:

| benchmark   | base     | head     | change |           |
|-------------|---------:|---------:|-------:|-----------|
| parse_small |  41.2 µs |  42.5 µs |  +3.1% |           |
| blur_hd     | 1312 µs  | 1409 µs  |  +7.4% |           |
| diff_large  |  883 µs  | 1084 µs  | +22.8% | 🔴 slower |
| encode_png  | 5120 µs  | 4557 µs  | −11.0% | 🟢 faster |

Over time, the comments become a record of the module’s performance history in the places where decisions were made — the pull requests themselves — which is more useful than a dashboard nobody opens.

Gotchas

  • Every run flags something. Thresholds are tighter than the noise. Use per-benchmark noise in the threshold and more rounds.
  • Base build is cached, head is not. Build both with the same cache state, or the comparison includes compile-time effects.
  • Benchmarks take too long for every pull request. Run a fast subset on pull requests and the full suite nightly.
  • Different runner CPUs between base and head. Avoided by building and running both in one job — never compare across jobs.

Performance note

With three interleaved rounds of 60 samples each, the noise floor on standard GitHub-hosted runners was 2–4% for most benchmarks and about 9% for the one that allocated heavily. That made 5% a workable default threshold and explained why the allocation-heavy benchmark needed a looser one.

Measured noise per benchmark on shared CI runners Three times the relative median absolute deviation for each benchmark, across many identical runs on standard GitHub-hosted runners. Allocation-heavy benchmarks are noisier and need looser thresholds. noise as % of median (3 × MAD) parse_small 2 % diff_large 3.2 % encode_png 2.4 % blur_hd (allocates per call) 9.5 %

Frequently Asked Questions

Should benchmarks run on dedicated hardware? If performance is central to the product, a dedicated, quiet runner reduces noise dramatically and allows tighter thresholds. For most projects, same-job comparison on shared runners is good enough.

Can I count instructions instead of timing? Instruction counts under a deterministic runtime are far less noisy and catch many regressions, though they miss cache and memory effects. Some teams track both.

Browser or Node? Node for every pull request — it is fast and stable. A browser run nightly adds confidence that the engine your users run behaves the same.

How many benchmarks should run per pull request? Few enough that the job finishes in a few minutes — a handful of representative hot paths. A suite of fifty microbenchmarks on every pull request mostly produces noise and waiting; run the full set nightly instead.

What about compile time and startup? Track them as separate benchmarks: module size, compile time and first-call time regress independently of steady-state speed.

← Back to Wasm Performance Benchmarking