Catching Size Regressions in CI

This guide answers one task: make a pipeline fail when a WebAssembly module grows more than it should, so payload increases are a decision someone made rather than something discovered six months later.

Prerequisites

  • [ ] A reproducible release build that produces the same bytes from the same commit.
  • [ ] brotli, because compressed size is what users download.
  • [ ] A place to store a baseline — a committed file is the simplest.
  • [ ] A tolerance you have agreed on, rather than one invented during an incident.

Measure what the user downloads

Three numbers exist and only one of them matters for a gate. The raw .wasm size is what the build produces. The compressed size is what crosses the network. The in-memory size is neither, and is not what this check is about.

Gate on the compressed size, computed the same way every time, and record the raw size alongside it for diagnosis — a change that leaves compressed size flat while raw size grows usually means more of something highly compressible, such as repeated error strings.

RAW=$(stat -c%s dist/engine.wasm)
GZ=$(brotli -q 11 -c dist/engine.wasm | wc -c)
echo "raw=$RAW compressed=$GZ"

Include the glue JavaScript if your users download it. A module that shrinks while its generated bindings grow has not improved anything, and a gate that measures only the binary will report success.

Gate on the number users pay The raw module size is a build artifact, the compressed size is the transfer cost, and the generated glue adds to it. A gate that ignores compression or the glue measures something other than what the user experiences. raw .wasm 148 kB — a build number compressed 52 kB — gate on this + glue 64 kB total transfer Reporting all three makes a regression diagnosable: which one grew tells you whether it was code, data or bindings. Gating on all three is overkill; gate on the total transfer and record the others.

A baseline and a tolerance

The baseline is the size at the last accepted state; the tolerance is how much growth is allowed without a conversation. A committed file is the simplest store and has the useful property that updating it appears in the diff, so growth is reviewed rather than absorbed.

# .size-baseline
wasm_compressed=53248
glue_compressed=11776
#!/usr/bin/env bash
# scripts/check-size.sh
set -euo pipefail
source .size-baseline

TOL=${SIZE_TOLERANCE:-0.03}          # 3% by default
wasm_now=$(brotli -q 11 -c dist/engine_bg.wasm | wc -c)
glue_now=$(brotli -q 11 -c dist/engine.js | wc -c)

check() {
  local name=$1 now=$2 base=$3
  awk -v n="$now" -v b="$base" -v t="$TOL" -v name="$name" 'BEGIN {
    delta = (n - b) / b
    printf "%-16s %8d → %8d  (%+.2f%%)\n", name, b, n, delta * 100
    if (delta > t) { printf "  regression: %s exceeds %.0f%% tolerance\n", name, t * 100; exit 1 }
  }'
}

check wasm "$wasm_now" "$wasm_compressed"
check glue "$glue_now" "$glue_compressed"

A three percent tolerance is a reasonable default: large enough that a compiler update does not fail the build, small enough that a real addition is visible. Make it configurable so a deliberate large change can pass with an explicit override and a note in the commit message.

Reporting rather than only failing

A gate that fails is useful; a gate that reports on every change is better, because it makes size a thing people see rather than a thing they trip over.

# emit a markdown table for a pull request comment
{
  echo "| artifact | baseline | this build | change |"
  echo "| --- | ---: | ---: | ---: |"
  printf "| engine.wasm | %d | %d | %+.2f%% |\n" \
    "$wasm_compressed" "$wasm_now" \
    "$(awk -v n=$wasm_now -v b=$wasm_compressed 'BEGIN{print (n-b)/b*100}')"
} > size-report.md
| artifact    | baseline | this build | change  |
| ----------- | -------: | ---------: | ------: |
| engine.wasm |    53248 |      54912 |  +3.13% |
| engine.js   |    11776 |      11776 |  +0.00% |

Posting that on every pull request changes behaviour. A developer who sees “+3.13%” next to their change asks whether the dependency they added was necessary, which is a conversation that does not happen when the number is invisible.

Finding what grew

When the gate fires, the question is which change caused it, and two tools answer it quickly.

twiggy attributes bytes to functions and data, and its diff mode compares two builds directly:

twiggy top -n 20 dist/engine.wasm
twiggy diff baseline.wasm dist/engine.wasm -n 20
 Delta Bytes │ Item
─────────────┼───────────────────────────────────────────
       +8912 │ core::fmt::Formatter::pad
       +3104 │ ::fmt
       +1288 │ data[3]
        -412 │ engine::parse_records

That output is typical and immediately diagnostic: a formatting or error-display path was pulled in, which usually means a format! or a Display implementation reached code that had previously avoided it.

Bisecting works when the attribution is unclear:

git bisect start HEAD HEAD~30
git bisect run sh -c 'cargo build --release --target wasm32-unknown-unknown &&
  test $(brotli -q 11 -c target/wasm32-unknown-unknown/release/engine.wasm | wc -c) -lt 54000'
What the gate is protecting against Small increases individually pass a tolerance check but accumulate into a large regression over time. Updating the baseline on every accepted change keeps the comparison honest against the last agreed state. commits → unchecked: +62% over a quarter baseline, updated deliberately Every step on the upper line passed a per-change tolerance. The gate catches them only if the baseline moves when a change is accepted, not automatically.

The changes that grow a module

Knowing the usual culprits makes a regression report much faster to act on, because most growth comes from a small set of causes.

Formatting and error display are the most common in Rust. A single format! in a code path that previously had none can pull in tens of kilobytes of core::fmt machinery, and a Display implementation on an error type does the same. Returning error codes rather than formatted messages from the module, and formatting on the JavaScript side, keeps that machinery out entirely.

A new dependency is the second. Crates vary enormously in how much they bring: a small numeric crate may add a kilobyte, while one with a builder API and rich errors can add fifty. Check the delta before and after adding one rather than after a month of accumulated changes.

Panic unwinding is the third. Without panic = "abort" the binary carries landing pads and unwinding tables throughout, which for a mid-sized module is commonly 15–25% of it.

Debug information is the fourth and the easiest to fix: a name or .debug_* custom section left in a release build inflates it substantially and does nothing for users. wasm-opt --strip-debug or the toolchain’s own strip flag removes it.

# see what a candidate dependency costs before committing to it
brotli -q 11 -c dist/engine.wasm | wc -c        # before
cargo add some-crate && cargo build --release --target wasm32-unknown-unknown
brotli -q 11 -c dist/engine.wasm | wc -c        # after

Running that two-line comparison before adding a dependency takes a minute and has, on more than one project, been the reason a convenient crate was replaced with thirty lines of hand-written code.

Updating the baseline honestly

The baseline should move when growth is accepted, and moving it should be visible.

Update it in the same commit as the change that grew the module, with the reason in the message. Do not regenerate it automatically on the main branch — an auto-updating baseline ratchets upward one accepted change at a time and reports no regression ever, which is the failure mode the dashed line in the figure above describes.

# accepted a deliberate increase
./scripts/update-baseline.sh
git add .size-baseline
git commit -m "Add AVIF decode path (+18 kB compressed, accepted)"

A baseline with a history is also a useful artifact in itself: git log -p .size-baseline reads as a record of every deliberate payload decision the project has made.

Expected output

A passing run is quiet and a failing one is specific:

./scripts/check-size.sh
wasm                53248 →    53696  (+0.84%)
glue                11776 →    11776  (+0.00%)
artifact ok
./scripts/check-size.sh
wasm                53248 →    61440  (+15.38%)
  regression: wasm exceeds 3% tolerance
Where the budget bites The measured binary sits inside a budget with two thresholds: a warning band that comments on the pull request and a hard ceiling that fails the job. current build 181 KB compressed headroom warning at +5% 190 KB — comment on the pull request fail at +15% 208 KB — the job fails and the merge is blocked Budget the compressed size, because that is what users wait for and what a compression-friendly change moves. Store the baseline on the default branch, not in the repository, or every merge rewrites the number it checks.

Gotchas

  • Gating on raw size. Compression ratios vary; the user pays for compressed bytes.
  • Ignoring the glue. A module that shrinks while its bindings grow has not improved.
  • Auto-updating the baseline. Guarantees the gate never fires.
  • A tolerance set during an incident. Ends up at 25% and stops meaning anything.
  • Different brotli versions between local and CI. Produces different numbers; pin it.
  • Non-reproducible builds. Makes every comparison noisy and the gate untrustworthy.

Performance note

The whole check — build, compress, compare — adds about 3 s to a pipeline for a 150 kB module, most of it the maximum-quality compression. Running brotli -q 5 instead is ten times faster and gives a number within a few percent, which is fine for a relative comparison provided the baseline was produced the same way. For the number you quote publicly, use quality 11.

Frequently Asked Questions

What tolerance should I use? Three percent for an established module, wider while a project is young and changing shape. The important part is that it is agreed in advance rather than adjusted to make a build pass.

Should the gate block a merge? Yes, with an easy override. A warning that does not block gets ignored within a month; a block with a documented way to accept the growth produces exactly the conversation the gate exists to start.

Where should the report be published? Wherever the team already looks — a pull request comment is usually enough, and a long-run chart is a bonus rather than a requirement.

Does this work for a multi-artifact build? Yes — extend the baseline file with one entry per artifact and loop. A build producing baseline and SIMD variants should track both, since they grow for different reasons.

The gate is five minutes of setup and it is the difference between knowing your payload and guessing it.

← Back to Testing & Verifying Wasm Builds