Comparing wasm-opt Optimization Levels

This page answers one task: you run Binaryen’s wasm-opt on your module, but you picked the level by habit — and you want to measure what each level does to size and speed on your code, so you can choose deliberately.

Prerequisites

  • [ ] A release build of the module, compiled by your toolchain with its normal optimisation.
  • [ ] Binaryen’s wasm-opt (a recent version; levels improve between releases).
  • [ ] A benchmark that exercises the module’s hot paths, and brotli for compressed sizes.

What the levels mean

wasm-opt runs a pipeline of passes over the module. The level selects which passes run and how aggressively: -O1 runs quick cleanups, -O2 adds most standard optimisations, -O3 adds more expensive ones, and -O4 adds flattening-based passes that are slow and occasionally find extra speed. -Os and -Oz run the -O2-style pipeline with shrink levels 1 and 2, which make inlining and other size-increasing transformations conservative. In Binaryen terms, the levels set optimizeLevel (0–4) and shrinkLevel (0–2); you can set them separately with --optimize-level and --shrink-level and add passes by name.

Because the input is already optimised by LLVM or another compiler, wasm-opt does not start from scratch. Its gains come from Wasm-specific work the compiler does not do — local and global simplification, removing redundant local.get/local.set pairs, merging blocks, function-level dead-code elimination, and inlining decisions made with the whole module visible. Typical results are 5–20% smaller and a few percent faster than the compiler’s output, but the spread between levels is often smaller than people expect, which is why measuring matters.

wasm-opt levels at a glance O1 runs quick cleanups. O2 runs the standard pipeline. O3 adds expensive passes. O4 adds flattening passes that are slow to run. Os and Oz use the standard pipeline with shrink levels 1 and 2, limiting inlining for size. flag optimize level shrink level intent -O1 1 0 fast cleanups -O2 2 0 standard pipeline -O3 3 0 more expensive passes -O4 4 0 + flattening, slow -Os 2 1 smaller code -Oz 2 2 smallest code

Step 1 — produce one build per level

Start from the same input every time — the toolchain’s output before any wasm-opt run. wasm-pack and Emscripten run wasm-opt themselves, so disable that (wasm-opt = false in [package.metadata.wasm-pack.profile.release], or take the pre-optimisation file) to get a clean baseline:

IN=target/wasm32-unknown-unknown/release/app.wasm
for lvl in O1 O2 O3 O4 Os Oz; do
  /usr/bin/time -f "$lvl %es" wasm-opt -$lvl "$IN" -o "out/app.$lvl.wasm"
done

Record how long each run took — -O4 can take minutes on large modules, which matters for CI and incremental builds.

Step 2 — measure raw and compressed size

for f in out/app.*.wasm; do
  printf "%-16s %8d %8d\n" "$(basename $f)" "$(stat -c%s $f)" "$(brotli -c $f | wc -c)"
done

Compressed size is what users download. Differences between levels shrink after compression, because the extra code from inlining is often repetitive; a 12% raw difference between -O3 and -Oz might be 7% compressed.

Step 3 — measure run time on real workloads

Load each variant in the same benchmark harness, warm up, and run the hot paths:

for (const lvl of ["O1", "O2", "O3", "O4", "Os", "Oz"]) {
  const mod = await load(`out/app.${lvl}.wasm`);
  warmUp(mod);
  console.log(lvl, median(runs(() => mod.process(input), 50)).toFixed(2), "ms");
}

Measure in the engines your users run; inlining decisions that help in one engine’s tier-up may matter less in another. See building a reproducible Wasm benchmark harness for a harness that controls warm-up and variance.

Compressed size by wasm-opt level for an image-processing module Brotli-compressed kilobytes of the same module after the compiler alone and after wasm-opt at each level. Oz is smallest, O3 and O4 are largest, and the spread across levels is about 15 percent. KB compressed compiler output 238 KB -O1 226 KB -O2 219 KB -O3 224 KB -O4 225 KB -Os 205 KB -Oz 197 KB

Step 4 — try converge and extra passes

--converge repeats the pipeline until the module stops shrinking, typically saving another 1–3% at the cost of running the pipeline two to four times. Some passes are worth adding by name when a module has a specific shape: --gufa (global type-flow analysis) for modules with many indirect calls or GC types, --inline-functions-with-loops or --flexible-inline-max-function-size to tune inlining, and --strip-debug --strip-producers to remove sections not needed in production. Measure each addition separately; most give nothing on most modules.

wasm-opt -Oz --converge --strip-debug --strip-producers "$IN" -o out/app.Oz-converge.wasm

Step 5 — choose per module and record the choice

With the measurements in a table, the decision is usually clear. If -Oz is within noise of -O3 on the benchmarks, ship -Oz. If hot paths are measurably faster with -O3, and the size cost is acceptable, ship -O3. For applications with several modules, choose per module: a startup-critical UI module can be -Oz while a compute kernel loaded later is -O3. Write the chosen level and the measurement date into the build configuration with a comment, and re-measure when upgrading Binaryen or the compiler.

Interaction with compiler optimisation levels

wasm-opt levels and compiler levels compound. A Rust build with opt-level = "z" followed by wasm-opt -O3 can re-inline code the compiler deliberately kept out of line; a build with opt-level = 3 followed by wasm-opt -Oz keeps much of the compiler’s inlining. Test the combinations that make sense — usually opt-level = "s" or "z" with -Oz for size, and opt-level = 3 with -O3 for speed — rather than mixing intents. The speed cost of size optimisation on the compiler side is examined in measuring the speed cost of -Oz.

Optimisation time in development

wasm-opt at high levels slows builds, and in the development loop size does not matter. Skip wasm-opt for development builds, or use -O1 for a quick cleanup that keeps debugging usable. Running it only in release builds and in CI keeps iteration fast. Caching wasm-opt output keyed by the input file’s hash also helps CI when the Wasm input has not changed.

Startup cost as a third dimension

Size and steady-state speed are the usual axes, but levels also affect how fast a module starts. Engines compile every function before or during first use; a module with more code — from aggressive inlining at -O3 — takes longer to compile in the baseline tier, and functions that inlining made very large can take disproportionately long in the optimising tier. For modules on the critical path of page load, measure time from fetch to first useful call for each level, not only download size and benchmark time. In practice -Os and -Oz often win startup twice over: fewer bytes to download and fewer bytes to compile. The difference is most visible on mid-range phones, where baseline compilation of a few megabytes takes hundreds of milliseconds; see measuring Wasm startup time end to end.

Automating the comparison

Running the comparison by hand once is useful; running it automatically whenever Binaryen, the compiler or major dependencies change is better. A small script that builds every level, records the four numbers — optimisation time, compressed size, startup time and hot-path time — and writes a Markdown table into the pull request turns the choice into a reviewed decision with data behind it. Keep the script in the repository next to the build configuration, so whoever upgrades the toolchain next can rerun it without reconstructing the method.

Expected output

A table of six levels with optimisation time, raw size, compressed size and median run time for three hot paths; a decision to ship the UI module at -Oz --converge (197 KB compressed, 2% slower than -O3 on the slowest path) and the filter module at -O3; and the chosen flags recorded in the build scripts with the Binaryen version used for measurement.

Gotchas

  • Optimising already-optimised output. wasm-pack or Emscripten may run wasm-opt already. Measure from the unoptimised input.
  • Comparing raw sizes only. Compressed differences are smaller. Measure what users download.
  • Assuming -O4 is fastest. It often is not, and it is slow to run.
  • Mixed intents. Size-optimised compiler output with speed-optimised wasm-opt undoes choices. Pair them deliberately.
  • Stale measurements. Binaryen improves; re-measure on upgrades.
  • Ignoring startup. Larger, more inlined code compiles more slowly in the engine. Measure time to first call too.

Performance note

For the image-processing module, -O3 ran the hottest filter 6% faster than -Oz and 2% faster than -O2; the other two benchmarks were within 1% across all levels. -O4 took 74 s to run against 9 s for -O3, with no speed benefit.

Hottest filter run time by wasm-opt level Milliseconds for the slowest benchmark of the image-processing module after wasm-opt at O2, O3, O4, Os and Oz, measured as medians after warm-up. ms per run (median) -O2 41.2 ms -O3 40.4 ms -O4 40.5 ms -Os 42.1 ms -Oz 42.9 ms

Frequently Asked Questions

Should I always run wasm-opt? For release builds, yes — it almost always saves size and rarely costs speed.

Is -O3 safe? Yes. All levels preserve semantics; the difference is effort and trade-offs.

Why is -O3 larger than -O2? More inlining and loop transformations add code to save time.

Does wasm-opt help Emscripten builds? Emscripten already runs it at the matching level; extra runs rarely help.

Does the level affect startup time? Yes — more inlined code takes longer to compile in the engine, so size-oriented levels often start faster as well as downloading faster.

Can different functions in one module use different levels? Not with a single wasm-opt run; split hot code into a separate module or tune inlining with pass options instead.

← Back to Wasm Optimization Flags & Size Reduction