Measuring the Speed Cost of -Oz

This page answers one question: -Oz and opt-level = "z" make WebAssembly modules smaller — how much slower do they make them, for your code, and what can you do when the answer is “too much”?

Prerequisites

  • [ ] A module with a CPU-bound function you can call repeatedly: a filter, a parser, a solver.
  • [ ] A benchmark harness that warms up and reports a median, such as the one in building a reproducible Wasm benchmark harness.
  • [ ] The ability to build the same code with two optimization levels side by side.

What size optimization gives up

Optimizing for size and for speed agree on most things: removing dead code, simplifying expressions and propagating constants make code both smaller and faster. They disagree on a few transformations that trade bytes for time. Loop unrolling duplicates a loop body to reduce branch overhead and expose parallelism. Inlining copies a function body into its callers to remove call overhead and enable further optimization in context. Vectorization emits SIMD code alongside a scalar fallback. And speed-oriented code generation prefers longer instruction sequences that avoid expensive operations.

-Oz turns most of those down or off; -Os turns them down less aggressively; -O2 and -O3 turn them up. For code that spends its time in tight loops over arrays — image kernels, numeric solvers, codecs — the lost unrolling and vectorization can cost a lot. For code dominated by branching, allocation, or calls across the JavaScript boundary, the difference is often within measurement noise, and the smaller binary also compiles faster in the browser.

Transformations that differ between size and speed levels Optimizations that -O3 applies freely and -Oz restricts, the kind of code each one helps, and the typical speed effect of losing it. transformation -O3 -Oz helps code like loop unrolling aggressive mostly off tight array loops inlining generous only tiny functions small hot helpers auto-vectorization on reduced kernels with +simd128 dead code removal on on everything constant propagation on on everything

Step 1 — build both variants of the same code

Keep everything identical except the optimization level, including wasm-opt:

# speed-optimized
CARGO_PROFILE_RELEASE_OPT_LEVEL=3 cargo build --release --target wasm32-unknown-unknown
wasm-bindgen target/wasm32-unknown-unknown/release/app.wasm --out-dir build/o3 --target web
wasm-opt -O3 build/o3/app_bg.wasm -o build/o3/app_bg.wasm

# size-optimized
CARGO_PROFILE_RELEASE_OPT_LEVEL=z cargo build --release --target wasm32-unknown-unknown
wasm-bindgen target/wasm32-unknown-unknown/release/app.wasm --out-dir build/oz --target web
wasm-opt -Oz build/oz/app_bg.wasm -o build/oz/app_bg.wasm

Both should use the same LTO and codegen-unit settings; otherwise you are measuring two variables at once. Record the sizes before benchmarking — the speed result only means something next to what you saved.

Step 2 — benchmark each hot function, not the whole app

An end-to-end timing hides where the difference comes from. Benchmark each CPU-heavy export separately with realistic input sizes:

import { bench } from "./harness.js";

for (const variant of ["o3", "oz"]) {
  const m = await import(`./build/${variant}/app.js`);
  await m.default();
  const img = makeTestImage(1920, 1080);
  console.log(variant, {
    blur:  await bench(() => m.gaussian_blur(img, 3.0)),
    parse: await bench(() => m.parse_document(sampleDoc)),
    diff:  await bench(() => m.text_diff(a, b)),
  });
}

Run each benchmark in a fresh page or worker per variant, so the engine’s optimizing tier treats them identically. Report medians over many runs; a single measurement of a few milliseconds is noise.

Step 3 — read the result per function

Typical results show the pattern from the table above: numeric kernels lose more than branchy code.

Slowdown of -Oz relative to -O3, by function Median run time of three exported functions built at -Oz, as a percentage slower than the same functions at -O3. The blur kernel loses most because loop unrolling and vectorization are reduced; the parser and text diff are within a few percent. percent slower at -Oz than -O3 (lower is better) gaussian_blur (tight loops) 31 % parse_document (branchy) 4 % text_diff (allocations) 2 % Sizes for the same module were 412 KB at -O3 and 298 KB at -Oz before compression — a 28% saving.

That combination — one function far slower, the rest barely affected — is the common case, and it argues against choosing one level for the whole module.

Step 4 — mix levels by crate

Cargo lets each crate in the dependency graph have its own optimization level. Put hot numeric code in its own crate and optimize it for speed, and everything else for size:

# Cargo.toml (workspace or top-level crate)
[profile.release]
opt-level = "z"
lto = "fat"
codegen-units = 1

[profile.release.package.image-kernels]
opt-level = 3

With fat LTO, the per-package level still applies to that crate’s functions during the whole-program optimization. For a single hot function inside a size-optimized crate, #[inline] on its small helpers recovers much of the inlining -Oz would skip, without changing the rest of the crate.

A final wasm-opt -Oz pass then runs over everything and can undo some of the per-crate decisions, since it does not know which functions you wanted fast. Use wasm-opt -Os for the final pass in mixed builds; it removes nearly as many bytes and is less inclined to shrink hot loops.

Step 5 — decide with both numbers on the page

The right choice depends on how the module is used. A module that loads on every page view and runs a few milliseconds of logic should be as small as possible: the download and compile dominate. A module that loads once and then runs a heavy kernel thousands of times should be fast. Most real modules are a mix, which is why the per-crate split usually wins.

Picking optimization levels from how the module is used A decision tree. Modules dominated by load time should be size-optimized; modules dominated by repeated heavy computation should be speed-optimized; mixed modules should optimize hot crates for speed and the rest for size. What dominates the module's cost to the user? download + compile opt-level z, wasm-opt -Oz smallest bytes, fastest start repeated heavy kernels opt-level 3, wasm-opt -O3 unrolling and SIMD kept both z overall, 3 for hot crates final pass -Os

Why the answer differs between projects

It would be convenient to quote a single number — “-Oz costs fifteen percent” — but the honest answer depends almost entirely on the shape of the hot code, and it is worth understanding why before trusting anyone’s benchmark, including the one above.

Loop-heavy numeric code is where size optimization costs the most. A blur kernel, a matrix multiply or an audio filter spends nearly all its time in a few small loops over arrays. Speed optimization unrolls those loops, keeps more values in locals, and — when SIMD is enabled — processes several elements per instruction. Size optimization does less of all three, so each iteration does less useful work per instruction executed. The result is the thirty percent seen for the blur.

Branch-heavy code behaves differently. A parser or a state machine spends its time deciding what to do next, and the cost is dominated by branches, memory loads and calls that neither optimization level can remove. Unrolling does not help code whose loop body is a large switch, and inlining large functions mostly makes them larger. Here the levels converge, and the smaller binary may even be faster because more of it fits in the CPU’s instruction cache.

Code dominated by allocation or by calls across the JavaScript boundary converges too, because the expensive part is outside the code being optimized: in the allocator, the glue, or the browser. That is why measuring each hot function separately, as in step 2, matters more than any general rule — and why the per-crate split is so often the right answer.

Expected output

A summary table from the harness for the mixed build:

variant    size     brotli   blur(ms)  parse(ms)  diff(ms)
o3         412 KB   141 KB   18.4      6.10       2.33
oz         298 KB   104 KB   24.1      6.34       2.38
mixed      317 KB   110 KB   18.9      6.31       2.37

The mixed build keeps nearly all of the size saving and nearly all of the blur speed.

Gotchas

  • Measuring the first call. The engine’s baseline tier runs first; a few calls in, optimized code takes over. Warm up before timing, or the comparison measures compile tiers, not your code.
  • Different wasm-opt levels on the two builds. wasm-opt -O3 on one and -Oz on the other is fine if intended; running it on only one of them makes the comparison meaningless.
  • SIMD only on one side. If +simd128 is enabled for one build and not the other, the vectorization difference swamps the optimization level. Keep target features identical.
  • Per-package levels ignored. The package name in [profile.release.package.NAME] must match the crate name exactly, with hyphens as written in its Cargo.toml.

Performance note

Smaller modules also compile faster. On a mid-range phone, the -Oz build above compiled in 190 ms against 260 ms for -O3. For a page that runs the blur once, the size-optimized build was faster end to end despite the slower kernel; for an editor that applies filters interactively, the speed-optimized kernel won within the first ten operations.

Frequently Asked Questions

Is -Os a reasonable compromise? Often. -Os keeps more inlining than -Oz and usually lands within a few percent of -O3 speed with most of the size saving. Measure all three on your hot functions.

Does Emscripten behave the same way? Yes — -Oz, -Os, -O2 and -O3 have the same trade-offs, and per-file optimization levels give the same mixed-build option.

Can the engine’s optimizer make up the difference? Partly. The browser’s optimizing tier inlines and optimizes again, but it does not unroll or vectorize the way LLVM does, and it works under a compile-time budget. Code that LLVM did not optimize for speed stays slower.

What about -O4 in wasm-opt? It runs more passes, including flattening, and occasionally finds extra speed for compute-heavy modules at the cost of much longer optimization time. Worth trying for a kernel crate, rarely for a whole application.

← Back to Wasm Optimization Flags & Size Reduction