Measuring the Speed Cost of -Oz
This page answers one question: -Oz and opt-level = "z" make WebAssembly modules smaller — how much slower do they make
them, for your code, and what can you do when the answer is “too much”?
Prerequisites
- [ ] A module with a CPU-bound function you can call repeatedly: a filter, a parser, a solver.
- [ ] A benchmark harness that warms up and reports a median, such as the one in building a reproducible Wasm benchmark harness.
- [ ] The ability to build the same code with two optimization levels side by side.
What size optimization gives up
Optimizing for size and for speed agree on most things: removing dead code, simplifying expressions and propagating constants make code both smaller and faster. They disagree on a few transformations that trade bytes for time. Loop unrolling duplicates a loop body to reduce branch overhead and expose parallelism. Inlining copies a function body into its callers to remove call overhead and enable further optimization in context. Vectorization emits SIMD code alongside a scalar fallback. And speed-oriented code generation prefers longer instruction sequences that avoid expensive operations.
-Oz turns most of those down or off; -Os turns them down less aggressively; -O2 and -O3 turn them up. For code that
spends its time in tight loops over arrays — image kernels, numeric solvers, codecs — the lost unrolling and vectorization
can cost a lot. For code dominated by branching, allocation, or calls across the JavaScript boundary, the difference is
often within measurement noise, and the smaller binary also compiles faster in the browser.
Step 1 — build both variants of the same code
Keep everything identical except the optimization level, including wasm-opt:
# speed-optimized
CARGO_PROFILE_RELEASE_OPT_LEVEL=3 cargo build --release --target wasm32-unknown-unknown
wasm-bindgen target/wasm32-unknown-unknown/release/app.wasm --out-dir build/o3 --target web
wasm-opt -O3 build/o3/app_bg.wasm -o build/o3/app_bg.wasm
# size-optimized
CARGO_PROFILE_RELEASE_OPT_LEVEL=z cargo build --release --target wasm32-unknown-unknown
wasm-bindgen target/wasm32-unknown-unknown/release/app.wasm --out-dir build/oz --target web
wasm-opt -Oz build/oz/app_bg.wasm -o build/oz/app_bg.wasm
Both should use the same LTO and codegen-unit settings; otherwise you are measuring two variables at once. Record the sizes before benchmarking — the speed result only means something next to what you saved.
Step 2 — benchmark each hot function, not the whole app
An end-to-end timing hides where the difference comes from. Benchmark each CPU-heavy export separately with realistic input sizes:
import { bench } from "./harness.js";
for (const variant of ["o3", "oz"]) {
const m = await import(`./build/${variant}/app.js`);
await m.default();
const img = makeTestImage(1920, 1080);
console.log(variant, {
blur: await bench(() => m.gaussian_blur(img, 3.0)),
parse: await bench(() => m.parse_document(sampleDoc)),
diff: await bench(() => m.text_diff(a, b)),
});
}
Run each benchmark in a fresh page or worker per variant, so the engine’s optimizing tier treats them identically. Report medians over many runs; a single measurement of a few milliseconds is noise.
Step 3 — read the result per function
Typical results show the pattern from the table above: numeric kernels lose more than branchy code.
That combination — one function far slower, the rest barely affected — is the common case, and it argues against choosing one level for the whole module.
Step 4 — mix levels by crate
Cargo lets each crate in the dependency graph have its own optimization level. Put hot numeric code in its own crate and optimize it for speed, and everything else for size:
# Cargo.toml (workspace or top-level crate)
[profile.release]
opt-level = "z"
lto = "fat"
codegen-units = 1
[profile.release.package.image-kernels]
opt-level = 3
With fat LTO, the per-package level still applies to that crate’s functions during the whole-program optimization. For a
single hot function inside a size-optimized crate, #[inline] on its small helpers recovers much of the inlining -Oz would
skip, without changing the rest of the crate.
A final wasm-opt -Oz pass then runs over everything and can undo some of the per-crate decisions, since it does not know
which functions you wanted fast. Use wasm-opt -Os for the final pass in mixed builds; it removes nearly as many bytes and
is less inclined to shrink hot loops.
Step 5 — decide with both numbers on the page
The right choice depends on how the module is used. A module that loads on every page view and runs a few milliseconds of logic should be as small as possible: the download and compile dominate. A module that loads once and then runs a heavy kernel thousands of times should be fast. Most real modules are a mix, which is why the per-crate split usually wins.
Why the answer differs between projects
It would be convenient to quote a single number — “-Oz costs fifteen percent” — but the honest answer depends almost
entirely on the shape of the hot code, and it is worth understanding why before trusting anyone’s benchmark, including the one
above.
Loop-heavy numeric code is where size optimization costs the most. A blur kernel, a matrix multiply or an audio filter spends nearly all its time in a few small loops over arrays. Speed optimization unrolls those loops, keeps more values in locals, and — when SIMD is enabled — processes several elements per instruction. Size optimization does less of all three, so each iteration does less useful work per instruction executed. The result is the thirty percent seen for the blur.
Branch-heavy code behaves differently. A parser or a state machine spends its time deciding what to do next, and the cost is dominated by branches, memory loads and calls that neither optimization level can remove. Unrolling does not help code whose loop body is a large switch, and inlining large functions mostly makes them larger. Here the levels converge, and the smaller binary may even be faster because more of it fits in the CPU’s instruction cache.
Code dominated by allocation or by calls across the JavaScript boundary converges too, because the expensive part is outside the code being optimized: in the allocator, the glue, or the browser. That is why measuring each hot function separately, as in step 2, matters more than any general rule — and why the per-crate split is so often the right answer.
Expected output
A summary table from the harness for the mixed build:
variant size brotli blur(ms) parse(ms) diff(ms)
o3 412 KB 141 KB 18.4 6.10 2.33
oz 298 KB 104 KB 24.1 6.34 2.38
mixed 317 KB 110 KB 18.9 6.31 2.37
The mixed build keeps nearly all of the size saving and nearly all of the blur speed.
Gotchas
- Measuring the first call. The engine’s baseline tier runs first; a few calls in, optimized code takes over. Warm up before timing, or the comparison measures compile tiers, not your code.
- Different
wasm-optlevels on the two builds.wasm-opt -O3on one and-Ozon the other is fine if intended; running it on only one of them makes the comparison meaningless. - SIMD only on one side. If
+simd128is enabled for one build and not the other, the vectorization difference swamps the optimization level. Keep target features identical. - Per-package levels ignored. The package name in
[profile.release.package.NAME]must match the crate name exactly, with hyphens as written in itsCargo.toml.
Performance note
Smaller modules also compile faster. On a mid-range phone, the -Oz build above compiled in 190 ms against 260 ms for -O3.
For a page that runs the blur once, the size-optimized build was faster end to end despite the slower kernel; for an editor
that applies filters interactively, the speed-optimized kernel won within the first ten operations.
Frequently Asked Questions
Is -Os a reasonable compromise?
Often. -Os keeps more inlining than -Oz and usually lands within a few percent of -O3 speed with most of the size saving.
Measure all three on your hot functions.
Does Emscripten behave the same way?
Yes — -Oz, -Os, -O2 and -O3 have the same trade-offs, and per-file optimization levels give the same mixed-build
option.
Can the engine’s optimizer make up the difference? Partly. The browser’s optimizing tier inlines and optimizes again, but it does not unroll or vectorize the way LLVM does, and it works under a compile-time budget. Code that LLVM did not optimize for speed stays slower.
What about -O4 in wasm-opt?
It runs more passes, including flattening, and occasionally finds extra speed for compute-heavy modules at the cost of much
longer optimization time. Worth trying for a kernel crate, rarely for a whole application.
Related
- Tuning LTO and codegen-units for Wasm — the other half of the release profile.
- Benchmarking SIMD vs scalar Wasm kernels — measuring the vectorization effect directly.
- Avoiding JIT warm-up errors in Wasm benchmarks — measuring correctly.
- Shrinking Rust Wasm with Cargo profiles — profile basics.
← Back to Wasm Optimization Flags & Size Reduction