Tuning LTO and codegen-units for Wasm

This page answers one question: which lto and codegen-units settings should a Rust WebAssembly release build use, and what do they actually buy in bytes and cost in build time?

Prerequisites

  • [ ] A Rust crate built for wasm32-unknown-unknown with a non-trivial dependency tree.
  • [ ] A way to measure size after the full pipeline — wasm-bindgen and wasm-opt — not just after cargo.
  • [ ] A few minutes to run four or five release builds back to back.

What the two settings control

Rust compiles each crate separately, and within a crate it splits the code into codegen units that LLVM optimizes in parallel. Both boundaries limit what the optimizer can see. A function in one codegen unit cannot be inlined into another unless it is marked #[inline] or is generic, and nothing in your crate can be inlined from a dependency crate except through the same mechanisms. Code that is never called across those boundaries cannot always be proven dead, because the compiler only sees one piece at a time.

Link-time optimization removes the crate boundary. With lto = "fat" (or true), LLVM merges the code of every crate into one module at link time and optimizes it as a whole: cross-crate inlining, whole-program dead code elimination, constant propagation across crates. lto = "thin" achieves much of the same with a summary-based approach that keeps work parallel. codegen-units = 1 removes the boundary inside each crate, so LLVM sees each crate as a single unit.

For WebAssembly these settings matter more than on native targets, for two reasons. Size is usually the primary goal, and whole-program dead code elimination is one of the largest size levers available. And the WebAssembly linker cannot do much optimization of its own; anything LLVM does not remove before linking stays in the module until wasm-opt, which works on already-lowered code and sees less structure.

What the optimizer sees under each setting Without LTO, each crate and each codegen unit is optimized in isolation. Thin LTO shares summaries across crates for cross-crate inlining in parallel. Fat LTO merges everything into one module, and with codegen-units equal to one the optimizer sees the whole program at once. lto = false (default) each crate optimized alone 16 codegen units per crate dead code kept across boundaries fastest build, largest output lto = "thin" cross-crate summaries parallel optimization kept most inlining wins good middle ground fat + codegen-units = 1 one module, whole-program view maximum dead code removal serial, slowest link smallest release output

Step 1 — measure your baseline honestly

Measure size where it matters: after wasm-bindgen and wasm-opt, and ideally compressed. A setting that removes 20% before wasm-opt may remove only 5% after it, because wasm-opt would have removed some of the same code anyway.

#!/usr/bin/env bash
# scripts/size.sh — build and report the shipped size for the current profile
set -euo pipefail
cargo build --release --target wasm32-unknown-unknown
wasm-bindgen target/wasm32-unknown-unknown/release/app.wasm --out-dir /tmp/sz --target web
wasm-opt -Oz --strip-debug /tmp/sz/app_bg.wasm -o /tmp/sz/app.opt.wasm
raw=$(wc -c < /tmp/sz/app.opt.wasm); br=$(brotli -q 11 -c /tmp/sz/app.opt.wasm | wc -c)
echo "raw=$raw brotli=$br"

Run it once with the default release profile and record both numbers. Build time is the other half of the trade, so time a clean build too: cargo clean -p app && time ./scripts/size.sh keeps dependencies cached and measures your crate plus the link.

Step 2 — try the combinations

Override the profile from the command line so you do not have to edit Cargo.toml between runs:

for lto in false thin fat; do
  for cgu in 16 1; do
    export CARGO_PROFILE_RELEASE_LTO=$lto CARGO_PROFILE_RELEASE_CODEGEN_UNITS=$cgu
    cargo clean --release --target wasm32-unknown-unknown -q
    start=$(date +%s)
    out=$(./scripts/size.sh)
    echo "lto=$lto cgu=$cgu $out build=$(( $(date +%s) - start ))s"
  done
done

A full clean is used here so that dependency code is rebuilt with each setting — LTO affects how dependencies are compiled, and a cached dependency built without LTO bitcode will not participate.

Shipped size by LTO and codegen-units setting The same crate's module after wasm-bindgen and wasm-opt -Oz, Brotli-compressed, for six combinations of lto and codegen-units. Fat LTO with one codegen unit is smallest; most of the gain arrives with thin LTO. Brotli-compressed bytes (KB), after wasm-opt -Oz lto=false, cgu=16 148 KB lto=false, cgu=1 139 KB lto=thin, cgu=16 121 KB lto=thin, cgu=1 117 KB lto=fat, cgu=16 112 KB lto=fat, cgu=1 108 KB Clean release build time rose from 38 s (no LTO, 16 units) to 71 s (fat LTO, 1 unit) on the same machine.

Step 3 — write the profile

For most WebAssembly projects that ship to browsers, the smallest output is worth the slower release build, because release builds happen in CI and size affects every user:

# Cargo.toml
[profile.release]
opt-level = "z"        # or "s"; see the -Oz cost discussion
lto = "fat"
codegen-units = 1
panic = "abort"
strip = "debuginfo"    # keep the name section until after wasm-bindgen
The release profile, line by line An annotated Cargo release profile for a browser-shipped Wasm module, showing what each setting contributes. [profile.release] opt-level = "z" optimize every function for size lto = "fat" whole-program inlining and dead code removal codegen-units = 1 one unit per crate, no split boundaries panic = "abort" no unwinding tables or landing pads strip = "debuginfo" keep names for wasm-bindgen, drop DWARF

If release builds are part of a tight loop — a benchmark you iterate on, a staging deploy on every push — add a second profile that inherits from release with cheaper settings, and use it there:

[profile.release-fast]
inherits = "release"
lto = "thin"
codegen-units = 16
cargo build --profile release-fast --target wasm32-unknown-unknown

Custom profiles are covered more broadly in shrinking Rust Wasm with Cargo profiles.

Step 4 — check speed as well as size

LTO usually makes code faster as well as smaller, because cross-crate inlining removes call overhead in hot paths. But codegen-units = 1 combined with aggressive inlining occasionally produces a function so large that the engine’s optimizing tier takes noticeably longer to compile it, which shows up as slower startup on phones. Measure the hot path and the startup time before and after:

const t0 = performance.now();
const { instance } = await WebAssembly.instantiateStreaming(fetch("app_bg.wasm"), imports);
const t1 = performance.now();
for (let i = 0; i < 50; i++) instance.exports.process(ptr, len);   // warm up
const t2 = performance.now();
for (let i = 0; i < 500; i++) instance.exports.process(ptr, len);
console.log({ instantiate: t1 - t0, perCall: (performance.now() - t2) / 500 });

The warm-up loop matters — see avoiding JIT warm-up errors in Wasm benchmarks. In the measurements behind the chart above, fat LTO with one codegen unit was 6% faster per call than the default and started 3 ms slower on a mid-range phone, both small enough to ignore in favour of the size win.

Step 5 — keep dependencies LTO-friendly

LTO only applies to code compiled as LLVM bitcode in the same build. Two things defeat it quietly. Precompiled static libraries — C code built separately and linked as .a files — participate only if they were built with -flto by a compatible clang. And crates that force #[inline(never)] or use extern "C" boundaries internally limit what inlining can achieve even with LTO on. Neither is usually worth fighting, but if a dependency dominates the module in a twiggy report, check whether it is crossing such a boundary.

Finally, re-run the size sweep after any large dependency change. The relative value of LTO depends on how much cross-crate dead code there is, which changes as the dependency tree changes. A crate that gained a large dependency may benefit more from fat LTO than it did before; one that shed dependencies may find thin LTO is now within a percent.

Expected output

With the profile from step 3, the size script reports the smallest numbers from the sweep:

raw=287344 brotli=110592

And cargo build prints a noticeably longer final step — the link — where LLVM now does the whole-program optimization that used to be spread across parallel units.

Gotchas

  • No size change after enabling LTO. Dependencies were cached from a build without LTO. Clean the target directory for the Wasm target and rebuild.
  • strip = true breaks wasm-bindgen. Stripping symbols removes the names wasm-bindgen relies on. Use strip = "debuginfo" and let wasm-opt --strip-debug finish the job after bindgen.
  • Release builds take minutes in CI. Fat LTO is serial at link time. Cache dependency builds, and use the thin profile for jobs that do not ship.
  • Measuring before wasm-opt. Differences shrink after wasm-opt; decisions made on cargo output alone overstate them.

Performance note

Across four projects the pattern was consistent: thin LTO delivered 60–75% of the size reduction of fat LTO with one codegen unit, at roughly half the extra build time. For browser-shipped modules the remaining few percent were worth it; for server-side WASI modules where size barely matters, thin LTO with default codegen units was the better default.

Frequently Asked Questions

Does opt-level = "z" make LTO less useful? No — they compound. opt-level decides how each function is optimized; LTO decides how much of the program the optimizer can see. The smallest builds use both.

Is fat LTO safe? Yes. It is a mature LLVM feature. The only common problem is mixing in C objects compiled by an incompatible LLVM version, which fails at link time rather than silently.

What about -C linker-plugin-lto? It enables cross-language LTO with C/C++ code compiled by clang into bitcode. It helps mixed Rust and C modules, as in linking C and Rust objects into one module, provided the clang and rustc LLVM versions match.

Why does incremental compilation not help release builds here? Incremental compilation is disabled by default in release profiles and does little with fat LTO, because the whole program is re-optimized at link time on every build. That is another reason to keep a faster profile for iteration and reserve the full settings for builds that ship.

Does Emscripten have equivalents? Yes: -flto on compile and link enables LTO for C and C++, with similar size gains and the same build-time cost.

← Back to Wasm Optimization Flags & Size Reduction