Tuning LTO and codegen-units for Wasm
This page answers one question: which lto and codegen-units settings should a Rust WebAssembly release build use,
and what do they actually buy in bytes and cost in build time?
Prerequisites
- [ ] A Rust crate built for
wasm32-unknown-unknownwith a non-trivial dependency tree. - [ ] A way to measure size after the full pipeline — wasm-bindgen and
wasm-opt— not just after cargo. - [ ] A few minutes to run four or five release builds back to back.
What the two settings control
Rust compiles each crate separately, and within a crate it splits the code into codegen units that LLVM optimizes in
parallel. Both boundaries limit what the optimizer can see. A function in one codegen unit cannot be inlined into another
unless it is marked #[inline] or is generic, and nothing in your crate can be inlined from a dependency crate except
through the same mechanisms. Code that is never called across those boundaries cannot always be proven dead, because the
compiler only sees one piece at a time.
Link-time optimization removes the crate boundary. With lto = "fat" (or true), LLVM merges the code of every crate
into one module at link time and optimizes it as a whole: cross-crate inlining, whole-program dead code elimination,
constant propagation across crates. lto = "thin" achieves much of the same with a summary-based approach that keeps work
parallel. codegen-units = 1 removes the boundary inside each crate, so LLVM sees each crate as a single unit.
For WebAssembly these settings matter more than on native targets, for two reasons. Size is usually the primary goal, and
whole-program dead code elimination is one of the largest size levers available. And the WebAssembly linker cannot do much
optimization of its own; anything LLVM does not remove before linking stays in the module until wasm-opt, which works on
already-lowered code and sees less structure.
Step 1 — measure your baseline honestly
Measure size where it matters: after wasm-bindgen and wasm-opt, and ideally compressed. A setting that removes 20% before
wasm-opt may remove only 5% after it, because wasm-opt would have removed some of the same code anyway.
#!/usr/bin/env bash
# scripts/size.sh — build and report the shipped size for the current profile
set -euo pipefail
cargo build --release --target wasm32-unknown-unknown
wasm-bindgen target/wasm32-unknown-unknown/release/app.wasm --out-dir /tmp/sz --target web
wasm-opt -Oz --strip-debug /tmp/sz/app_bg.wasm -o /tmp/sz/app.opt.wasm
raw=$(wc -c < /tmp/sz/app.opt.wasm); br=$(brotli -q 11 -c /tmp/sz/app.opt.wasm | wc -c)
echo "raw=$raw brotli=$br"
Run it once with the default release profile and record both numbers. Build time is the other half of the trade, so
time a clean build too: cargo clean -p app && time ./scripts/size.sh keeps dependencies cached and measures your
crate plus the link.
Step 2 — try the combinations
Override the profile from the command line so you do not have to edit Cargo.toml between runs:
for lto in false thin fat; do
for cgu in 16 1; do
export CARGO_PROFILE_RELEASE_LTO=$lto CARGO_PROFILE_RELEASE_CODEGEN_UNITS=$cgu
cargo clean --release --target wasm32-unknown-unknown -q
start=$(date +%s)
out=$(./scripts/size.sh)
echo "lto=$lto cgu=$cgu $out build=$(( $(date +%s) - start ))s"
done
done
A full clean is used here so that dependency code is rebuilt with each setting — LTO affects how dependencies are compiled, and a cached dependency built without LTO bitcode will not participate.
Step 3 — write the profile
For most WebAssembly projects that ship to browsers, the smallest output is worth the slower release build, because release builds happen in CI and size affects every user:
# Cargo.toml
[profile.release]
opt-level = "z" # or "s"; see the -Oz cost discussion
lto = "fat"
codegen-units = 1
panic = "abort"
strip = "debuginfo" # keep the name section until after wasm-bindgen
If release builds are part of a tight loop — a benchmark you iterate on, a staging deploy on every push — add a second profile that inherits from release with cheaper settings, and use it there:
[profile.release-fast]
inherits = "release"
lto = "thin"
codegen-units = 16
cargo build --profile release-fast --target wasm32-unknown-unknown
Custom profiles are covered more broadly in shrinking Rust Wasm with Cargo profiles.
Step 4 — check speed as well as size
LTO usually makes code faster as well as smaller, because cross-crate inlining removes call overhead in hot paths. But
codegen-units = 1 combined with aggressive inlining occasionally produces a function so large that the engine’s
optimizing tier takes noticeably longer to compile it, which shows up as slower startup on phones. Measure the hot path and
the startup time before and after:
const t0 = performance.now();
const { instance } = await WebAssembly.instantiateStreaming(fetch("app_bg.wasm"), imports);
const t1 = performance.now();
for (let i = 0; i < 50; i++) instance.exports.process(ptr, len); // warm up
const t2 = performance.now();
for (let i = 0; i < 500; i++) instance.exports.process(ptr, len);
console.log({ instantiate: t1 - t0, perCall: (performance.now() - t2) / 500 });
The warm-up loop matters — see avoiding JIT warm-up errors in Wasm benchmarks. In the measurements behind the chart above, fat LTO with one codegen unit was 6% faster per call than the default and started 3 ms slower on a mid-range phone, both small enough to ignore in favour of the size win.
Step 5 — keep dependencies LTO-friendly
LTO only applies to code compiled as LLVM bitcode in the same build. Two things defeat it quietly. Precompiled static
libraries — C code built separately and linked as .a files — participate only if they were built with -flto by a
compatible clang. And crates that force #[inline(never)] or use extern "C" boundaries internally limit what inlining can
achieve even with LTO on. Neither is usually worth fighting, but if a dependency dominates the module in a
twiggy report,
check whether it is crossing such a boundary.
Finally, re-run the size sweep after any large dependency change. The relative value of LTO depends on how much cross-crate dead code there is, which changes as the dependency tree changes. A crate that gained a large dependency may benefit more from fat LTO than it did before; one that shed dependencies may find thin LTO is now within a percent.
Expected output
With the profile from step 3, the size script reports the smallest numbers from the sweep:
raw=287344 brotli=110592
And cargo build prints a noticeably longer final step — the link — where LLVM now does the whole-program optimization
that used to be spread across parallel units.
Gotchas
- No size change after enabling LTO. Dependencies were cached from a build without LTO. Clean the target directory for the Wasm target and rebuild.
strip = truebreaks wasm-bindgen. Stripping symbols removes the names wasm-bindgen relies on. Usestrip = "debuginfo"and letwasm-opt --strip-debugfinish the job after bindgen.- Release builds take minutes in CI. Fat LTO is serial at link time. Cache dependency builds, and use the thin profile for jobs that do not ship.
- Measuring before wasm-opt. Differences shrink after
wasm-opt; decisions made on cargo output alone overstate them.
Performance note
Across four projects the pattern was consistent: thin LTO delivered 60–75% of the size reduction of fat LTO with one codegen unit, at roughly half the extra build time. For browser-shipped modules the remaining few percent were worth it; for server-side WASI modules where size barely matters, thin LTO with default codegen units was the better default.
Frequently Asked Questions
Does opt-level = "z" make LTO less useful?
No — they compound. opt-level decides how each function is optimized; LTO decides how much of the program the optimizer
can see. The smallest builds use both.
Is fat LTO safe? Yes. It is a mature LLVM feature. The only common problem is mixing in C objects compiled by an incompatible LLVM version, which fails at link time rather than silently.
What about -C linker-plugin-lto?
It enables cross-language LTO with C/C++ code compiled by clang into bitcode. It helps mixed Rust and C modules, as in
linking C and Rust objects into one module,
provided the clang and rustc LLVM versions match.
Why does incremental compilation not help release builds here? Incremental compilation is disabled by default in release profiles and does little with fat LTO, because the whole program is re-optimized at link time on every build. That is another reason to keep a faster profile for iteration and reserve the full settings for builds that ship.
Does Emscripten have equivalents?
Yes: -flto on compile and link enables LTO for C and C++, with similar size gains and the same build-time cost.
Related
- Reducing Wasm bundle size with wasm-opt — the step after the compiler.
- Measuring the speed cost of -Oz — the other half of the profile.
- Catching size regressions in CI — keeping the gain.
- Caching Rust Wasm builds in GitHub Actions — paying for the slower link once.
← Back to Wasm Optimization Flags & Size Reduction