Linking Wasm Objects with LTO
This page answers one task: you want the optimiser to see the whole WebAssembly program at link time — inlining across files and languages, removing more dead code — and you need to enable link-time optimisation correctly, confirm it ran, and judge whether its build-time cost is worth the result.
Prerequisites
- [ ] clang/wasi-sdk or Emscripten for C/C++, and/or Rust.
- [ ] A release build you can measure for size and speed.
- [ ] Matching LLVM versions between compilers if you mix C and Rust.
What LTO changes
Normally each source file (or Rust codegen unit) is compiled to machine code — here, WebAssembly — independently, and the linker joins the results. The
optimiser therefore cannot inline a function from one file into another, cannot see that a function exported from one file is only ever called with a
constant, and cannot remove code that is reachable only through paths it never examines together. Link-time optimisation defers code generation: objects
contain LLVM bitcode instead of WebAssembly, and wasm-ld runs the optimiser over all of them together before generating the final module.
Full LTO merges all bitcode into one module and optimises it as a whole — the best results, the slowest builds, and little parallelism. Thin LTO keeps modules separate but shares summaries between them so cross-module inlining and dead-code elimination still happen, in parallel — most of the benefit with much shorter link times. For WebAssembly, LTO usually shrinks binaries noticeably because cross-module dead-code elimination and inlining remove abstraction layers that separate compilation leaves in place.
Step 1 — enable LTO for C and C++
Compile and link with -flto (full) or -flto=thin:
clang --target=wasm32-wasip1 -O2 -flto=thin -c src/a.c -o build/a.o
clang --target=wasm32-wasip1 -O2 -flto=thin -c src/b.c -o build/b.o
clang --target=wasm32-wasip1 -O2 -flto=thin build/a.o build/b.o -o build/app.wasm
With Emscripten, -flto works the same way, and Emscripten can also use LTO-built system libraries (-flto on the link command selects them). The flag
must be on both compile and link steps; objects compiled without it are linked normally and miss out.
Step 2 — enable LTO for Rust
In Cargo.toml, set the release profile:
[profile.release]
lto = "fat" # or "thin"
codegen-units = 1 # fewer units = more optimisation within the crate graph
opt-level = "s" # or "z" / 3, depending on size versus speed goals
lto = "fat" performs full LTO across all crates; "thin" performs thin LTO; false still does thin LTO within the crate graph’s codegen units by default.
codegen-units = 1 lets the optimiser see each crate as one unit, which helps size. The interplay with other size settings is covered in
tuning LTO and codegen-units for Wasm.
Step 3 — cross-language LTO between C and Rust
When a Rust module links C code, ordinary builds optimise each side separately. Cross-language LTO lets LLVM inline C helpers into Rust callers and vice
versa. It requires that rustc and clang use compatible LLVM versions, that C objects are compiled with -flto=thin, and that rustc is told to emit and
link bitcode:
export RUSTFLAGS="-C linker-plugin-lto -C linker=clang -C link-arg=-fuse-ld=lld"
export CFLAGS="-flto=thin"
cargo build --release --target wasm32-wasip1
LLVM version mismatches produce errors like “Invalid bitcode version” or silently fall back. Check rustc -vV and clang --version; their LLVM major
versions should match. The linking mechanics of mixed modules are in
linking C and Rust objects into one module.
Step 4 — verify that LTO actually ran
Builds can silently skip LTO when a flag is missing on one step. Check the intermediate objects: LTO objects are LLVM bitcode, not WebAssembly.
file build/a.o # "LLVM IR bitcode" when -flto applied
llvm-nm build/a.o | head # works on bitcode too
Then compare the final module with and without LTO: function counts (wasm-objdump -h), size, and the presence of small helper functions that LTO should
have inlined away. A link map, as in
reading wasm-ld map files,
shows LTO output attributed to a single synthetic object, another sign it ran.
Step 5 — measure and decide
Measure size, speed on your benchmarks, and build time for no LTO, thin and full. Typical results for WebAssembly: thin LTO cuts size by 5–20% versus none,
full LTO a few percent more; speed effects are usually small and occasionally negative when aggressive inlining enlarges hot loops. Build time grows
substantially, especially for full LTO with codegen-units = 1. A common configuration is no LTO for development, thin LTO for CI builds, and full LTO only
for tagged releases.
LTO and wasm-opt together
LTO and wasm-opt overlap partly but not entirely. LTO operates on LLVM’s intermediate representation with full type information and can make decisions
across the whole program before WebAssembly is generated. wasm-opt works on the final WebAssembly and applies optimisations specific to it — local and
global reordering, constant-merging in data, code-size passes tuned for the binary format. Running both gives the best result; the order is fixed
(LTO during linking, wasm-opt afterwards). With LTO enabled, wasm-opt’s additional size reduction is smaller than without it, because LTO has already
removed much of what it would find, so measure the combination rather than assuming each adds its usual saving.
Debug information and profiling with LTO
LTO rearranges and inlines code across files, which makes debug information less precise: breakpoints in inlined functions may not hit, and stack traces show fewer frames. Keep LTO off in debug builds for a good debugging experience, and remember that production stack traces from LTO builds attribute inlined code to its caller. Profiling still works, but hot functions in profiles may be large merged functions rather than the small helpers in the source. When investigating a performance problem in an LTO build, rebuilding without LTO helps attribute costs to source functions before applying fixes.
Caching LTO builds
LTO moves most of the compile work into the link step, which defeats incremental compilation: a change to one file re-runs optimisation for the whole
program under full LTO. Thin LTO mitigates this with a cache of per-module optimisation results. wasm-ld accepts --thinlto-cache-dir=<dir> (passed
through clang as -Wl,--thinlto-cache-dir=…), and rustc maintains its own incremental caches; keeping these directories between CI runs, keyed on the
toolchain version, makes repeated thin-LTO links much faster when only a few modules change. Full LTO has no equivalent, which is another reason to reserve
it for release builds. Watch cache size: thin-LTO caches grow with every distinct build and should be pruned, which wasm-ld supports with
--thinlto-cache-policy.
Profile-guided optimisation
LTO decides what to inline and how to lay out code from static heuristics. Profile-guided optimisation (PGO) adds measurements: an instrumented build
records which branches and calls are hot on a representative workload, and the final build uses the profile to inline and optimise accordingly. PGO for
WebAssembly is possible with clang and rustc — collecting the profile by running the instrumented module under a WASI runtime or in Node — but the tooling
is less polished than for native targets, and gains vary. Consider it only after LTO and wasm-opt are in place and a profile shows the remaining time
is in branch-heavy code, and measure on the real engines, since browser JIT tiers make their own layout decisions.
Expected output
Objects report “LLVM IR bitcode”; the release module shrinks from 520 KB without LTO to 455 KB with thin LTO and 438 KB with full LTO; benchmarks change by under 2%; and CI uses thin LTO while tagged releases use full LTO, documented in the build scripts.
Gotchas
-fltoon compile but not link (or vice versa). LTO silently does not happen. Use it on both.- LLVM version mismatch for cross-language LTO. Bitcode errors or fallback. Match versions.
- Expecting speed gains. Size gains are reliable; speed effects vary. Measure.
- Full LTO in development. Builds become slow. Reserve it for releases.
- Debugging LTO builds. Inlining confuses debuggers. Debug without LTO.
Performance note
For a mixed C and Rust module, release link time was 6 s without LTO, 14 s with thin LTO and 48 s with full LTO and codegen-units = 1. Size fell by 12.5% with
thin LTO and 15.8% with full LTO.
Frequently Asked Questions
Does wasm-pack enable LTO?
It uses Cargo’s release profile; set lto there. wasm-pack itself does not change it.
Is LTO safe? Yes for correct code. It can expose undefined behaviour that separate compilation happened to hide.
Does Emscripten’s -O3 imply LTO?
No — LTO is separate; add -flto explicitly.
Can I use LTO with dynamic linking? Within each module, yes; it cannot optimise across dynamically linked modules.
Can thin LTO builds be cached between CI runs? Yes — keep the ThinLTO cache directory between runs, keyed on the toolchain version.
Related
- Tuning LTO and codegen-units for Wasm — Rust profile settings.
- Controlling dead-code elimination in wasm-ld — what LTO removes further.
- Comparing wasm-opt optimization levels — the post-link stage.
- Shrinking Rust Wasm with cargo profiles — profile-wide settings.
← Back to Linking Wasm Objects with wasm-ld