Reducing Generic Monomorphization Bloat in Rust
This page answers one task: a Rust WebAssembly module is much larger than its logic suggests, and size analysis shows the same function appearing many times with different type parameters — so you want to reduce the duplication without giving up the generic APIs that made the code pleasant to write.
Prerequisites
- [ ] A Rust crate compiled to Wasm with a size problem, and
twiggyinstalled. - [ ] A build that keeps function names for analysis.
- [ ] Benchmarks for hot paths, so you can check that changes do not slow them down.
How monomorphization multiplies code
Rust compiles generic functions by monomorphization: for every distinct set of type arguments used in the program, the compiler generates a separate
copy of the function specialised for those types. fn process<T: Into<String>>(x: T) called with &str, String and Cow<str> becomes three
functions. That is why generic code is fast — each copy is optimised for its concrete types, with calls inlined and no indirection — and why it is big.
The effect compounds: a generic function that calls other generic functions instantiates all of them per type, and generic containers such as
HashMap<K, V> instantiate their entire implementation for each key and value pair in use.
In native binaries a few hundred extra kilobytes rarely matter. In a WebAssembly module downloaded on every visit, they do. Common sources are serde
derives (a serializer and deserializer per type, per format), iterator chains with closures (each closure is a distinct type), impl AsRef<Path> or
impl Into<String> convenience parameters, and numeric code generic over f32 and f64 that is used with both.
Step 1 — find duplicated instantiations
twiggy monos app_bg.wasm | head -30
# Apprx. Bloat Bytes │ Apprx. Bloat % │ Bytes │ % │ Monomorphizations
# ────────────────────┼────────────────┼───────┼───────┼─────────────────────────────
# 21304 ┊ 5.20% ┊ 28406 ┊ 6.93% ┊ app::geometry::simplify
# ┊ 14203 ┊ 3.46% ┊ app::geometry::simplify<f64>
# ┊ 14203 ┊ 3.46% ┊ app::geometry::simplify<f32>
# 11820 ┊ 2.88% ┊ 15760 ┊ 3.84% ┊ hashbrown::raw::RawTable<T>::reserve_rehash
twiggy monos groups instantiations of the same generic function and estimates the “bloat” — bytes that could be saved if only one copy existed. The
top entries are the candidates. cargo llvm-lines --release --target wasm32-unknown-unknown gives a similar view from the compiler’s side, counting LLVM IR
lines per generic function, which tracks compile time and code size.
Step 2 — apply the inner non-generic function pattern
The most effective fix keeps the generic signature for convenience but moves the body into a non-generic inner function, so only a thin conversion is duplicated:
// before: the whole body is instantiated for every T
pub fn load<P: AsRef<Path>>(path: P) -> io::Result<Config> {
let text = std::fs::read_to_string(path.as_ref())?;
parse_config(&text) // plus everything inlined here, per T
}
// after: one copy of the body, tiny generic shims
pub fn load<P: AsRef<Path>>(path: P) -> io::Result<Config> {
fn inner(path: &Path) -> io::Result<Config> {
let text = std::fs::read_to_string(path)?;
parse_config(&text)
}
inner(path.as_ref())
}
The standard library uses this pattern throughout (std::fs::read is written exactly this way). It changes no behaviour and no API, and the compiler
usually inlines the shim, so performance is unchanged.
Step 3 — use trait objects where speed does not depend on specialisation
For code that is generic over behaviour rather than over data layout — callbacks, writers, visitors — dyn Trait gives one copy of the code with an
indirect call instead of one copy per implementation:
// one instantiation per closure type
pub fn walk<F: FnMut(&Node)>(root: &Node, mut f: F) { /* ... */ }
// one copy; each call goes through a vtable
pub fn walk(root: &Node, f: &mut dyn FnMut(&Node)) { /* ... */ }
Indirect calls in WebAssembly are call_indirect through a table, which costs a bounds and signature check plus the call — a few nanoseconds. In code
called per element in a hot loop that can matter; in code called per document, per frame or per request it does not. Use benchmarks to decide per
function.
Step 4 — reduce the set of types
Sometimes the simplest fix is fewer type arguments. If geometry code is used with both f32 and f64 only because different callers chose differently,
standardise on one. If a function accepts impl Into<String> but every caller already has a String, take String. If serde derives produce code for
types that are only ever serialized, derive only Serialize. If two HashMap instantiations differ only in value type, consider storing values in a
Vec and mapping keys to indices with one map type. Each removed type argument removes an entire tree of instantiations.
Step 5 — measure size and speed together
Make one change at a time and record compressed size and the relevant benchmarks:
cargo build --release --target wasm32-unknown-unknown
wasm-opt -Oz -o out.wasm target/wasm32-unknown-unknown/release/app.wasm
brotli -c out.wasm | wc -c
twiggy monos out.wasm | head -5
Duplicated generic code compresses well, because the copies are similar, so compressed savings are typically half or less of raw savings. Still, they add up, and fewer instantiations also shorten compile times — often noticeably in large crates.
Serde and other derive-heavy code
Serde is a frequent source of bloat in Wasm modules: each derived type gets serializer and deserializer code for each data format, and deserializers
are large because they handle every field order, missing fields and error reporting. Options include deserializing into fewer, flatter types;
using serde-wasm-bindgen to convert directly between JavaScript values and Rust types instead of going through JSON text; using a compact binary
format with a smaller implementation; or, for small fixed configurations, a hand-written parser. miniserde and nanoserde trade features for
dramatically smaller code and suit simple data. Measure with twiggy top filtered to serde names to see which types cost the most.
Iterator chains and closures
Long iterator chains with closures are monomorphized per closure type, and closures are unique types even when their bodies are identical. In
performance-critical code that is a feature; elsewhere, prefer plain loops for one-off transformations in cold code, and share helper functions
instead of repeating similar closures at many call sites. Optimisation at opt-level = "z" reduces inlining, which helps, but the instantiations
remain distinct functions.
Designing new APIs to avoid bloat
Retrofitting is harder than designing for size from the start. For crates that will be compiled to WebAssembly, a few conventions keep instantiations
in check. Accept concrete borrowed types — &str, &[u8], &Path — in internal functions, and reserve generic conversion parameters for the outermost
public API, where the inner-function pattern applies. Make generic types thin wrappers over non-generic cores: a Cache<K, V> can store keys and values
as bytes or behind a trait object in a non-generic engine, with the generic layer only converting at the edges. Use generics for data whose layout the
algorithm depends on — numeric kernels, containers in hot loops — and trait objects for behaviour plugged in from outside. Review cargo llvm-lines
output when adding a dependency or a new generic abstraction, and treat a large jump the same way you would treat a size-budget failure.
When duplication is worth keeping
Not every duplicate should go. A numeric kernel specialised for f32 and f64 may run twice as fast as a version that converts, and a SIMD path
specialised per lane type is the point of the code. Keep instantiations where benchmarks show they pay for themselves, and document why, so a later
size pass does not undo a deliberate choice.
Expected output
twiggy monos shows the top duplicated function reduced from two copies to one; load and four similar convenience APIs use inner functions; the
tree walker takes &mut dyn FnMut; geometry code is standardised on f64; the compressed module shrank from 186 KB to 151 KB; and benchmarks for the
hot paths are within noise of the previous build.
Gotchas
- Converting hot inner loops to
dyn. Indirect calls per element cost measurable time. Benchmark first. - Optimising by raw bytes only. Duplicates compress well. Measure compressed size.
- Forgetting closures are unique types. Identical-looking closures still instantiate separately.
- Generic convenience APIs everywhere. Each
impl Into<…>parameter multiplies callers’ code. Use the inner-function pattern. - Serde derives on every type. Derive only what is needed, and only for the formats used.
Performance note
Applying the inner-function pattern to 14 public functions, switching two callbacks to dyn, and standardising on f64 removed 35 KB compressed. The
tree-walk benchmark slowed by 1.2% from the indirect callback, within the project’s tolerance.
Frequently Asked Questions
Does LTO remove duplicate instantiations? It merges identical code in some cases, but instantiations for different types differ and remain.
Does opt-level = "z" help?
It inlines less, which reduces size, but does not remove the instantiations themselves.
Is dyn always slower?
Only by the indirect call; in code not dominated by call overhead the difference is unmeasurable.
Can wasm-opt merge duplicates?
wasm-opt’s duplicate-function elimination merges byte-identical functions, which some instantiations are — but most are not.
How do I see which generic functions cost the most compile time and size?
cargo llvm-lines lists LLVM IR lines and copies per generic function; it tracks both closely.
Related
- Finding the largest functions in a Wasm binary — where to look first.
- Removing panic and formatting bloat from Rust Wasm — another common source.
- Shrinking Rust Wasm with Cargo profiles — profile settings.
- Tuning LTO and codegen units for Wasm — whole-program options.
← Back to Wasm Optimization Flags & Size Reduction