Reducing Generic Monomorphization Bloat in Rust

This page answers one task: a Rust WebAssembly module is much larger than its logic suggests, and size analysis shows the same function appearing many times with different type parameters — so you want to reduce the duplication without giving up the generic APIs that made the code pleasant to write.

Prerequisites

  • [ ] A Rust crate compiled to Wasm with a size problem, and twiggy installed.
  • [ ] A build that keeps function names for analysis.
  • [ ] Benchmarks for hot paths, so you can check that changes do not slow them down.

How monomorphization multiplies code

Rust compiles generic functions by monomorphization: for every distinct set of type arguments used in the program, the compiler generates a separate copy of the function specialised for those types. fn process<T: Into<String>>(x: T) called with &str, String and Cow<str> becomes three functions. That is why generic code is fast — each copy is optimised for its concrete types, with calls inlined and no indirection — and why it is big. The effect compounds: a generic function that calls other generic functions instantiates all of them per type, and generic containers such as HashMap<K, V> instantiate their entire implementation for each key and value pair in use.

In native binaries a few hundred extra kilobytes rarely matter. In a WebAssembly module downloaded on every visit, they do. Common sources are serde derives (a serializer and deserializer per type, per format), iterator chains with closures (each closure is a distinct type), impl AsRef<Path> or impl Into<String> convenience parameters, and numeric code generic over f32 and f64 that is used with both.

One generic function, many copies A generic function written once is instantiated by the compiler for each type it is called with. Calls with str, String and Cow produce three specialised copies, and every generic function those copies call is instantiated again for each type, multiplying the code. fn parse<T: AsRef<str>> written once called with 3 types &str, String, Cow 3 specialised copies each fully optimised callees instantiated too per type again module grows duplicate bodies

Step 1 — find duplicated instantiations

twiggy monos app_bg.wasm | head -30
#  Apprx. Bloat Bytes │ Apprx. Bloat % │ Bytes │ %     │ Monomorphizations
# ────────────────────┼────────────────┼───────┼───────┼─────────────────────────────
#               21304 ┊          5.20% ┊ 28406 ┊ 6.93% ┊ app::geometry::simplify
#                                      ┊ 14203 ┊ 3.46% ┊   app::geometry::simplify<f64>
#                                      ┊ 14203 ┊ 3.46% ┊   app::geometry::simplify<f32>
#               11820 ┊          2.88% ┊ 15760 ┊ 3.84% ┊ hashbrown::raw::RawTable<T>::reserve_rehash

twiggy monos groups instantiations of the same generic function and estimates the “bloat” — bytes that could be saved if only one copy existed. The top entries are the candidates. cargo llvm-lines --release --target wasm32-unknown-unknown gives a similar view from the compiler’s side, counting LLVM IR lines per generic function, which tracks compile time and code size.

Step 2 — apply the inner non-generic function pattern

The most effective fix keeps the generic signature for convenience but moves the body into a non-generic inner function, so only a thin conversion is duplicated:

// before: the whole body is instantiated for every T
pub fn load<P: AsRef<Path>>(path: P) -> io::Result<Config> {
    let text = std::fs::read_to_string(path.as_ref())?;
    parse_config(&text)                     // plus everything inlined here, per T
}

// after: one copy of the body, tiny generic shims
pub fn load<P: AsRef<Path>>(path: P) -> io::Result<Config> {
    fn inner(path: &Path) -> io::Result<Config> {
        let text = std::fs::read_to_string(path)?;
        parse_config(&text)
    }
    inner(path.as_ref())
}

The standard library uses this pattern throughout (std::fs::read is written exactly this way). It changes no behaviour and no API, and the compiler usually inlines the shim, so performance is unchanged.

Step 3 — use trait objects where speed does not depend on specialisation

For code that is generic over behaviour rather than over data layout — callbacks, writers, visitors — dyn Trait gives one copy of the code with an indirect call instead of one copy per implementation:

// one instantiation per closure type
pub fn walk<F: FnMut(&Node)>(root: &Node, mut f: F) { /* ... */ }

// one copy; each call goes through a vtable
pub fn walk(root: &Node, f: &mut dyn FnMut(&Node)) { /* ... */ }

Indirect calls in WebAssembly are call_indirect through a table, which costs a bounds and signature check plus the call — a few nanoseconds. In code called per element in a hot loop that can matter; in code called per document, per frame or per request it does not. Use benchmarks to decide per function.

Generic parameter versus dyn trait object A generic parameter creates one specialised, inlinable copy per type, which is fastest but multiplies code size. A dyn trait object compiles the function once and dispatches through a table call, adding a small per-call cost while keeping a single copy. F: FnMut(&Node) (generic) one copy per closure type calls inlined, fastest size grows with call sites hot inner loops &mut dyn FnMut(&Node) one copy in total call_indirect per call size independent of call sites everything else

Step 4 — reduce the set of types

Sometimes the simplest fix is fewer type arguments. If geometry code is used with both f32 and f64 only because different callers chose differently, standardise on one. If a function accepts impl Into<String> but every caller already has a String, take String. If serde derives produce code for types that are only ever serialized, derive only Serialize. If two HashMap instantiations differ only in value type, consider storing values in a Vec and mapping keys to indices with one map type. Each removed type argument removes an entire tree of instantiations.

Step 5 — measure size and speed together

Make one change at a time and record compressed size and the relevant benchmarks:

cargo build --release --target wasm32-unknown-unknown
wasm-opt -Oz -o out.wasm target/wasm32-unknown-unknown/release/app.wasm
brotli -c out.wasm | wc -c
twiggy monos out.wasm | head -5

Duplicated generic code compresses well, because the copies are similar, so compressed savings are typically half or less of raw savings. Still, they add up, and fewer instantiations also shorten compile times — often noticeably in large crates.

Serde and other derive-heavy code

Serde is a frequent source of bloat in Wasm modules: each derived type gets serializer and deserializer code for each data format, and deserializers are large because they handle every field order, missing fields and error reporting. Options include deserializing into fewer, flatter types; using serde-wasm-bindgen to convert directly between JavaScript values and Rust types instead of going through JSON text; using a compact binary format with a smaller implementation; or, for small fixed configurations, a hand-written parser. miniserde and nanoserde trade features for dramatically smaller code and suit simple data. Measure with twiggy top filtered to serde names to see which types cost the most.

Iterator chains and closures

Long iterator chains with closures are monomorphized per closure type, and closures are unique types even when their bodies are identical. In performance-critical code that is a feature; elsewhere, prefer plain loops for one-off transformations in cold code, and share helper functions instead of repeating similar closures at many call sites. Optimisation at opt-level = "z" reduces inlining, which helps, but the instantiations remain distinct functions.

Designing new APIs to avoid bloat

Retrofitting is harder than designing for size from the start. For crates that will be compiled to WebAssembly, a few conventions keep instantiations in check. Accept concrete borrowed types — &str, &[u8], &Path — in internal functions, and reserve generic conversion parameters for the outermost public API, where the inner-function pattern applies. Make generic types thin wrappers over non-generic cores: a Cache<K, V> can store keys and values as bytes or behind a trait object in a non-generic engine, with the generic layer only converting at the edges. Use generics for data whose layout the algorithm depends on — numeric kernels, containers in hot loops — and trait objects for behaviour plugged in from outside. Review cargo llvm-lines output when adding a dependency or a new generic abstraction, and treat a large jump the same way you would treat a size-budget failure.

When duplication is worth keeping

Not every duplicate should go. A numeric kernel specialised for f32 and f64 may run twice as fast as a version that converts, and a SIMD path specialised per lane type is the point of the code. Keep instantiations where benchmarks show they pay for themselves, and document why, so a later size pass does not undo a deliberate choice.

Expected output

twiggy monos shows the top duplicated function reduced from two copies to one; load and four similar convenience APIs use inner functions; the tree walker takes &mut dyn FnMut; geometry code is standardised on f64; the compressed module shrank from 186 KB to 151 KB; and benchmarks for the hot paths are within noise of the previous build.

Gotchas

  • Converting hot inner loops to dyn. Indirect calls per element cost measurable time. Benchmark first.
  • Optimising by raw bytes only. Duplicates compress well. Measure compressed size.
  • Forgetting closures are unique types. Identical-looking closures still instantiate separately.
  • Generic convenience APIs everywhere. Each impl Into<…> parameter multiplies callers’ code. Use the inner-function pattern.
  • Serde derives on every type. Derive only what is needed, and only for the formats used.

Performance note

Applying the inner-function pattern to 14 public functions, switching two callbacks to dyn, and standardising on f64 removed 35 KB compressed. The tree-walk benchmark slowed by 1.2% from the indirect callback, within the project’s tolerance.

Compressed module size after each change Brotli-compressed kilobytes after applying the inner non-generic function pattern, switching callbacks to dyn trait objects, and standardising numeric code on f64. KB compressed before 186 KB inner non-generic functions 168 KB dyn callbacks 160 KB f64 only 151 KB

Frequently Asked Questions

Does LTO remove duplicate instantiations? It merges identical code in some cases, but instantiations for different types differ and remain.

Does opt-level = "z" help? It inlines less, which reduces size, but does not remove the instantiations themselves.

Is dyn always slower? Only by the indirect call; in code not dominated by call overhead the difference is unmeasurable.

Can wasm-opt merge duplicates? wasm-opt’s duplicate-function elimination merges byte-identical functions, which some instantiations are — but most are not.

How do I see which generic functions cost the most compile time and size? cargo llvm-lines lists LLVM IR lines and copies per generic function; it tracks both closely.

← Back to Wasm Optimization Flags & Size Reduction