Finding the Largest Functions in a Wasm Binary

This page answers one task: a WebAssembly module is larger than you want, and before changing flags or rewriting code you need to know where the bytes are — which sections, which functions, which dependencies — so effort goes to the biggest wins first.

Prerequisites

  • [ ] The release .wasm file you ship, plus a build of it with the name section kept.
  • [ ] WABT’s wasm-objdump, and twiggy (cargo install twiggy).
  • [ ] Optionally Binaryen’s wasm-opt and wasm-tools.

Where the bytes in a module live

A module is a sequence of sections. The code section holds function bodies and is almost always the largest. The data section holds static data — string literals, lookup tables, embedded assets. The name and other custom sections hold debug information and names; they can be large but are usually stripped for release. Smaller sections — types, imports, exports, the function table — rarely matter, though a module with thousands of exports pays for every export name.

Within the code section, size concentrates in a few functions. A typical Rust or C++ module has a long tail of small functions and a handful of large ones: a formatting routine, a monomorphised generic instantiated for many types, a parser’s state machine, an inlined allocator. Finding them gives a ranked list of candidates; the next question for each is why it exists and whether it is needed.

Section sizes of a typical Rust module Kilobytes per section in a 410 KB release module before stripping. The code section dominates, followed by the name section and static data. Types, imports, exports and the function table together are small. KB per section code 268 KB name (custom) 92 KB data 41 KB types + imports + exports + table 9 KB

Step 1 — list section sizes

wasm-objdump -h app_bg.wasm
#     Type start=0x0000000a end=0x000002f4 (size=0x000002ea) count: 64
#   Import start=...                       (size=0x00000412) count: 38
#     Code start=0x00002c71 end=0x00044d36 (size=0x0004209c) count: 1214
#     Data start=...                       (size=0x0000a3f1) count: 3
#   Custom start=...                       (size=0x00016f20) "name"

If custom sections are large in the file you ship, stripping them is the cheapest win: wasm-opt --strip-debug --strip-producers, wasm-strip, or strip = true in the Cargo profile. If data is large, look for embedded assets or tables that could be loaded separately. Usually, though, the code section is the story.

Step 2 — rank functions with twiggy top

twiggy needs function names to be useful, so analyse a build that keeps the name section (or the unstripped artefact before your final strip step):

twiggy top -n 20 app_bg.wasm
#  Shallow Bytes │ Shallow % │ Item
# ───────────────┼───────────┼──────────────────────────────────────────────
#          18422 ┊     4.49% ┊ core::fmt::Formatter::pad_integral
#          12904 ┊     3.15% ┊ serde_json::de::Deserializer::parse_value
#          11730 ┊     2.86% ┊ dlmalloc::dlmalloc::Dlmalloc::malloc
#           9812 ┊     2.39% ┊ app::render::<impl Renderer<f64>>::draw
#           9611 ┊     2.34% ┊ app::render::<impl Renderer<f32>>::draw
#           ...

Shallow size is the function’s own body. The list above already suggests stories: formatting machinery, a JSON parser, the allocator, and the same generic draw instantiated twice.

Step 3 — use retained size to find what pulls code in

A small function can keep a large amount of code alive — a single call to format! keeps the whole formatting machinery. twiggy’s retained size counts a function plus everything only reachable through it:

twiggy dominators app_bg.wasm | head -40
twiggy paths app_bg.wasm 'core::fmt::Formatter::pad_integral'   # who calls it?

The dominator tree shows which items, if removed, would let the most code disappear. twiggy paths answers “why is this here?” by listing call paths from exports to the function — often revealing that one debug-only to_string() or an unwrap() with a formatted panic message keeps kilobytes alive.

What the render export retains The export render retains 61 percent of the code. Within it, the config loader retains serde_json at 19 percent, the panic handler retains the formatting machinery at 14 percent, and two instantiations of the generic draw function retain 12 and 11 percent. Removing a retaining item removes its whole share. export render retains 61% of code load_config 19%: serde_json panic handler 14%: core::fmt draw<f64> 12% draw<f32> 11%

Step 4 — attribute size to crates or libraries

Summing function sizes by crate turns a list of symbols into a list of dependencies to question. twiggy can emit CSV or JSON for a short script:

twiggy top -n 100000 --format json app_bg.wasm \
  | jq -r '.[] | [(.name | capture("^(<?(?<c>[a-z_0-9]+)::)").c // "other"), .shallow_size] | @tsv' \
  | awk '{s[$1]+=$2} END {for (c in s) print s[c], c}' | sort -rn | head

For C and C++ modules built with Emscripten, the names show library prefixes (std::__2::, _ZNSt), and --emit-symbol-map gives a mapping for minified names. In practice, a crate-level summary often shows two or three dependencies responsible for half the code.

Step 5 — decide what to cut first

Rank candidates by size times feasibility. The common big wins, roughly in order of effort: strip custom sections; remove formatting from panic paths (removing panic and formatting bloat from Rust Wasm); replace heavy dependencies (serde_json for a small config, a full regex engine for one pattern); collapse duplicated generic instantiations (reducing generic monomorphization bloat in Rust); switch allocator; and finally tune optimisation levels. Re-run twiggy after each change — removing one function sometimes reveals that the next largest item was retained by it too.

Comparing two builds

When a module grows unexpectedly, compare the two versions directly:

twiggy diff old_bg.wasm new_bg.wasm | head -30

The diff lists items that appeared, disappeared or changed size, which usually points straight at the dependency or code change responsible. Run it in CI next to a size budget, as described in catching size regressions in CI, and post the top of the diff in the pull request when the budget fails.

Measuring what users download

Raw sizes rank candidates, but users download compressed bytes. Highly repetitive code — tables, many similar generic instantiations — compresses well, so its cost after Brotli is lower than its raw share suggests, while dense unique code compresses less. After identifying candidates, measure the compressed size of the whole module before and after each change (brotli -c app_bg.wasm | wc -c) and report that number; a change that saves 30 KB raw might save 6 KB compressed, and a change that saves 15 KB raw might save 9 KB.

Reading names in optimised builds

Optimisation changes what you see. Inlining merges small functions into their callers, so a dependency’s code may show up under your function’s name rather than its own; wasm-opt may merge identical functions, so one entry stands for several; and outlined or split functions get synthetic names. For attribution by crate, analyse the module after wasm-opt runs, because that is what ships, but keep in mind that a large application function may contain inlined library code. When a result is surprising, build a variant with less inlining (-C inline-threshold is gone in recent Rust, but opt-level = "s" and -Os inline less than -O3) and analyse that too — it shows which libraries contributed code before inlining blurred the boundaries. For Emscripten builds, --profiling-funcs keeps function names in an optimised build without full debug info, which is exactly what size analysis needs.

Data segments and embedded assets

When the data section is large, inspect it rather than guessing. wasm-objdump -x -j Data app_bg.wasm lists segments with their offsets and sizes, and wasm-objdump -s -j Data dumps contents, where long runs of readable text reveal embedded strings: panic messages with file paths, help text, error message tables, Unicode tables pulled in by a text library. File paths in panic messages are a common surprise — every unwrap() location embeds the source path — and are removed along with formatting by the panic-bloat techniques. Large binary tables, such as fonts or lookup tables, are often better fetched as separate files at runtime, compressed and cached independently of the code.

Expected output

A ranked table of the top 20 functions with shallow and retained sizes; a crate summary showing serde_json at 19%, core::fmt at 14% and the allocator at 4%; a dominator path showing that one formatted panic message retains the formatting code; and a list of three changes ordered by expected compressed savings.

Gotchas

  • Running twiggy on a stripped module. Every function is anonymous. Analyse a build with names.
  • Optimising by shallow size only. Retained size shows what actually disappears.
  • Ignoring the data section. Embedded assets can be large. Check section sizes first.
  • Measuring raw bytes only. Users download compressed bytes. Measure both.
  • Fixing everything at once. Change one thing, re-measure, repeat.

Performance note

On the analysed module, the three highest-ranked changes — stripping names, removing panic formatting and replacing serde_json with a small hand-written parser for the config — reduced the compressed size from 142 KB to 79 KB, without touching optimisation flags.

Compressed size after each ranked change Brotli-compressed kilobytes of the module after applying, in order, stripping custom sections, removing panic formatting and replacing serde_json for configuration loading. KB compressed initial release build 142 KB stripped custom sections 118 KB no panic formatting 96 KB serde_json replaced 79 KB

Frequently Asked Questions

Does twiggy work on non-Rust modules? Yes — it analyses any Wasm module; names come from the name section regardless of language.

What about wasm-opt --metrics? It reports counts of instruction kinds per function or module, useful alongside twiggy for spotting unusual code shapes.

Can I see sizes in the browser? Chrome DevTools shows function names in profiles but not sizes; use twiggy offline.

Why does a tiny function have a huge retained size? It is the only path to a large subtree — usually formatting, parsing or an allocator.

Why does my function look huge after wasm-opt? Inlining merged library code into it. Analyse a less-inlined build to see the original contributors.

← Back to Wasm Optimization Flags & Size Reduction