Reading wasm-ld Map Files

This page answers one task: a WebAssembly binary is larger than expected, contains a function nobody thought was linked, or resolved a symbol from the wrong library, and you want the linker’s own record of what went into the output and why.

Prerequisites

  • [ ] A build that links with wasm-ld — directly, through clang/emcc, or through rustc.
  • [ ] Access to the link flags (Cargo rustflags, Emscripten LDFLAGS, or the linker command).
  • [ ] Optionally twiggy or wasm-objdump for cross-checking sizes.

What a map file records

A link map is a text report the linker writes alongside its output. For every section of the final module — functions, data segments, globals — it lists each piece, the input it came from (an object file, or a member of a static archive such as libstd.rlib or libc.a), its size, and its address or offset in the output. It answers questions that the final binary alone cannot: which of the hundred archives contributed a function, why a 40 KB data segment exists, whether the code you thought was dead was actually kept.

wasm-ld writes a map with --Map=<file> (also spelled -Map). It reflects the result of linking — after dead-code elimination by section garbage collection — but before any post-link optimisation by wasm-opt. So it describes what the linker produced, which is usually the right level for questions about dependencies and symbol resolution; for final-size questions, combine it with analysis of the optimised binary.

Excerpt of a wasm-ld map file The map lists output sections with their offsets and sizes, then the input sections inside them with the object file or archive member each came from, down to individual symbols. VA Off Size Out In Symbol header - 142 1b2f4 CODE code section: 110 KB - 146 2a8e libapp.rlib(app.o):(.text.parse) your code - 2bd4 9a12 libstd.rlib(std.o):(.text.fmt) std formatting - c5e6 1f40 libc.a(qsort.o):(.text.qsort) libc member 10000 1b44a 3c10 DATA data section: 15 KB

Step 1 — generate the map

Pass the flag through whichever driver you use:

# clang / wasi-sdk
clang --target=wasm32-wasip1 src/*.c -O2 -Wl,--Map=build/app.map -o build/app.wasm

# Emscripten
emcc src/*.c -O2 -Wl,--Map=build/app.map -o build/app.js

# Rust (wasm32-unknown-unknown) — in .cargo/config.toml
# [target.wasm32-unknown-unknown]
# rustflags = ["-C", "link-arg=--Map=target/app.map"]
cargo build --release --target wasm32-unknown-unknown

Rust’s release builds with LTO combine most code into one object before linking, so the map attributes much of it to a single LTO object; build without LTO (or with -C codegen-units=16 and lto = false) temporarily when you want per-crate attribution.

Step 2 — find the big contributors

The map is plain text, sorted by output address. Sorting input sections by size gives a ranked list of contributors:

awk '$4 ~ /^[0-9a-f]+$/ && $6 != "" { printf "%8d  %s\n", strtonum("0x"$3), $5 }' build/app.map \
  | sort -rn | head -20

Group by archive to see which library is responsible for how much: everything from libstd.rlib(...) versus your own crate versus libc.a. A typical surprise is formatting code (core::fmt) pulled in by a single panic! with a message or a {:?} in an error path, discussed in removing panic and formatting bloat from Rust Wasm.

Step 3 — explain why something was kept

When an unexpected function appears in the map, the question is what references it. wasm-ld can tell you directly:

clang ... -Wl,--why-extract=build/why.txt        # which symbol caused each archive member to be extracted
clang ... -Wl,--trace-symbol=qsort               # every input that defines or references qsort

--why-extract writes, for each archive member pulled into the link, the symbol whose reference caused it — for example libc.a(printf.o) extracted because log_error references fprintf. Following that chain back usually leads to one call site that, once removed or replaced, frees a whole cluster of library code.

Tracing why a library function is in the binary The map shows a large function from an archive member. The why-extract report names the symbol that pulled that member in. Trace-symbol shows which of your objects referenced it. Removing or replacing that reference lets section garbage collection drop the whole chain. map: big libc member printf.o, 9 KB --why-extract pulled in by fprintf --trace-symbol referenced from log.o replace the call simpler logging relink member gone

Step 4 — check symbol resolution

When the same symbol is defined in more than one input — a custom malloc and libc’s, two versions of a vendored library — the map shows which definition was used. The linker takes the first definition it encounters in object files, and extracts archive members only for undefined symbols, so link order decides the winner. If the map shows libc’s malloc when you expected your allocator, your object was not on the command line or came after the archive. Related errors are covered in fixing undefined symbol errors from wasm-ld.

Step 5 — compare maps between builds

Keep the map as a build artefact in CI and diff it between commits when size changes. A diff of the sorted contributor list shows exactly what was added: a new dependency, a monomorphised generic, a newly reachable library function. Combined with the size budgets in catching size regressions in CI, this turns “the binary grew by 30 KB” into “this commit started using std::collections::BTreeMap in the parser”.

Map files versus post-link analysis

The map describes the linker’s output, not the final shipped binary. wasm-opt then inlines, merges and removes code, often shrinking the binary by a quarter or more, so absolute sizes in the map overstate the final cost of each function. Use the map for attribution — which input, which archive, why it was linked — and tools that read the final binary for final sizes: twiggy top and twiggy dominators on the optimised module, as in analyzing Wasm size with twiggy, or wasm-objdump -h for section sizes. Names connect the two: keep the name section in the binary you analyse (strip it only from the shipped copy) so function names in twiggy’s output match those in the map.

Data segments in the map

Code is only half the story. The map also lists data segments with their sources: string literals, lookup tables, Unicode tables from the standard library, static arrays. Large read-only tables are common contributors in C and Rust programs — a 64 KB CRC table, Unicode case-mapping data from char::to_lowercase, embedded fonts or test fixtures compiled in by accident with include_bytes!. The data section’s breakdown shows them immediately. Alternatives include computing tables at startup instead of embedding them, loading large assets separately at runtime, or using smaller algorithms that do not need full Unicode tables when inputs are known to be ASCII.

Mapping Rust crates to their cost

In Rust projects, the question is usually “which crate costs how much”. Object and archive names in the map follow crate names — libserde_json-…rlib, libregex_syntax-…rlib — so grouping input sections by the archive name before the first hyphen gives a per-crate total. A short script that sums sizes per crate and prints a sorted table makes dependency decisions concrete: whether regex is worth 180 KB for a validation feature, whether chrono can be replaced by a few lines of date arithmetic, whether a serialisation format crate is pulling in its whole feature set. Generic code complicates attribution, because monomorphised copies are emitted into the crate that instantiates them; a large generic instantiated from your crate appears under your crate’s name even though the code came from a dependency. The function names in the map, which include the generic parameters, reveal that.

Expected output

The map shows that 38 KB of the code section comes from core::fmt via an error path’s format!; --why-extract shows libc.a(printf.o) pulled in by a debug fprintf; removing both shrinks the linked output by 52 KB and the optimised binary by 31 KB; and the map is stored as a CI artefact for future diffs.

Gotchas

  • Reading sizes from an LTO build. Everything is attributed to one LTO object. Build without LTO for attribution.
  • Treating map sizes as final. wasm-opt changes them. Use twiggy on the optimised binary for final sizes.
  • Forgetting link order. The first definition wins. Check which definition the map shows.
  • Ignoring data segments. Tables and strings can dominate. Read the data section too.
  • Losing the map. Keep it as a CI artefact to diff later.
  • Monomorphised generics misattributed. They appear under the instantiating crate. Read the full symbol names.

Performance note

Generating the map added under a second to the link of a 2 MB module and produced a 1.4 MB text file. Two hours spent following its largest contributors cut the optimised binary from 612 KB to 448 KB.

Size before and after acting on the map Kilobytes of optimised Wasm before investigation, after removing formatting code from error paths, and after also replacing a libc printf dependency and an embedded table. KB of optimised .wasm before 612 KB after removing fmt from error paths 531 KB after printf + table changes 448 KB

Frequently Asked Questions

Does Emscripten produce a map automatically? No — pass -Wl,--Map=file explicitly. Emscripten’s --emit-symbol-map is different: it maps minified function names, not inputs.

Can I get a map for Rust builds with wasm-pack? Yes, through rustflags link arguments; wasm-pack passes them to cargo.

Is the map format stable? It resembles LLD’s ELF maps but is not a formal format; scripts should be tolerant of small changes.

What is --why-extract output format? A tab-separated list of archive member, referencing input and symbol; easy to grep.

How do I see cost per Rust crate? Group the map’s input sections by archive name and sum sizes; remember monomorphised generics appear under the instantiating crate.

Can the map show which exports keep code alive? Indirectly — exported functions are roots; combine the map with --why-extract and the export list to see what each export pulls in.

← Back to Linking Wasm Objects with wasm-ld