Profiling Wasm Hot Paths with perf
This page answers one task: find which functions in a WebAssembly module consume the time when it runs under wasmtime on Linux,
using perf — the same profiler you would use for native code — with Wasm function names intact.
Prerequisites
- [ ] Linux with
perfinstalled (linux-tools-$(uname -r)or your distribution’s equivalent). - [ ]
wasmtimeCLI 20+ or an embedding built with profiling support. - [ ] A module with its
namesection intact — do not strip it for profiling builds. - [ ] Permission to profile:
kernel.perf_event_paranoidset to 1 or lower, or root.
Why a native profiler needs help with Wasm
perf samples the CPU’s instruction pointer many times per second and maps each sample to a function using the symbol tables of
loaded binaries. That works for native programs because their code lives in files with symbols. WebAssembly code does not: the
runtime compiles it at load time into anonymous executable memory, so perf sees samples landing in addresses no file describes,
and reports them as [unknown] or as a raw hex address. The profile is useless.
The fix is for the runtime to tell perf what it compiled and where. wasmtime supports two mechanisms. perfmap writes a simple
text file, /tmp/perf-<pid>.map, listing address ranges and function names — enough for function-level attribution. jitdump
writes a richer binary file containing the compiled code itself, which lets perf annotate individual instructions and survives
code being moved or recompiled. Both rely on the module’s name section to know what to call each function.
Step 1 — build a profiling-friendly module
Profile an optimized build — profiling a debug build finds the wrong hot spots — but keep names:
[profile.profiling]
inherits = "release"
debug = 1 # line tables, for annotation
strip = "none"
cargo build --profile profiling --target wasm32-wasip1
wasm-objdump -h target/wasm32-wasip1/profiling/app.wasm | grep -E '"name"|debug_line'
If you run wasm-opt, add --debuginfo so it preserves the name section. Release builds that strip names produce profiles full
of wasm-function[1832], which you can map back by hand but should not have to.
Step 2 — record with perfmap for a quick function profile
perf record -g -F 999 -k mono \
wasmtime run --profile=perfmap --dir data::/data target/wasm32-wasip1/profiling/app.wasm
perf report --sort symbol --no-children | head -25
# Overhead Symbol
31.42% [.] app::transform::apply_rules
18.07% [.] app::parser::Lexer::next_token
11.85% [.] core::str::from_utf8
7.22% [.] dlmalloc::dlmalloc::Dlmalloc::malloc
5.90% [.] hashbrown::raw::RawTable::find
4.11% [.] wasmtime_runtime::libcalls::memory32_grow
That is a function-level profile with names, which answers the main question in most investigations: where does the time go? Here a third of it is in rule application, a fifth in the lexer, and a surprising 12% in UTF-8 validation — a hint that strings are being validated more often than necessary.
Step 3 — use jitdump for instruction-level annotation
When a function is hot and you need to know which part of it, record with jitdump and inject the JIT code into the profile:
perf record -k mono -g \
wasmtime run --profile=jitdump --dir data::/data target/wasm32-wasip1/profiling/app.wasm
perf inject --jit --input perf.data --output perf.jit.data
perf annotate --input perf.jit.data --stdio -s 'app::transform::apply_rules' | head -40
The -k mono flag selects the monotonic clock that jitdump records use, so samples line up with the code records. The annotation
shows machine instructions with the percentage of samples on each, interleaved with source lines when the module has line tables.
A loop where most samples sit on a single load instruction is waiting on memory; samples spread across arithmetic indicate
compute. That distinction connects directly to
benchmarking memory bandwidth in Wasm.
Step 4 — read hardware counters
perf stat counts hardware events for the whole run, which helps classify a bottleneck before you dig into functions:
perf stat -e cycles,instructions,cache-misses,branch-misses \
wasmtime run --dir data::/data target/wasm32-wasip1/profiling/app.wasm
3,912,480,117 cycles
7,105,330,980 instructions # 1.82 insn per cycle
41,227,811 cache-misses
28,915,402 branch-misses
An instructions-per-cycle figure below about one usually means the CPU is stalled on memory; above two, the code is executing efficiently and further speed must come from doing less work. Compare these numbers between two builds or two runtimes to see whether a change altered the amount of work or how efficiently it runs.
Step 5 — render a flame graph
A flame graph shows call stacks, which perf report presents only as nested text. Convert the recorded data:
perf script --input perf.data > out.perf
stackcollapse-perf.pl out.perf | grep -v '^wasmtime_cranelift' > out.folded
flamegraph.pl --title "app.wasm under wasmtime" out.folded > flame.svg
Filtering out compiler frames focuses the graph on execution; leave them in if compilation time is what you are investigating. The same technique for browser profiles is in generating flame graphs for Wasm.
Interpreting a server-side profile
Profiles of embedded WebAssembly mix three kinds of time, and it is worth separating them before optimizing. Time in your module’s
functions is the work you control directly. Time in runtime libcalls — memory32_grow, table operations, trap handling — is
the module asking the runtime for something; a large share usually means a design issue such as growing memory in small steps or
excessive indirect calls. Time in host functions — WASI implementations, your own imports — is the boundary; a large share means
the module calls out too often or with too much data per call. Each kind has a different fix, and the profile shows which one
you have. In the example above, 4% in memory32_grow was cut to almost nothing by reserving memory up front, before any code in
the module was touched.
Making profiling routine
The barrier to profiling server-side WebAssembly is mostly setup, so remove it once. Keep the profiling cargo profile in the
repository, add a script that builds with it and runs the workload under perf record with the right flags, and document the
one-line sysctl change developers need. With that in place, profiling a regression flagged by CI is a single command rather
than an afternoon of rediscovering flags.
It also pays to profile the production-like configuration, not just a convenient one. The same module under a different wasmtime version, a different compiler backend, or with fuel metering enabled can have a noticeably different profile, because those settings change the code the runtime generates. If production enables fuel or epoch interruption for safety, profile with it enabled; its cost appears in the profile as extra instructions in loop headers, and that is part of the real picture.
Finally, keep old profiles. A folded-stack file from a known-good release is a cheap artifact to store, and comparing it with a new one — a differential flame graph — is the quickest way to see which functions grew when a release got slower.
Expected output
A flame graph SVG in which the widest towers under app::transform::apply_rules and app::parser::Lexer::next_token show which
callers lead to the hot spots, with Wasm function names rather than addresses.
Gotchas
- Everything shows as
[unknown]. wasmtime was run without--profile, orperfcould not read the map file. Check that/tmp/perf-<pid>.mapwas written and that you are profiling the same process. - Names are
wasm-function[N]. The module was stripped. Rebuild with the name section kept. perf injectfinds no JIT code.-k monowas missing fromperf record, so timestamps do not match the jitdump records.- Permission denied. Lower
perf_event_paranoidfor the session (sudo sysctl kernel.perf_event_paranoid=1) or profile as root.
Performance note
Profiling with perfmap at 999 Hz slowed the workload by about 2%, small enough to profile production-like runs. jitdump recording added about 6%, mostly from writing code records at compile time. In the investigation above, the profile led to two fixes — caching validated strings and reserving memory up front — that together cut run time by 19%.
Frequently Asked Questions
Does this work with Wasmer or WasmEdge?
Both have profiling integrations of varying maturity; wasmtime’s perfmap and jitdump support is the most established for perf.
Can I profile an embedded wasmtime host the same way?
Yes. Enable profiling in the engine configuration with Config::profiler(ProfilingStrategy::PerfMap) or JitDump, then run the host
under perf record.
How long should a recording be? Ten to thirty seconds of the workload is usually plenty at 999 Hz: tens of thousands of samples resolve any function that takes more than a fraction of a percent of the time.
What about macOS?
perf is Linux-only. wasmtime’s --profile=guest mode writes a profile you can open in the Firefox Profiler on any platform, which
is a good cross-platform alternative.
Can I profile inside Node instead?
Node can write a perf map for V8-compiled code with --perf-basic-prof, which includes Wasm functions; the results look similar.
Related
- Comparing Wasm runtimes on the same workload — explaining a runtime’s numbers.
- Profiling allocation hot spots in a Wasm module — when the profile points at malloc.
- Embedding wasmtime in a Rust application — enabling profiling in an embedding.
- Reading the name custom section — where the function names come from.
← Back to Wasm Performance Benchmarking