Profiling Wasm Hot Paths with perf

This page answers one task: find which functions in a WebAssembly module consume the time when it runs under wasmtime on Linux, using perf — the same profiler you would use for native code — with Wasm function names intact.

Prerequisites

  • [ ] Linux with perf installed (linux-tools-$(uname -r) or your distribution’s equivalent).
  • [ ] wasmtime CLI 20+ or an embedding built with profiling support.
  • [ ] A module with its name section intact — do not strip it for profiling builds.
  • [ ] Permission to profile: kernel.perf_event_paranoid set to 1 or lower, or root.

Why a native profiler needs help with Wasm

perf samples the CPU’s instruction pointer many times per second and maps each sample to a function using the symbol tables of loaded binaries. That works for native programs because their code lives in files with symbols. WebAssembly code does not: the runtime compiles it at load time into anonymous executable memory, so perf sees samples landing in addresses no file describes, and reports them as [unknown] or as a raw hex address. The profile is useless.

The fix is for the runtime to tell perf what it compiled and where. wasmtime supports two mechanisms. perfmap writes a simple text file, /tmp/perf-<pid>.map, listing address ranges and function names — enough for function-level attribution. jitdump writes a richer binary file containing the compiled code itself, which lets perf annotate individual instructions and survives code being moved or recompiled. Both rely on the module’s name section to know what to call each function.

How perf learns the names of JIT-compiled Wasm functions wasmtime compiles the module, then writes either a perf map or a jitdump file describing each compiled function's address and name, taken from the module's name section. perf record samples the process, and perf report or perf inject uses the file to attribute samples to Wasm function names. module + name section function names kept wasmtime compiles code in anonymous memory perfmap / jitdump address → name perf record samples instruction pointers perf report time per Wasm function

Step 1 — build a profiling-friendly module

Profile an optimized build — profiling a debug build finds the wrong hot spots — but keep names:

[profile.profiling]
inherits = "release"
debug = 1              # line tables, for annotation
strip = "none"
cargo build --profile profiling --target wasm32-wasip1
wasm-objdump -h target/wasm32-wasip1/profiling/app.wasm | grep -E '"name"|debug_line'

If you run wasm-opt, add --debuginfo so it preserves the name section. Release builds that strip names produce profiles full of wasm-function[1832], which you can map back by hand but should not have to.

Step 2 — record with perfmap for a quick function profile

perf record -g -F 999 -k mono \
  wasmtime run --profile=perfmap --dir data::/data target/wasm32-wasip1/profiling/app.wasm
perf report --sort symbol --no-children | head -25
# Overhead  Symbol
    31.42%  [.] app::transform::apply_rules
    18.07%  [.] app::parser::Lexer::next_token
    11.85%  [.] core::str::from_utf8
     7.22%  [.] dlmalloc::dlmalloc::Dlmalloc::malloc
     5.90%  [.] hashbrown::raw::RawTable::find
     4.11%  [.] wasmtime_runtime::libcalls::memory32_grow

That is a function-level profile with names, which answers the main question in most investigations: where does the time go? Here a third of it is in rule application, a fifth in the lexer, and a surprising 12% in UTF-8 validation — a hint that strings are being validated more often than necessary.

Step 3 — use jitdump for instruction-level annotation

When a function is hot and you need to know which part of it, record with jitdump and inject the JIT code into the profile:

perf record -k mono -g \
  wasmtime run --profile=jitdump --dir data::/data target/wasm32-wasip1/profiling/app.wasm
perf inject --jit --input perf.data --output perf.jit.data
perf annotate --input perf.jit.data --stdio -s 'app::transform::apply_rules' | head -40

The -k mono flag selects the monotonic clock that jitdump records use, so samples line up with the code records. The annotation shows machine instructions with the percentage of samples on each, interleaved with source lines when the module has line tables. A loop where most samples sit on a single load instruction is waiting on memory; samples spread across arithmetic indicate compute. That distinction connects directly to benchmarking memory bandwidth in Wasm.

Step 4 — read hardware counters

perf stat counts hardware events for the whole run, which helps classify a bottleneck before you dig into functions:

perf stat -e cycles,instructions,cache-misses,branch-misses \
  wasmtime run --dir data::/data target/wasm32-wasip1/profiling/app.wasm
     3,912,480,117      cycles
     7,105,330,980      instructions              #    1.82  insn per cycle
        41,227,811      cache-misses
        28,915,402      branch-misses

An instructions-per-cycle figure below about one usually means the CPU is stalled on memory; above two, the code is executing efficiently and further speed must come from doing less work. Compare these numbers between two builds or two runtimes to see whether a change altered the amount of work or how efficiently it runs.

What perf's outputs tell you about a Wasm workload The perf commands used on a Wasm module, what each one shows, and the question it answers. command shows answers perf report (perfmap) time per Wasm function where does time go? perf annotate (jitdump) time per instruction which part of the hot function? perf stat cycles, IPC, misses memory- or compute-bound? flame graph from perf script call stacks by width which callers lead there?

Step 5 — render a flame graph

A flame graph shows call stacks, which perf report presents only as nested text. Convert the recorded data:

perf script --input perf.data > out.perf
stackcollapse-perf.pl out.perf | grep -v '^wasmtime_cranelift' > out.folded
flamegraph.pl --title "app.wasm under wasmtime" out.folded > flame.svg

Filtering out compiler frames focuses the graph on execution; leave them in if compilation time is what you are investigating. The same technique for browser profiles is in generating flame graphs for Wasm.

Interpreting a server-side profile

Profiles of embedded WebAssembly mix three kinds of time, and it is worth separating them before optimizing. Time in your module’s functions is the work you control directly. Time in runtime libcalls — memory32_grow, table operations, trap handling — is the module asking the runtime for something; a large share usually means a design issue such as growing memory in small steps or excessive indirect calls. Time in host functions — WASI implementations, your own imports — is the boundary; a large share means the module calls out too often or with too much data per call. Each kind has a different fix, and the profile shows which one you have. In the example above, 4% in memory32_grow was cut to almost nothing by reserving memory up front, before any code in the module was touched.

Making profiling routine

The barrier to profiling server-side WebAssembly is mostly setup, so remove it once. Keep the profiling cargo profile in the repository, add a script that builds with it and runs the workload under perf record with the right flags, and document the one-line sysctl change developers need. With that in place, profiling a regression flagged by CI is a single command rather than an afternoon of rediscovering flags.

It also pays to profile the production-like configuration, not just a convenient one. The same module under a different wasmtime version, a different compiler backend, or with fuel metering enabled can have a noticeably different profile, because those settings change the code the runtime generates. If production enables fuel or epoch interruption for safety, profile with it enabled; its cost appears in the profile as extra instructions in loop headers, and that is part of the real picture.

Finally, keep old profiles. A folded-stack file from a known-good release is a cheap artifact to store, and comparing it with a new one — a differential flame graph — is the quickest way to see which functions grew when a release got slower.

Expected output

A flame graph SVG in which the widest towers under app::transform::apply_rules and app::parser::Lexer::next_token show which callers lead to the hot spots, with Wasm function names rather than addresses.

Gotchas

  • Everything shows as [unknown]. wasmtime was run without --profile, or perf could not read the map file. Check that /tmp/perf-<pid>.map was written and that you are profiling the same process.
  • Names are wasm-function[N]. The module was stripped. Rebuild with the name section kept.
  • perf inject finds no JIT code. -k mono was missing from perf record, so timestamps do not match the jitdump records.
  • Permission denied. Lower perf_event_paranoid for the session (sudo sysctl kernel.perf_event_paranoid=1) or profile as root.

Performance note

Profiling with perfmap at 999 Hz slowed the workload by about 2%, small enough to profile production-like runs. jitdump recording added about 6%, mostly from writing code records at compile time. In the investigation above, the profile led to two fixes — caching validated strings and reserving memory up front — that together cut run time by 19%.

Run time before and after fixes found with perf Wall-clock time for the workload under wasmtime before profiling, after avoiding repeated UTF-8 validation, and after reserving memory up front instead of growing it repeatedly. ms per run under wasmtime before profiling 1,310 ms + cached UTF-8 validation 1,162 ms + memory reserved up front 1,061 ms

Frequently Asked Questions

Does this work with Wasmer or WasmEdge? Both have profiling integrations of varying maturity; wasmtime’s perfmap and jitdump support is the most established for perf.

Can I profile an embedded wasmtime host the same way? Yes. Enable profiling in the engine configuration with Config::profiler(ProfilingStrategy::PerfMap) or JitDump, then run the host under perf record.

How long should a recording be? Ten to thirty seconds of the workload is usually plenty at 999 Hz: tens of thousands of samples resolve any function that takes more than a fraction of a percent of the time.

What about macOS? perf is Linux-only. wasmtime’s --profile=guest mode writes a profile you can open in the Firefox Profiler on any platform, which is a good cross-platform alternative.

Can I profile inside Node instead? Node can write a perf map for V8-compiled code with --perf-basic-prof, which includes Wasm functions; the results look similar.

← Back to Wasm Performance Benchmarking