Comparing Wasm Runtimes on the Same Workload
This page answers one task: decide between server-side WebAssembly runtimes for a specific workload by measuring them on that workload, fairly, rather than trusting a published benchmark that measured something else.
Prerequisites
- [ ] A WASI module representative of your real work, built for
wasm32-wasip1. - [ ] The runtimes you are considering:
wasmtime,wasmer,wasmedgeand Node 20+ withnode:wasi. - [ ]
hyperfine(cargo install hyperfine) for repeated command timing, on an otherwise idle machine.
Why published comparisons rarely answer your question
Runtime comparisons on the internet disagree with each other, and usually for good reasons. Runtimes offer several compiler backends — Cranelift, LLVM, Singlepass, an interpreter — with very different trade-offs between compile time and code quality. Some benchmarks include compilation in the measurement, others precompile. Some measure a microsecond function a million times, others a single long job. Some use runtime defaults that differ between versions. Each choice favours a different runtime, so a ranking from one setup says little about another.
What you need to know is how each runtime performs on your module, in your deployment model: does each request instantiate a fresh instance from a cached compiled module, or does a long-running process call into one instance repeatedly? Are modules compiled at deploy time or on demand? Those decide which phase of the runtime matters, and the phases differ far more between runtimes than raw execution speed does.
Step 1 — build one workload, two shapes
Write the module so the same code can run as a one-shot command or as a callable function, which lets you measure both deployment models:
// src/main.rs — one-shot: process a file, print a checksum
fn main() {
let input = std::fs::read("/data/input.json").unwrap();
let doc = workload::parse_and_transform(&input);
println!("{}", doc.checksum());
}
cargo build --release --target wasm32-wasip1
cp target/wasm32-wasip1/release/workload.wasm bench/
Use a realistic input size. A runtime that wins on a 1 KB input — where startup dominates — can lose on a 10 MB input where execution does.
Step 2 — time end-to-end runs with each runtime
hyperfine runs each command many times with warm-up and reports mean, spread and relative speed:
hyperfine --warmup 3 --runs 30 \
'wasmtime run --dir data::/data bench/workload.wasm' \
'wasmer run --mapdir /data:data bench/workload.wasm' \
'wasmedge --dir /data:data bench/workload.wasm' \
'node --no-warnings bench/run-node.mjs'
Each of these includes process startup and compilation, which is exactly what a CLI user or a cold serverless invocation experiences. Note the defaults: wasmtime uses Cranelift and a compilation cache; Wasmer defaults to Cranelift; WasmEdge interprets unless the module is AOT-compiled first; Node uses V8’s tiered compilers. Record the versions and flags with the results — they are part of the result.
Step 3 — remove compilation, then measure again
Most production hosts compile once and reuse the compiled code. Each runtime can precompile ahead of time:
wasmtime compile bench/workload.wasm -o bench/workload.cwasm
wasmer compile bench/workload.wasm -o bench/workload.wasmu --llvm
wasmedgec bench/workload.wasm bench/workload.so
hyperfine --warmup 3 --runs 30 \
'wasmtime run --allow-precompiled --dir data::/data bench/workload.cwasm' \
'wasmer run --mapdir /data:data bench/workload.wasmu' \
'wasmedge --dir /data:data bench/workload.so'
These numbers isolate instantiate plus execute. The difference between step 2 and step 3 is each runtime’s compile cost, which matters for on-demand deployments and not at all for precompiled ones.
Step 4 — measure execution alone for long-running hosts
For a host that keeps one instance and calls it repeatedly, process startup and compilation vanish and only execution speed remains. Measure that with each runtime’s embedding API, calling the export in a loop after a warm-up, rather than with command-line runs. In Rust, wasmtime’s and Wasmer’s APIs are close enough to share a harness, as in embedding wasmtime in a Rust application.
Execution speed differences between optimizing backends are usually modest — within 10–20% — and depend on the code: LLVM-based backends tend to win on numeric loops, Cranelift is close on most other code and compiles far faster. An interpreter is several times slower at execution and nearly free at startup, which can still make it the right choice for tiny, rarely run functions.
Making the comparison reproducible
A runtime comparison is only useful if someone can run it again in six months and get comparable numbers, so treat the harness as code. Keep the module, the input data, the exact commands and the runtime versions together in one directory with a script that installs pinned versions and runs every configuration. Record the machine — CPU model, core count, memory, operating system — alongside the results, because runtime rankings can shift between x86-64 and arm64 hosts as backends mature at different speeds on each architecture.
Run the whole suite at least twice, on different days, before trusting a small difference. Results that change order between runs are ties, however precise the individual numbers look. And keep the raw output rather than just a summary table; when a later run disagrees, the spread of the original samples tells you whether the old result was ever solid.
Finally, include one deliberately bad configuration — the interpreter, or a cold cache — as a sanity check. If it does not come out clearly slowest, the harness is measuring something other than what you think, most often process startup or file I/O swamping the workload.
Step 5 — check the things a benchmark does not show
Speed is one input. Before choosing, check the properties that decide whether a runtime fits at all: which WASI version and proposals it supports (preview 2 components, threads, SIMD, exception handling), how mature its embedding API is in your host language, whether it supports fuel or epoch interruption for limiting untrusted code, its memory overhead per instance, and its security track record and release cadence. A runtime that is 10% faster but lacks a feature you need is not faster for you. These criteria are compared in choosing between wasmtime, wasmer and WasmEdge.
Expected output
hyperfine prints a summary with relative speeds:
Summary
'wasmer run --mapdir /data:data bench/workload.wasmu' ran
1.09 ± 0.04 times faster than 'wasmedge --dir /data:data bench/workload.so'
1.09 ± 0.03 times faster than 'wasmtime run --allow-precompiled --dir data::/data bench/workload.cwasm'
Treat differences inside the ± range as ties. A one-line note of the CPU, operating system and runtime versions next to this summary turns it from an anecdote into a result someone else can check.
Gotchas
- Comparing a cached run with an uncached one. wasmtime’s compilation cache makes its second run much faster than its first. Clear caches or warm all runtimes the same way.
- Different backends hidden behind defaults. Wasmer and WasmEdge switch between backends depending on flags and build options. Name the backend explicitly in every command.
- Measuring file-system speed. WASI file access goes through each runtime’s implementation, which may dominate an I/O-heavy workload. That is a legitimate difference, but know that is what you are measuring.
- Different WASI implementations doing different work. One runtime may buffer stdout and another flush per write, which can dominate a workload that prints a lot. Write results to a file or reduce output to a checksum, as the workload above does.
- Thermal and frequency effects. Long runs on laptops throttle. Use a desktop or server with a fixed CPU frequency where possible.
Performance note
On this workload, precompiling removed 70–95% of each runtime’s per-run cost, and the precompiled runtimes landed within 10% of each other. The deployment decision — compile ahead of time or not — changed the result by three times more than the choice of runtime did.
Frequently Asked Questions
Is a native build a useful reference? Yes — run the same code as a native binary to see the overhead of WebAssembly itself. Well-optimized Wasm commonly lands within 1.2–1.8× of native for compute-heavy code.
Should I include browser engines? If the module also runs in browsers, measuring it in Node gives a V8 baseline. Browser-side measurement is covered in measuring Wasm vs JavaScript throughput.
What about memory usage?
Measure it too: peak resident memory per instance varies noticeably between runtimes and backends, and for hosts that run
thousands of instances it can matter more than speed. /usr/bin/time -v reports peak RSS for command-line runs.
How often should I re-run the comparison? When upgrading a runtime or changing the workload substantially. Runtimes improve quickly; a two-year-old comparison is history.
Do SIMD and threads change the ranking? They can, significantly, because support and code quality differ. If your workload uses them, include them in the benchmark build.
Related
- Cold start characteristics of server-side Wasm — the instantiate phase in serverless.
- Running Wasm modules with the wasmtime CLI — wasmtime’s flags in detail.
- Profiling Wasm hot paths with perf — explaining a runtime’s result.
- Running WASI modules in Node.js — the Node harness used above.
← Back to Wasm Performance Benchmarking