Pinning a Compiler Tier for Benchmarks
This page answers one task: make a WebAssembly benchmark measure exactly one compiler tier — baseline or optimized — so that results are stable, comparable between builds, and attributable to code changes rather than to when tier-up happened.
Prerequisites
- [ ] Node 20+ (for V8 flags), Firefox (for SpiderMonkey preferences) and optionally wasmtime.
- [ ] A benchmark harness that runs the function many times, as in building a reproducible Wasm benchmark harness.
Why pin a tier at all
Under normal tiering, a function’s speed changes during a benchmark: it starts on baseline code and switches to optimized code once the engine decides it is hot and finishes compiling it in the background. The moment of the switch depends on call counts, thread scheduling and machine load, so two runs of the same benchmark can disagree simply because the switch happened at different times. Comparing two builds is worse: if one build’s function tiers up earlier — because it is smaller, or called slightly differently — the comparison measures tiering, not code quality.
Pinning a tier removes that variable. Optimized-only runs measure the code quality your compiler and the engine’s optimizer produce together — the right number for comparing two builds of a hot kernel. Baseline-only runs measure what code runs like before tier-up — the right number for understanding startup and short-lived pages. Neither is what users see in general; users get a mixture. Pinning is a measurement technique, not a deployment setting.
Step 1 — pin tiers in V8 with Node
# optimized only: compile everything with TurboFan up front
node --no-liftoff bench.mjs
# baseline only: Liftoff, never tier up
node --liftoff --no-wasm-tier-up bench.mjs
# also disable lazy compilation so compile time is not in the first calls
node --no-liftoff --no-wasm-lazy-compilation bench.mjs
The first form makes every call run optimized code from the start, at the cost of a much longer initial compile — exclude compile time from the
measurement. The second keeps every function on Liftoff for the whole run. Flag names occasionally change; node --v8-options | grep -E 'liftoff|tier|lazy' lists what your version accepts. The same flags can be passed to Chromium with --js-flags="…" for browser-based benchmarks
in a dedicated profile.
Step 2 — pin tiers in Firefox
Firefox exposes the WebAssembly compilers as preferences in about:config:
javascript.options.wasm_baselinejit true/false baseline compiler
javascript.options.wasm_optimizingjit true/false Ion
Disable the baseline compiler to force optimized-only compilation; disable the optimizing compiler to stay on baseline code. Use a dedicated profile for benchmarking so you never browse with tiers disabled. Firefox’s tiering behaviour, which optimizes eagerly in the background, is described in how SpiderMonkey compiles Wasm.
Step 3 — understand wasmtime
wasmtime compiles every function ahead of time with Cranelift at a chosen optimization level, so there is no tier-up within a run. The equivalent of pinning is choosing the optimization level explicitly:
wasmtime run -O opt-level=2 app.wasm # default: optimized
wasmtime run -O opt-level=0 app.wasm # minimal optimization
Server-side benchmarks therefore vary less from run to run, which makes wasmtime a convenient place to compare builds of compute kernels — with the caveat that Cranelift’s code quality differs from browser engines’, so ratios between builds transfer better than absolute numbers.
Step 4 — exclude compilation and keep the rest of the method
Pinning changes which code runs; it does not replace good benchmarking practice. Separate compile time from execution time — with --no-liftoff
the initial compile is long and should not be averaged into per-call numbers. Still warm up, because CPU caches, branch predictors and first-touch
memory effects are independent of the compiler tier. Still take the median of many iterations and interleave variants. And always pin the same way for
every variant you compare: an optimized-only run of build A against a default-tiering run of build B measures nothing useful.
// bench.mjs — compile separately, warm up, measure, report median
const t0 = performance.now();
const module = await WebAssembly.compile(bytes);
const compileMs = performance.now() - t0;
const { exports } = await WebAssembly.instantiate(module, imports);
for (let i = 0; i < 20; i++) exports.kernel(ptr, n); // warm caches and memory
const samples = [];
for (let i = 0; i < 200; i++) { const t = performance.now(); exports.kernel(ptr, n); samples.push(performance.now() - t); }
samples.sort((a, b) => a - b);
console.log({ compileMs: compileMs.toFixed(1), medianMs: samples[100].toFixed(3) });
Step 5 — report what was pinned
Make the configuration part of every result: engine and version, tier pinning flags, compile time, warm-up, and the statistic reported. “3.1 ms” is
meaningless next to “3.1 ms median, V8 12.4 --no-liftoff, after 20 warm-up calls”. When the purpose is to decide between two builds, report the
optimized-only comparison, and add one default-tiering end-to-end number to confirm the difference survives in realistic conditions.
Building a small pinned benchmark suite
Pinning pays off most when it is part of a routine, not a one-off experiment. A practical setup is a script that runs each benchmark three times — optimized-only, baseline-only and default tiering — for the current build and the base build, and prints a compact table. The optimized-only column answers “did the code get faster?”, the baseline-only column answers “did the pre-tier-up experience change?”, and the default column checks that the first two add up to what users get. Running the three configurations in Node takes seconds for most kernels and needs no browser.
Keep the inputs fixed and committed, so changes between runs come from code rather than data. Store the raw medians per configuration as CI artifacts, which makes it possible to look back at when a regression first appeared in each tier. And keep the suite small: three or four kernels that represent the module’s hot paths are more useful than fifty microbenchmarks nobody reads. The CI side of this is covered in tracking benchmark results in CI, where pinned measurements are a straightforward way to lower the noise floor.
When pinned results mislead
Pinned results can steer you wrong in two specific ways. A change that improves optimized-only speed may make startup worse — a larger module from more inlining, or a function so big the optimizer takes longer — and a pinned benchmark will never show it. And a short-lived workload may never reach optimized code in real use, so optimized-only improvements to it are invisible to users. Pair every pinned comparison with a check of what users experience: default tiering, cold start, on a representative device. The tier-up observation techniques in watching Wasm tier-up in Chrome tell you which tier a real session actually spends its time in, which decides which pinned number matters.
Expected output
For the same kernel and two builds:
build A build B
--no-liftoff (median) 1.31 ms 1.12 ms ← B's code is 15% faster
--liftoff --no-tier-up 4.05 ms 4.20 ms ← B is slightly slower before tier-up
default tiering, 10 s 1.33 ms 1.14 ms ← the optimized gain survives end to end
Gotchas
- Including compile time in per-call averages. Optimized-only compilation is slow; measure it separately.
- Pinning only one variant. Both builds must run under identical flags.
- Browsing with tiers disabled. Use a separate browser profile for benchmarks.
- Comparing across machines. Pinned numbers still depend on the CPU. Compare builds on one machine.
- Treating pinned numbers as user-facing. Users experience tiering. Confirm with a default-tiering run.
Performance note
Pinning reduced run-to-run variation of a kernel benchmark in Node from about 9% to under 2%, because every run measured the same code. That made a 4% improvement between two builds clearly visible, where under default tiering it had been lost in the noise of when tier-up occurred.
Frequently Asked Questions
Can a web page pin tiers? No. Tier selection is an engine setting, available through command-line flags, browser preferences and runtime options, not page APIs.
Is optimized-only the same as “after warm-up” in default tiering? Close, but not identical. In default tiering some functions may still be on baseline code after warm-up, and inlining decisions can differ.
Should CI benchmarks pin tiers? For comparing builds, yes — it makes regressions easier to see. Keep at least one default-tiering benchmark so you notice startup regressions too.
What about Safari? JavaScriptCore’s tiers are not exposed as user preferences. Benchmark Safari with long warm-ups and report steady state, and use the pinned results from other engines to compare builds.
Does pinning affect SIMD or threads? No. Features are independent of tiers; both compilers support the same instructions.
Related
- Avoiding JIT warm-up errors in Wasm benchmarks — benchmarking under default tiering.
- Tracking benchmark results in CI — where pinned benchmarks reduce noise.
- Comparing Wasm runtimes on the same workload — server-side comparisons.
- Measuring the speed cost of -Oz — a build comparison that benefits from pinning.
← Back to Engine Tiering & JIT Compilation