Pinning a Compiler Tier for Benchmarks

This page answers one task: make a WebAssembly benchmark measure exactly one compiler tier — baseline or optimized — so that results are stable, comparable between builds, and attributable to code changes rather than to when tier-up happened.

Prerequisites

Why pin a tier at all

Under normal tiering, a function’s speed changes during a benchmark: it starts on baseline code and switches to optimized code once the engine decides it is hot and finishes compiling it in the background. The moment of the switch depends on call counts, thread scheduling and machine load, so two runs of the same benchmark can disagree simply because the switch happened at different times. Comparing two builds is worse: if one build’s function tiers up earlier — because it is smaller, or called slightly differently — the comparison measures tiering, not code quality.

Pinning a tier removes that variable. Optimized-only runs measure the code quality your compiler and the engine’s optimizer produce together — the right number for comparing two builds of a hot kernel. Baseline-only runs measure what code runs like before tier-up — the right number for understanding startup and short-lived pages. Neither is what users see in general; users get a mixture. Pinning is a measurement technique, not a deployment setting.

What each pinned measurement tells you Optimized-only runs isolate code quality and are best for comparing builds of hot code. Baseline-only runs show pre-tier-up performance relevant to startup. Default tiering is what users experience but mixes both. optimized only every call on top-tier code stable, low variance compares compiler output fairly comparing builds of hot code baseline only every call on baseline code shows pre-tier-up speed relevant to short-lived work startup and one-shot tasks default tiering mix that shifts during the run higher variance what users actually get end-to-end reality checks

Step 1 — pin tiers in V8 with Node

# optimized only: compile everything with TurboFan up front
node --no-liftoff bench.mjs

# baseline only: Liftoff, never tier up
node --liftoff --no-wasm-tier-up bench.mjs

# also disable lazy compilation so compile time is not in the first calls
node --no-liftoff --no-wasm-lazy-compilation bench.mjs

The first form makes every call run optimized code from the start, at the cost of a much longer initial compile — exclude compile time from the measurement. The second keeps every function on Liftoff for the whole run. Flag names occasionally change; node --v8-options | grep -E 'liftoff|tier|lazy' lists what your version accepts. The same flags can be passed to Chromium with --js-flags="…" for browser-based benchmarks in a dedicated profile.

Step 2 — pin tiers in Firefox

Firefox exposes the WebAssembly compilers as preferences in about:config:

javascript.options.wasm_baselinejit     true/false   baseline compiler
javascript.options.wasm_optimizingjit   true/false   Ion

Disable the baseline compiler to force optimized-only compilation; disable the optimizing compiler to stay on baseline code. Use a dedicated profile for benchmarking so you never browse with tiers disabled. Firefox’s tiering behaviour, which optimizes eagerly in the background, is described in how SpiderMonkey compiles Wasm.

Step 3 — understand wasmtime

wasmtime compiles every function ahead of time with Cranelift at a chosen optimization level, so there is no tier-up within a run. The equivalent of pinning is choosing the optimization level explicitly:

wasmtime run -O opt-level=2 app.wasm        # default: optimized
wasmtime run -O opt-level=0 app.wasm        # minimal optimization

Server-side benchmarks therefore vary less from run to run, which makes wasmtime a convenient place to compare builds of compute kernels — with the caveat that Cranelift’s code quality differs from browser engines’, so ratios between builds transfer better than absolute numbers.

Tier pinning options by engine How to force optimized-only or baseline-only compilation in V8, SpiderMonkey and wasmtime. engine optimized only baseline only V8 (Node, Chromium) --no-liftoff --liftoff --no-wasm-tier-up SpiderMonkey (Firefox) wasm_baselinejit = false wasm_optimizingjit = false wasmtime (Cranelift) -O opt-level=2 (default) -O opt-level=0

Step 4 — exclude compilation and keep the rest of the method

Pinning changes which code runs; it does not replace good benchmarking practice. Separate compile time from execution time — with --no-liftoff the initial compile is long and should not be averaged into per-call numbers. Still warm up, because CPU caches, branch predictors and first-touch memory effects are independent of the compiler tier. Still take the median of many iterations and interleave variants. And always pin the same way for every variant you compare: an optimized-only run of build A against a default-tiering run of build B measures nothing useful.

// bench.mjs — compile separately, warm up, measure, report median
const t0 = performance.now();
const module = await WebAssembly.compile(bytes);
const compileMs = performance.now() - t0;
const { exports } = await WebAssembly.instantiate(module, imports);
for (let i = 0; i < 20; i++) exports.kernel(ptr, n);            // warm caches and memory
const samples = [];
for (let i = 0; i < 200; i++) { const t = performance.now(); exports.kernel(ptr, n); samples.push(performance.now() - t); }
samples.sort((a, b) => a - b);
console.log({ compileMs: compileMs.toFixed(1), medianMs: samples[100].toFixed(3) });

Step 5 — report what was pinned

Make the configuration part of every result: engine and version, tier pinning flags, compile time, warm-up, and the statistic reported. “3.1 ms” is meaningless next to “3.1 ms median, V8 12.4 --no-liftoff, after 20 warm-up calls”. When the purpose is to decide between two builds, report the optimized-only comparison, and add one default-tiering end-to-end number to confirm the difference survives in realistic conditions.

Building a small pinned benchmark suite

Pinning pays off most when it is part of a routine, not a one-off experiment. A practical setup is a script that runs each benchmark three times — optimized-only, baseline-only and default tiering — for the current build and the base build, and prints a compact table. The optimized-only column answers “did the code get faster?”, the baseline-only column answers “did the pre-tier-up experience change?”, and the default column checks that the first two add up to what users get. Running the three configurations in Node takes seconds for most kernels and needs no browser.

Keep the inputs fixed and committed, so changes between runs come from code rather than data. Store the raw medians per configuration as CI artifacts, which makes it possible to look back at when a regression first appeared in each tier. And keep the suite small: three or four kernels that represent the module’s hot paths are more useful than fifty microbenchmarks nobody reads. The CI side of this is covered in tracking benchmark results in CI, where pinned measurements are a straightforward way to lower the noise floor.

When pinned results mislead

Pinned results can steer you wrong in two specific ways. A change that improves optimized-only speed may make startup worse — a larger module from more inlining, or a function so big the optimizer takes longer — and a pinned benchmark will never show it. And a short-lived workload may never reach optimized code in real use, so optimized-only improvements to it are invisible to users. Pair every pinned comparison with a check of what users experience: default tiering, cold start, on a representative device. The tier-up observation techniques in watching Wasm tier-up in Chrome tell you which tier a real session actually spends its time in, which decides which pinned number matters.

Expected output

For the same kernel and two builds:

                          build A        build B
--no-liftoff  (median)    1.31 ms        1.12 ms     ← B's code is 15% faster
--liftoff --no-tier-up    4.05 ms        4.20 ms     ← B is slightly slower before tier-up
default tiering, 10 s     1.33 ms        1.14 ms     ← the optimized gain survives end to end

Gotchas

  • Including compile time in per-call averages. Optimized-only compilation is slow; measure it separately.
  • Pinning only one variant. Both builds must run under identical flags.
  • Browsing with tiers disabled. Use a separate browser profile for benchmarks.
  • Comparing across machines. Pinned numbers still depend on the CPU. Compare builds on one machine.
  • Treating pinned numbers as user-facing. Users experience tiering. Confirm with a default-tiering run.

Performance note

Pinning reduced run-to-run variation of a kernel benchmark in Node from about 9% to under 2%, because every run measured the same code. That made a 4% improvement between two builds clearly visible, where under default tiering it had been lost in the noise of when tier-up occurred.

Run-to-run variation with and without tier pinning Coefficient of variation across twenty runs of the same kernel benchmark in Node under default tiering and with --no-liftoff, showing that pinning removes most of the variation caused by tier-up timing. run-to-run variation (%) default tiering 9.1 % --no-liftoff (optimized only) 1.8 % --liftoff --no-wasm-tier-up 1.5 %

Frequently Asked Questions

Can a web page pin tiers? No. Tier selection is an engine setting, available through command-line flags, browser preferences and runtime options, not page APIs.

Is optimized-only the same as “after warm-up” in default tiering? Close, but not identical. In default tiering some functions may still be on baseline code after warm-up, and inlining decisions can differ.

Should CI benchmarks pin tiers? For comparing builds, yes — it makes regressions easier to see. Keep at least one default-tiering benchmark so you notice startup regressions too.

What about Safari? JavaScriptCore’s tiers are not exposed as user preferences. Benchmark Safari with long warm-ups and report steady state, and use the pinned results from other engines to compare builds.

Does pinning affect SIMD or threads? No. Features are independent of tiers; both compilers support the same instructions.

← Back to Engine Tiering & JIT Compilation