Using Engine Flags to Experiment with Tiering
This page answers one task: you want to understand how your WebAssembly module behaves under each of the engine’s compilation tiers — how fast the baseline code is, how much optimisation buys, how long compilation takes — and the normal browser hides that by switching tiers automatically. Engine flags let you pin and trace tiers in a test environment; you want to use them correctly and interpret the results.
Prerequisites
- [ ] A test Chrome or Chromium build you can launch from the command line, and/or Node.js.
- [ ] A benchmark harness for your module.
- [ ] Willingness to treat flags as experimental — they change between versions.
Why experiment with tiers
V8 runs WebAssembly through Liftoff (a fast single-pass baseline compiler) and TurboFan (an optimising compiler), deciding per function when to optimise. That adaptivity is good for users but makes measurement confusing: a benchmark may measure Liftoff code, TurboFan code, or a mix, depending on run length. Pinning a tier answers specific questions: How slow is my code before optimisation? (matters for first-interaction latency and short-lived workers.) How much faster does TurboFan make it? (tells you whether warm-up is worth engineering.) How long does optimising compilation take for my module? (matters for startup and battery.)
Flag names in the table are examples from recent V8 versions; list the flags your version supports with node --v8-options | grep wasm and check before relying on
any of them.
Step 1 — pass flags to Node
node --v8-options | grep -i wasm | less # what this version supports
node --liftoff-only bench.mjs # baseline tier only
node --no-liftoff bench.mjs # optimising tier for everything
node --no-wasm-lazy-compilation bench.mjs # eager compilation
Node is the easiest environment for tier experiments: same V8, scriptable, no browser noise. Results apply to Chrome’s V8 of the same version, though browser factors (other tabs, GPU process, power saving) are absent.
Step 2 — launch Chrome with flags
Start a separate Chrome instance with its own profile so flags do not affect your normal browsing:
google-chrome --user-data-dir=/tmp/chrome-tier-test \
--js-flags="--liftoff-only" \
http://localhost:8080/bench.html
Verify the flags took effect: some flags print a warning when unknown; differences in measured speed confirm the change. Never ask users to run with flags, and never depend on them in production — they are unsupported, may be removed, and can make browsers less secure or stable.
Step 3 — trace compilation
Tracing flags print when functions compile and in which tier (for example flags with trace and wasm in their names), which shows how many functions tier up
and when. Chrome’s Performance panel also shows compile tasks without flags. Use traces to confirm hypotheses — “the hot loop function tiers up after 200 ms”
— rather than to collect large logs.
Step 4 — interpret results
Typical findings: baseline code runs at roughly half to a third of optimised speed for compute-heavy kernels (varies widely), optimised compilation of a large module takes many times longer than baseline compilation, and adaptive tiering ends up close to optimised-only speed for long-running work while starting much faster. If baseline speed is far worse than optimised and your workload is short-lived (a single call after page load), invest in warm-up or reducing work; if the difference is small, tiering does not matter for you.
Step 5 — keep experiments reproducible
Record the browser or Node version, the exact flags, the machine, power state and the benchmark commit with every result. Flags that exist in one version may be renamed or change meaning in the next; a result without its version is not reproducible. Re-run key experiments when upgrading.
Other engines
Firefox exposes preferences in about:config for its Wasm compilers (for example toggling the baseline and optimising tiers, with names like
javascript.options.wasm_baselinejit and javascript.options.wasm_optimizingjit), and JavaScriptCore has environment variables and options for its tiers,
mostly intended for engine developers. They serve the same experimental purpose; consult each engine’s current documentation, since names change.
Experiments worth running
A few experiments answer most practical questions. Baseline versus optimised speed for your hottest functions tells you how much first-interaction latency and short-lived workers suffer before tier-up. Total compile time per tier for the whole module tells you what eager compilation would cost and how much work TurboFan does in the background — relevant on battery-powered devices. Time to tier-up for a hot loop under default settings tells you how long a benchmark must warm up and how long users wait for full speed. Lazy versus eager compilation for your startup path shows how much laziness saves at load time. Run each with the same inputs and harness, in fresh processes, and keep the scripts with the results so the experiment can be repeated after upgrades.
Using results to change code, not flags
The point of experiments is to inform changes you can ship. If baseline code is much slower and users interact immediately, warm up the critical path during idle time or reduce the work done on first interaction. If optimised compilation of a few huge functions dominates background work, split them. If tier-up takes long for a hot loop, consider whether the loop is too large or polymorphic in ways the optimiser handles poorly — measured with optimised-only runs. None of these require flags in production; the flags only told you where to look.
Comparing engines with the same questions
Asking the same questions of SpiderMonkey and JavaScriptCore, through their own options, often reveals that a function’s baseline speed or compile time differs a lot between engines. That explains browser-specific complaints and helps prioritise: a pause that is 20 ms in Chrome and 120 ms in Safari deserves a fix even if Chrome users never notice it.
Automating tier experiments in CI
Tier behaviour changes with engine releases, so a one-off experiment goes stale. A small CI job that runs the benchmark under default, baseline-only and optimised-only configurations in Node — the easiest V8 to script — and records results per Node version catches shifts early: a new V8 that makes your baseline code much slower, or a toolchain change that produces code TurboFan optimises poorly. Keep the job tolerant of flags disappearing (skip a configuration with a clear message if a flag is not recognised) so engine upgrades do not break the pipeline, and treat its output as a trend to watch rather than a pass/fail gate.
Expected output
A tier experiment on the image filter shows 42 ms per frame under default tiering after warm-up, 96 ms with Liftoff only, and 41 ms with TurboFan only, while compile time is 35 ms (Liftoff) versus 410 ms (TurboFan) for the whole module; the team concludes that a warm-up call during idle time is worthwhile and records the Node and Chrome versions used.
Gotchas
- Flags in production or user instructions. Unsupported and risky. Experiments only.
- Using your main browser profile. Flags affect everything. Use a separate profile.
- Not checking flag names for your version. They change. List supported flags first.
- Comparing runs across versions. Engines change. Record versions.
- Over-interpreting single runs. Use the same harness rules and repeat.
- Experiments that break when flags are renamed. Detect unsupported flags and skip with a message.
Performance note
For a compute-heavy filter, baseline-only code ran at about 44% of optimised speed; adaptive tiering reached optimised speed after roughly 300 ms of steady use.
Frequently Asked Questions
Can a page detect which tier is running? No — tiering is invisible to code; only timing hints at it.
Are flags safe for automated tests? For dedicated test environments, yes; keep them out of user-facing runs.
Do flags affect correctness? They should not, but experimental flags can expose engine bugs; treat odd results with suspicion.
Should benchmarks pin a tier? For tier analysis, yes; for user-facing performance, measure default behaviour after warm-up.
What should change after a tier experiment? Code and loading strategy — warm-up, smaller functions, less work on first interaction — never production flags.
Can tier experiments run in CI? Yes, most easily in Node; record results per engine version and skip configurations whose flags no longer exist.
Which engine should I experiment in first? V8 through Node, because it is scriptable; then check important findings in Firefox and Safari with their own options.
How long should a tier-up experiment run? Long enough for timings to plateau under default settings — often a few hundred milliseconds to seconds of steady work.
Related
- Understanding lazy compilation in V8 — what flags change.
- Pinning a compiler tier for benchmarks — benchmark setups.
- Avoiding JIT warm-up errors in Wasm benchmarks — warm-up rules.
- Comparing baseline and optimized Wasm code — looking at the code.
← Back to Engine Tiering & JIT Compilation