Shipping a Ported Module Behind a Feature Flag

This page answers one task: a WebAssembly port of a performance-critical JavaScript function is ready and passes its tests, and you want to release it to real users without betting everything on it — able to compare it against the old code in production and switch back instantly.

Prerequisites

  • [ ] The old JavaScript implementation and the Wasm port behind the same function signature.
  • [ ] A feature-flag system (a vendor service, a config endpoint, or a simple percentage rollout keyed on a user id).
  • [ ] Telemetry for errors and timings, as in measuring Wasm performance with real-user monitoring.

Why ports need a careful rollout

A port changes two things at once: the code and the platform it runs on. Tests cover the inputs you have; production supplies the ones you do not — unusual encodings, enormous files, devices with little memory, browsers with missing features. A WebAssembly module also adds failure modes the JavaScript did not have: the binary can fail to load, compile or instantiate. A feature-flagged rollout turns those risks into measurements. Both implementations stay in the bundle (or the Wasm one loads lazily), a flag decides which one serves each user, and telemetry shows how each behaves — speed, errors and, during a shadow phase, whether their results agree.

Rollout stages for a ported module The rollout starts with shadow mode, where the old code serves results and the Wasm port runs alongside for comparison. Then the port serves a small percentage, then half, then everyone, with the old code kept as a fallback until it is finally removed. 0 day shadow: compare only 7 day 1% served by Wasm 14 day 10% 21 day 50% 28 day 100%, JS as fallback 42 day remove old code

Step 1 — put both implementations behind one interface

Callers should not know which implementation answered. Wrap them in a module that exposes the original API and chooses internally:

// parse.js — the only module callers import
import { parse as parseJs } from "./legacy/parse.js";
let wasmParse = null;                              // loaded lazily

async function loadWasm() {
  const m = await import("./wasm/parse.js");
  await m.default();
  return m.parse;
}

export async function parse(text, { flags = currentFlags() } = {}) {
  if (flags.wasmParse) {
    try {
      wasmParse ??= await loadWasm();
      return wasmParse(text);
    } catch (err) {
      reportFallback(err);                         // load or runtime failure: fall back, but record it
    }
  }
  return parseJs(text);
}

If the original API is synchronous, preload the Wasm module at startup when the flag is on and use it once ready, falling back to JavaScript until then; changing a synchronous API to an asynchronous one is a larger change than the port itself.

Step 2 — start in shadow mode

Before any user sees Wasm results, run both implementations for a sample of calls, return the JavaScript result, and compare:

export async function parseShadowed(text) {
  const result = parseJs(text);
  if (Math.random() < SHADOW_RATE) {
    queueMicrotask(async () => {
      try {
        const w = (wasmParse ??= await loadWasm())(text);
        if (!deepEqual(result, w)) reportMismatch({ size: text.length, kind: classify(result, w) });
        recordTiming("wasm-shadow", /* … */);
      } catch (err) { reportShadowError(err); }
    });
  }
  return result;
}

Report mismatches with the input’s shape — length, record count, presence of non-ASCII text — not its content. Run the shadow work off the critical path, ideally in a worker, so comparison never slows users down. The comparison logic is the same one built in keeping JavaScript and Wasm results identical.

Step 3 — ramp up by percentage

When shadow mode shows zero unexplained mismatches over enough traffic, start serving Wasm results to a small percentage of users. Bucket by a stable user or session id, so each user sees consistent behaviour:

function inRollout(userId, percent) {
  let h = 2166136261;
  for (const c of userId) h = Math.imul(h ^ c.charCodeAt(0), 16777619);
  return ((h >>> 0) % 100) < percent;
}

Watch three numbers per cohort: the error rate (including load failures and fallbacks), the operation’s latency percentiles, and user-facing outcomes such as task completion. Increase the percentage in steps — 1, 10, 50, 100 — holding each long enough to cover weekly traffic patterns.

Choosing an implementation for one call The flag decides whether a user is in the Wasm cohort. Users in the cohort load the module lazily; if loading or the call fails, the call falls back to JavaScript and the failure is reported. Users outside the cohort use JavaScript, with a sampled shadow comparison. call parse(text) same API flag: Wasm cohort? stable bucket per user load + call Wasm lazy, cached failure → JS fallback reported result returned user unaware

Step 4 — make the fallback automatic and visible

Some users will be unable to run the module: old browsers, strict Content Security Policies, extensions that interfere, devices that cannot allocate the memory. The wrapper must fall back to JavaScript for them without user-visible errors, and the telemetry must record why, so you can see whether the fallback rate is 0.1% (normal) or 8% (something is wrong with deployment). A kill switch — the same flag set to 0% — must take effect without a deploy, so a bad release can be switched off in minutes. Test the kill switch before you need it.

Step 5 — remove the old code

Two implementations are a maintenance cost: every bug fix must be made twice and every behaviour change kept in sync. Once the port has served everyone for a few weeks with a negligible fallback rate, decide what the fallback is for. If no supported browser needs it, delete the JavaScript implementation and the flag. If a small population still needs it — older devices, hosts that forbid WebAssembly — keep it as a documented fallback with its own tests, and stop adding features to it. Leaving the flag in place indefinitely “just in case” is how codebases accumulate dead paths nobody dares remove.

What to put in the dashboard

A rollout dashboard needs only a few panels, compared side by side for the JavaScript and Wasm cohorts: operation latency at the median and 95th percentile, error and fallback rates broken down by cause (load, compile, runtime trap, mismatch), the share of users in each cohort, and one user-facing outcome metric such as successful imports or completed searches. Segment latency by device class, because the Wasm benefit is usually largest on slow devices and that is easy to miss in an aggregate. Add the shadow-mode mismatch count while it runs. With these panels, each ramp-up decision becomes a quick look rather than a meeting, and a regression shows up as a divergence between two lines rather than a vague sense that something got worse.

Bundle size and loading strategy

Shipping both implementations costs bytes. Keep the JavaScript version in the main bundle (it is usually small) and load the Wasm module lazily, only for users in the cohort and only when the feature is used. That way users outside the rollout pay nothing, and users inside it pay the download only once, cached thereafter. Preloading the module after the page becomes idle, for users in the cohort, hides most of the download latency. When the rollout completes and the JavaScript version becomes a fallback, consider moving it out of the main bundle too, loaded only when the Wasm path fails. The techniques are covered in lazy loading Wasm on first use.

Server-side rollouts

The same pattern works when the port runs on a server — a Node service replacing a JavaScript hot path with a Wasm module. There, shadow mode is even easier: both implementations run on the same machine, inputs can be compared without privacy concerns about leaving the device, and latency is measured directly. Roll out per request or per tenant instead of per user, and watch CPU usage per instance as well as latency, since the main benefit on servers is often fewer machines for the same traffic.

Expected output

Shadow mode runs for a week with 0 unexplained mismatches across 2.3 million sampled calls; the rollout reaches 100% over four weeks; the 95th-percentile latency of the operation drops by 61% in the Wasm cohort; the fallback rate settles at 0.2%, all from browsers without the required features; and the kill switch, tested once, disables the port within two minutes.

Gotchas

  • Turning a synchronous API asynchronous. Callers break. Preload and keep the API shape.
  • Random bucketing per call. Users flip between implementations. Bucket on a stable id.
  • Silent fallbacks. Users are fine but you never learn why. Report every fallback with a cause.
  • Untested kill switch. Test it before the rollout.
  • Keeping both paths forever. Set a removal date when the rollout starts.

Performance note

Lazy loading kept the main bundle unchanged for users outside the cohort. Shadow comparisons at a 2% sampling rate, run in a worker, added no measurable main-thread time. In the Wasm cohort, p95 latency for imports fell from 3.4 s to 1.3 s on mid-range phones.

Import latency at the 95th percentile by cohort and device Seconds to import a large file at the 95th percentile for the JavaScript and Wasm cohorts on desktop and mid-range phones during the rollout. seconds (p95) JS cohort, desktop 0.9 s Wasm cohort, desktop 0.4 s JS cohort, mid-range phone 3.4 s Wasm cohort, mid-range phone 1.3 s

Frequently Asked Questions

Can I use an existing feature-flag service? Yes — any service that buckets users consistently and changes values without deploys works.

How much shadow traffic is enough? Enough to cover your input variety: typically days of traffic and hundreds of thousands of calls for a widely used function.

Should shadow comparison run on users’ devices? Only if it is cheap and off the main thread. Alternatively, compare on the server with recorded input shapes.

What if mismatches appear only on one browser? Investigate engine-specific behaviour — float functions, locale data — and decide per case.

How do I keep the two implementations in sync during rollout? Freeze behaviour changes to the function until the rollout completes, or apply each change to both with shared tests.

Does the pattern work for server-side ports? Yes — bucket per request or tenant, compare in-process, and watch CPU per instance as well as latency.

← Back to Porting JavaScript Hot Paths to Wasm