Testing Wasm in Multiple Browsers with Playwright

This page answers one task: a WebAssembly feature works in the browser you develop in, but users report failures in Safari or Firefox — so you want a single end-to-end suite that loads the real build in Chromium, Firefox and WebKit, exercises the Wasm paths, and fails when any engine misbehaves.

Prerequisites

  • [ ] A production build of the app with its .wasm files.
  • [ ] Node and @playwright/test, with browsers installed via npx playwright install --with-deps.
  • [ ] A static server you can configure with headers (Playwright’s webServer option can start it).

Why engine coverage matters for Wasm

WebAssembly is specified precisely, and the core instruction set behaves identically across engines. Differences appear around it: which post-MVP features each engine supports and in which version (SIMD, threads, exception handling, GC, memory64, JSPI), how strictly each enforces MIME types and Content Security Policy for compilation, memory limits on mobile devices, how SharedArrayBuffer and cross-origin isolation behave, and timing-sensitive behaviour such as streaming compilation and tier-up. A feature that uses a proposal shipped in Chrome but not yet in Safari, or a loader that relies on a Chromium-only API, fails only in other engines.

Playwright drives Chromium, Firefox and WebKit through one API, so the same test runs against all three. WebKit in Playwright is close to Safari’s engine but not identical to shipping Safari, and Chromium is not identical to Chrome; treat the matrix as strong evidence rather than proof, and keep a manual or device-cloud check on real Safari for releases.

What each engine run catches Chromium runs catch general regressions. Firefox runs catch differences in feature support and stricter or different header handling. WebKit runs catch missing post-MVP features, memory limits and Safari-specific loader issues closest to iOS users. engine typical Wasm failures caught Chromium regressions, Chromium-only APIs relied on Firefox feature gaps, header and CSP differences WebKit missing proposals, memory limits, Safari loader quirks

Step 1 — configure projects for three engines

// playwright.config.ts
import { defineConfig, devices } from "@playwright/test";

export default defineConfig({
  testDir: "e2e",
  use: { baseURL: "http://localhost:4173", trace: "retain-on-failure" },
  webServer: { command: "node e2e/serve.mjs", port: 4173, reuseExistingServer: !process.env.CI },
  projects: [
    { name: "chromium", use: { ...devices["Desktop Chrome"] } },
    { name: "firefox",  use: { ...devices["Desktop Firefox"] } },
    { name: "webkit",   use: { ...devices["Desktop Safari"] } },
  ],
});

Every test now runs three times, once per project. trace: "retain-on-failure" records a trace with network, console and DOM snapshots for failed tests.

Step 2 — serve the build like production

The test server must set Content-Type: application/wasm, the isolation headers if the app uses threads, and any CSP the production site uses — otherwise the tests exercise a different configuration from users:

// e2e/serve.mjs
import http from "node:http";
import { readFile } from "node:fs/promises";
import path from "node:path";

const types = { ".html": "text/html", ".js": "text/javascript", ".wasm": "application/wasm" };
http.createServer(async (req, res) => {
  const p = path.join("dist", new URL(req.url, "http://x").pathname.replace(/\/$/, "/index.html"));
  try {
    const body = await readFile(p);
    res.writeHead(200, {
      "Content-Type": types[path.extname(p)] ?? "application/octet-stream",
      "Cross-Origin-Opener-Policy": "same-origin",
      "Cross-Origin-Embedder-Policy": "require-corp",
    });
    res.end(body);
  } catch { res.writeHead(404).end(); }
}).listen(4173);

Step 3 — wait for readiness and fail on console errors

Have the app signal when the module is ready — a data-wasm-ready attribute or a performance.mark — and wait for that, rather than for an arbitrary timeout. Collect console errors and page errors in a fixture, so a trap or a failed instantiation fails the test even if the UI looks fine:

import { test as base, expect } from "@playwright/test";

const test = base.extend({
  page: async ({ page }, use) => {
    const errors: string[] = [];
    page.on("pageerror", (e) => errors.push(e.message));
    page.on("console", (m) => m.type() === "error" && errors.push(m.text()));
    await use(page);
    expect(errors, "console or page errors").toEqual([]);
  },
});

test("resizes an image with the Wasm codec", async ({ page }) => {
  await page.goto("/editor");
  await page.locator("[data-wasm-ready]").waitFor();
  await page.setInputFiles("input[type=file]", "e2e/fixtures/photo.jpg");
  await page.getByRole("button", { name: "Resize" }).click();
  await expect(page.getByTestId("output-size")).toHaveText("800 × 600");
});
One test across three engines Playwright starts the configured server, then runs the same test in Chromium, Firefox and WebKit. Each run loads the page, waits for the module's ready signal, performs the user action, checks the result and fails if any console error or uncaught exception occurred. start server MIME + COOP/COEP launch engine chromium, firefox, webkit wait for ready data-wasm-ready act + assert real user flow no console errors trap = failure

Step 4 — handle features some engines lack

When the app deliberately falls back on engines without a feature — no threads, no JSPI — test both paths and assert which one ran, instead of skipping engines:

test("uses threads where available", async ({ page, browserName }) => {
  await page.goto("/editor");
  await page.locator("[data-wasm-ready]").waitFor();
  const mode = await page.locator("[data-wasm-ready]").getAttribute("data-wasm-mode");
  const isolated = await page.evaluate(() => crossOriginIsolated);
  expect(isolated).toBe(true);
  expect(["threads", "single"]).toContain(mode);
  test.info().annotations.push({ type: "wasm-mode", description: `${browserName}: ${mode}` });
});

Use test.skip(browserName === "webkit", "reason") only for genuinely unsupported scenarios, with a reason that names the feature, so skips are visible and revisited when engines catch up. Feature detection in the app is covered in feature-detecting Wasm at startup.

Step 5 — run the matrix in CI

jobs:
  e2e:
    strategy:
      fail-fast: false
      matrix: { project: [chromium, firefox, webkit] }
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: npm ci && npm run build
      - run: npx playwright install --with-deps ${{ matrix.project }}
      - run: npx playwright test --project=${{ matrix.project }}
      - uses: actions/upload-artifact@v4
        if: failure()
        with: { name: "trace-${{ matrix.project }}", path: test-results }

fail-fast: false lets all engines finish, so one failure shows whether the problem is engine-specific. Upload traces for failed runs; Playwright’s trace viewer shows network responses (including Wasm headers) and console output at the moment of failure.

Reading engine-specific failures

When a test fails in only one engine, the failure message usually points at the category. CompileError mentioning an unknown opcode or section means the module uses a feature the engine does not support — check build flags and feature detection. TypeError: Incorrect response MIME type in one engine but not others means the server path differs or one engine is stricter. A RangeError on WebAssembly.Memory points to memory limits — WebKit’s are lower in some configurations. Errors referencing SharedArrayBuffer or Atomics mean isolation headers did not apply. Timeouts waiting for readiness in only one engine often hide an exception thrown before the ready signal; the console-error fixture usually reveals it.

Keeping the suite fast and stable

Three engines triple the run time, so keep cross-engine tests focused on what differs: loading, feature paths and one or two representative flows per Wasm-backed feature. Run broad UI tests on one engine. Avoid timing assertions in cross-engine tests — compilation speed differs widely between engines and CI machines — and test performance separately with budgets. Shard the suite across machines with --shard when it grows, and keep fixture files small so uploads and processing do not dominate the run.

Testing mobile viewports and constrained devices

Desktop engines on CI machines have generous memory and fast CPUs; phones do not. Playwright’s device descriptors (devices["iPhone 15"], devices["Pixel 7"]) emulate viewport, user agent and touch, but not memory limits or CPU speed. For Wasm features that allocate large memories, add a test that runs the heaviest realistic input and asserts the module’s peak memory.buffer.byteLength stays under a budget suited to mobile — a few hundred megabytes at most — because exceeding it is what crashes tabs on iOS. CPU throttling through the Chrome DevTools Protocol (Emulation.setCPUThrottlingRate) works only in Chromium, so use it for responsiveness checks there and accept that the other engines run at full speed. Real-device checks through a device cloud complement this for releases.

Covering service workers and caching

Many Wasm apps cache modules in a service worker or rely on the HTTP cache for repeat visits. Playwright can test both: load the page, wait for the service worker to activate (page.waitForFunction(() => navigator.serviceWorker.controller)), reload, and assert through page.on("response") that the .wasm response came from the service worker (response.fromServiceWorker()). Run this in every engine, because service-worker behaviour around Response streaming and instantiateStreaming has differed between them, and a cached response missing its Content-Type fails streaming compilation in some engines but not others.

Expected output

The suite runs nine tests per engine on every pull request; Chromium and Firefox report data-wasm-mode="threads", WebKit reports the mode its isolation support allows; no console errors are recorded; and a deliberate MIME-type misconfiguration fails all three projects with traces showing the response headers.

Gotchas

  • A test server without production headers. Tests pass on a different configuration. Mirror MIME types, isolation and CSP.
  • Fixed timeouts instead of readiness signals. Flaky on slower engines. Wait for an explicit signal.
  • Ignoring console errors. Traps go unnoticed. Fail tests on page and console errors.
  • Skipping engines without reasons. Gaps become permanent. Name the missing feature.
  • Treating Playwright WebKit as Safari. Close, not identical. Check real Safari before releases.

Performance note

The cross-engine job runs in 3 min 40 s with the three projects in parallel jobs, compared with 9 min 50 s when run serially in one job; the WebKit project is consistently the slowest due to longer browser startup on Linux runners.

Duration of the cross-engine suite in CI Minutes for a nine-test Playwright suite run against Chromium, Firefox and WebKit, serially in one job compared with three parallel jobs. minutes per CI run serial in one job 9.8 min three parallel jobs 3.7 min

Frequently Asked Questions

Can Playwright test Safari on iOS? Not directly; its WebKit build approximates Safari. Use a device cloud or a manual check for iOS-specific issues.

Should every test run in all engines? No — run Wasm-related flows everywhere and the rest in one engine.

How do I test a WebKit-only failure locally? npx playwright test --project=webkit --debug opens the inspector on Linux, macOS and Windows.

Do workers and threads work under Playwright? Yes, with the same headers as production; check crossOriginIsolated in a test.

Can Playwright limit memory to simulate a phone? Not directly; assert the module’s peak memory stays under a mobile budget instead, and confirm on real devices before releases.

← Back to Testing & Verifying Wasm Builds