Alerting on Wasm Load Failures
This page answers one task: a release goes out and a WebAssembly-powered feature silently stops working for some users — a CDN serves the wrong content type, a new build uses a feature an older Safari lacks, a glue file and module are out of sync. You want to detect such failures within minutes, from real users, with enough detail to know what broke.
Prerequisites
- [ ] Control over the code that loads the module (your loader or a wrapper around generated glue).
- [ ] An error-reporting or metrics pipeline (Sentry, a logging endpoint, RUM tooling).
- [ ] Release identifiers available in the client (version and module hash).
Why load failures deserve their own signal
Most error monitoring watches exceptions thrown by application code. Wasm load failures are different: they happen before any application logic runs, often for a whole class of users at once, and they are usually caused by deployment rather than code — headers, caches, CDNs, mismatched files, browser support. They also tend to be hidden by fallbacks: a feature that degrades gracefully when Wasm fails produces no visible error, just a worse experience. A dedicated “Wasm load success rate” per release and browser makes these failures visible as a number that should be near 100% and should not move when you deploy.
Step 1 — classify failures in the loader
Wrap loading so every failure is caught, classified and reported once:
async function loadModule(url, imports) {
const t0 = performance.now();
let stage = "fetch";
try {
const res = await fetch(url);
if (!res.ok) throw Object.assign(new Error(`HTTP ${res.status}`), { cls: "network" });
stage = "compile";
const { instance } = await WebAssembly.instantiateStreaming(res, imports);
report("wasm_load", { ok: true, ms: performance.now() - t0 });
return instance;
} catch (e) {
const cls = e.cls ?? classify(e, stage);
report("wasm_load", { ok: false, cls, stage, name: e.name, msg: String(e.message).slice(0, 200) });
throw e;
}
}
function classify(e, stage) {
if (e instanceof WebAssembly.CompileError) return /Content Security Policy|wasm-unsafe-eval/i.test(e.message) ? "csp" : "compile";
if (e instanceof WebAssembly.LinkError) return "link";
if (e instanceof RangeError) return "memory";
if (e instanceof TypeError && /MIME|Content-Type|application\/wasm/i.test(e.message)) return "mime";
if (stage === "fetch") return "network";
return "other";
}
Error messages differ between browsers, so classification by error type plus a few message patterns is approximate; include the raw (truncated) message for anything classified as “other” so new patterns can be added.
Step 2 — attach release context
Every report should carry the app version, the module’s content hash, the URL fetched (without query parameters that might contain user data), the browser family and major version, and whether a fallback ran. Content-hash tagging matters: when glue and module mismatch, the report shows which glue version tried to load which module.
Step 3 — compute a success rate and alert on change
Count successes as well as failures — a failure count alone cannot distinguish “more users” from “more failures”. In your metrics system, compute the load success rate per release and per browser family over short windows (5–15 minutes). Alert when the rate for the newest release drops below a threshold (for example 99%) or below the previous release’s rate by more than a margin, with a minimum sample size so small numbers do not trigger false alarms.
Step 4 — map failure classes to fixes
The dominant failure class usually identifies the cause:
- MIME: the server or CDN sends
application/octet-streamortext/plain— fix headers; see fixing incorrect response MIME type errors. - Compile with “magic word” messages: the response body was HTML (an error page or SPA fallback) or double-compressed.
- Compile in specific browsers only: the build uses a feature those versions lack; add feature detection and a fallback build.
- CSP: the page’s policy lacks
'wasm-unsafe-eval'. - Link: cached glue and new module (or the reverse) — fingerprint both files together and deploy them atomically.
- Memory: initial memory too large for some devices; reduce it or allow growth.
Step 5 — test the alert
Alerts that have never fired may not work. In a staging environment, deliberately break loading — serve the module with a wrong MIME type, deploy mismatched glue — and confirm reports arrive, the success rate drops, and the alert fires with the right class. Re-run this when the loader or the reporting pipeline changes.
Fallbacks and the signal
If the application falls back (to JavaScript, a server API, or a reduced feature) when Wasm fails, the user may not notice — but you should. Report fallback activations as part of the same signal, and track “share of sessions using the fallback” per release. A sudden increase is as much a regression as an error spike, even though nothing visibly broke.
Ad blockers, extensions and corporate networks
Some failures are environmental and not fixable by you: extensions blocking requests, corporate proxies stripping or altering responses, very old browsers. They appear as a steady background failure rate. Measure that baseline over a few releases, set thresholds relative to it, and look at changes rather than absolute numbers. Serving modules from your own origin (not a third-party CDN) reduces blocking by privacy extensions.
Load time as part of release health
Failures are not the only regression a release can introduce; a module that loads successfully but takes twice as long hurts users too. Report load duration with every success — total time and, where the loader separates them, fetch and compile phases — and track percentiles per release alongside the success rate. A jump in the 75th percentile after a deploy usually means the module grew, compression stopped working, or caching broke because file names or headers changed. Combine the two signals in one release-health view: success rate, p75 load time, and fallback share, each compared with the previous release over the same window. Rollout systems that ramp releases gradually (to 1%, 10%, 50% of users) can gate each step on these numbers, turning load health into an automatic brake rather than a dashboard someone has to remember to look at.
Correlating with server-side signals
Client reports tell you that loads fail; server and CDN logs often tell you why. Log requests for .wasm and glue files with status codes, response sizes and
content types, and look at them alongside client failure classes. A spike in 404s for the new module hash points to an incomplete upload; responses with
text/html content types for .wasm paths point to a routing fallback serving the app shell; a mix of old and new file versions served by different edge
locations points to a cache invalidation problem. Keeping client and server views next to each other shortens the time from alert to root cause considerably.
Expected output
Every load reports success or a classified failure with release, module hash and browser; the dashboard shows a 99.7% success rate baseline; a deploy that pushed new glue without the matching module triggers a “link” alert within eight minutes for the new release only; a staging drill confirms MIME and CSP failures alert correctly; and fallback usage is tracked alongside.
Gotchas
- Reporting only failures. Rates cannot be computed. Count successes too.
- No release context. You cannot tell which deploy broke. Attach version and module hash.
- Thresholds without minimum samples. Noise triggers alerts. Require enough events.
- Silent fallbacks. Regressions hide. Report fallback activations.
- Never testing alerts. They fail when needed. Run drills in staging.
- Watching failures but not load time. Slow loads are regressions too. Track percentiles per release.
Performance note
Reporting adds one small beacon per page load; with sampling of successes at 10% (failures always reported), the volume was about 0.1 events per session.
Frequently Asked Questions
Should success events be sampled? Yes, if volume matters — scale counts back up when computing rates; always send failures.
Can I detect failures without changing the loader? Partially, via global error handlers, but classification and context are much weaker.
What about server-side or edge Wasm? Track instantiation failures per deployment in the host’s metrics; the same classes apply.
How fast should alerts fire? Within the rollout window, so a bad release can be halted before reaching everyone.
Should load time be part of release health? Yes — track p75 load time per release next to the success rate; a jump usually means size, compression or caching changed.
Can rollouts stop automatically on load regressions? Yes — gate each rollout step on the new release’s success rate, p75 load time and fallback share compared with the previous release.
Where do I look first when an alert fires? At the dominant failure class, then CDN and server logs for the module and glue URLs of the new release.
Related
- Reporting Wasm crashes to an error tracker — runtime errors.
- Fixing Wasm streaming compile failed errors — common causes.
- Fixing CSP errors that block Wasm compilation — CSP failures.
- Versioning Wasm files with content hashes — avoiding mismatches.
← Back to Observability & Error Reporting