How Wasm Code Caching Works in Browsers
This page answers one question: when a user returns to your site, does the browser compile your WebAssembly module again, or reuse compiled machine code from the last visit? The answer depends on how you load the module and how you serve it. You want to understand the mechanism well enough to make repeat visits start fast, and to check that it is actually happening.
Prerequisites
- [ ] A module loaded with
WebAssembly.instantiateStreamingorcompileStreaming(or glue that does). - [ ] Control over response headers and file naming.
- [ ] Chrome DevTools for verification.
What gets cached, and where
Two caches matter. The HTTP cache stores the .wasm bytes, saving the download on repeat visits if headers allow it. The code cache stores the engine’s
compiled machine code for those bytes, saving compilation. In Chrome, the Wasm code cache is attached to the HTTP cache entry for the module’s response: when a
module is compiled from a streamed response, V8 can produce serialised compiled code, and Chrome stores it alongside the cached response. On the next load of the
same URL with the same response, Chrome finds the cached code, checks it matches the bytes, and deserialises it instead of compiling.
Several conditions apply in Chrome: the module must be loaded through the streaming APIs (so the engine knows the response it came from), the response must be cacheable, the module must exceed a size threshold (small modules compile quickly enough that caching is not worth it), and the code is written to the cache only after the module has been used enough for optimised code to exist — so a module used for a moment and abandoned may not be cached on the first visit.
Step 1 — load with streaming APIs
const { instance } = await WebAssembly.instantiateStreaming(fetch("/app.3f9c2e.wasm"), imports);
WebAssembly.instantiate(arrayBuffer) gives the engine bytes without a response identity, so they cannot be associated with an HTTP cache entry; code caching
does not apply. Most toolchain glue uses streaming when the server sends Content-Type: application/wasm; if the MIME type is wrong, glue often falls back to
the ArrayBuffer path silently, losing both streaming compilation and code caching.
Step 2 — make the response cacheable and stable
Serve the module with a long-lived cache policy and a content-hashed URL:
Cache-Control: public, max-age=31536000, immutable
Content-Type: application/wasm
A URL that changes on every deployment (or carries a cache-busting query string per page load) prevents reuse. A URL that stays the same while the content changes forces revalidation and can invalidate cached code. Content-hashed file names give both: a stable URL for unchanged modules and a new URL when the bytes change. See versioning Wasm files with content hashes.
Step 3 — understand when code is written
Because the cache entry is written after the engine has produced optimised code for the module, a brief first visit may not populate the cache. Later visits that use the module enough will. For applications where the second visit’s startup matters a lot, you cannot force the write directly, but you can make sure the module is loaded early and used, and you can measure across several visits rather than expecting the second visit to be perfect.
Step 4 — service workers
Modules served from a service worker’s Cache Storage can also benefit from code caching in Chrome, provided they are loaded with streaming APIs and the response
comes from the service worker’s cache. A service worker that constructs new Response objects without the right headers, or that serves bytes as
application/octet-stream, defeats streaming and caching. Keep cached responses intact, with their content type.
Step 5 — verify cache hits
In Chrome DevTools, the Performance panel shows compile tasks on a cold load and their absence (or short deserialisation tasks) on a warm load. Compare total time to instantiate across visits: a large drop indicates code reuse. Clearing site data resets both caches, which is how to reproduce the cold path. For automated checks, measure instantiation time across repeated loads in a fresh browser profile with a script.
Other engines
Engines differ. Chrome’s mechanism is described above. Firefox and Safari have their own approaches to compilation and caching, which have changed over time; for example, Firefox relies on fast baseline compilation of whole modules, and caching behaviour across visits may differ from Chrome’s. Measure startup on repeat visits in each browser you care about rather than assuming Chrome’s behaviour is universal.
Deployments and cache churn
Every deployment that changes a module’s bytes starts its users from a cold code cache for that module, even if only one function changed. Teams that deploy many times a day can therefore keep most users permanently on the cold path. Two strategies reduce churn. Make builds reproducible, so a deployment that does not change Rust or C source produces byte-identical modules and therefore the same hashed URL — timestamps, absolute paths and non-deterministic ordering in build outputs silently defeat this. And split large applications into several modules by how often they change: a stable core library and a frequently updated feature layer, so routine releases invalidate only the smaller module. Measuring how often the module’s URL actually changes across a month of deployments is a quick way to see whether churn is costing you.
Field measurement
Lab checks show the mechanism works; field data shows how often users benefit. Record, per page load, whether the module came from the HTTP cache (Resource
Timing’s transferSize of 0 with a non-zero decodedBodySize indicates a cache hit) and how long compilation or instantiation took. The distribution of
instantiation times split by cache status shows the gap between cold and warm loads for real users and devices, and the share of warm loads shows how well
caching is working. A low share points to churn, headers or URL patterns; a small gap between cold and warm suggests code caching is not engaging, often
because of the ArrayBuffer path or a MIME problem.
Code caching for workers
Modules compiled inside workers also benefit when loaded with streaming APIs from cacheable URLs. A common pattern — compile once on the main thread and post the
WebAssembly.Module to workers — means only the main thread’s load interacts with the cache, and workers reuse the compiled module directly.
Expected output
The module is loaded with instantiateStreaming from a content-hashed URL with a one-year immutable cache policy; a cold load in Chrome shows 180 ms of
compilation; after a few visits, warm loads show compilation replaced by deserialisation taking about 15 ms; and a deployment that changes the module’s bytes
produces a new URL, so the old cache entry is never misapplied.
Gotchas
- Instantiating from ArrayBuffers. No response identity, no code cache. Stream.
- Wrong MIME type. Glue falls back silently. Serve
application/wasm. - Cache-busting query strings per load. Every visit is cold. Hash file names instead.
no-storeheaders. Nothing is cached. Allow caching for immutable files.- Expecting the second visit to be warm. Code may be written later. Measure several visits.
- Non-reproducible builds. Identical source yields new URLs and cold caches. Make builds deterministic.
Performance note
For a 6 MB module in Chrome on a laptop, a cold load spent about 180 ms compiling; with a warm code cache, deserialising cached code took about 15 ms.
Frequently Asked Questions
Can I store compiled modules in IndexedDB myself?
Storing WebAssembly.Module objects in IndexedDB was supported in some browsers historically but is not reliably available; rely on the code cache.
Does the code cache work in incognito? Caches are per session there and discarded afterwards.
Is the cache shared across origins? No — caches are partitioned by site.
Do small modules get cached? Below the size threshold, compilation is cheap enough that caching is skipped.
Why do my users rarely get warm loads? Frequent deployments change the module’s URL; make builds reproducible and split rarely changing code into its own module.
How can I tell from field data whether a module came from cache?
Resource Timing shows a transferSize of 0 with a non-zero decodedBodySize for HTTP cache hits.
Should workers load the module themselves? Compile once on the main thread with streaming and post the Module to workers, so all reuse one compiled module.
Does splitting a module help caching? Yes — a stable core module keeps its cached code across releases while a frequently changing feature module is recompiled.
Related
- Setting cache-control headers for Wasm — headers.
- Caching Wasm with a service worker — offline.
- Understanding lazy compilation in V8 — what is compiled.
- Measuring cache hit rates for Wasm assets — field data.
← Back to Engine Tiering & JIT Compilation