Caching Model Files for Offline Inference
This page answers one task: an application runs machine-learning models in the browser — speech recognition, embeddings, a small language model — and their weights are tens to hundreds of megabytes. Downloading them on every visit is slow and expensive for users. You want the files stored locally once, verified, versioned, available offline, and cleaned up when models change.
Prerequisites
- [ ] Model files served from your origin or a CDN, with stable versioned URLs.
- [ ] A runtime that accepts model bytes (ONNX Runtime Web, whisper.cpp, llama.cpp-style runtimes, TensorFlow.js).
- [ ] Optionally a service worker, for offline loading of the app itself.
Where large files can live
Browsers offer three storage options suitable for large binaries. The Cache API stores Response objects keyed by request URL; it is simple, works in
pages, workers and service workers, and fits “download a URL once” naturally. OPFS (Origin Private File System) stores files that can be read and written in
chunks, including with synchronous access handles in workers; it suits very large files, partial writes and resumable downloads. IndexedDB can store
Blobs and ArrayBuffers, but large binary values are less convenient there. All three count against the origin’s storage quota, which is typically a
large fraction of free disk space but varies by browser and can be evicted under pressure unless the origin has persistent storage.
The HTTP cache is not a reliable substitute: browsers may skip caching very large responses, evict them early, or partition them, and the application cannot inspect or control what is there.
Step 1 — version and hash model URLs
Put a version or content hash in every model URL — /models/whisper-base.en-q5_1.v3.bin or /models/<sha256>.bin — and never change the bytes behind a URL.
The URL then doubles as the cache key: a new model is a new URL, and stale caches cannot serve outdated weights. Publish a small manifest listing each model’s
URL, size and SHA-256, so the client knows what to expect before downloading.
Step 2 — download with progress and store in the Cache API
async function fetchModel({ url, size, sha256 }, onProgress) {
const cache = await caches.open("models-v1");
const hit = await cache.match(url);
if (hit) return new Uint8Array(await hit.arrayBuffer());
const res = await fetch(url);
const reader = res.body.getReader();
const chunks = []; let received = 0;
for (;;) {
const { done, value } = await reader.read();
if (done) break;
chunks.push(value); received += value.length;
onProgress(received / size);
}
const bytes = new Uint8Array(received);
let off = 0; for (const c of chunks) { bytes.set(c, off); off += c.length; }
const digest = new Uint8Array(await crypto.subtle.digest("SHA-256", bytes));
if (toHex(digest) !== sha256) throw new Error("model integrity check failed");
await cache.put(url, new Response(bytes, { headers: { "Content-Type": "application/octet-stream" } }));
return bytes;
}
Verifying the hash before caching prevents a truncated or corrupted download from being stored and reused forever. For files of several hundred megabytes, holding the whole file in memory twice (chunks plus the assembled array) is costly; stream directly into OPFS instead (Step 3).
Step 3 — use OPFS for very large or resumable downloads
Write chunks to an OPFS file as they arrive, and resume with an HTTP Range request after interruptions:
// in a worker
const dir = await navigator.storage.getDirectory();
const handle = await dir.getFileHandle("llm-q4.gguf.part", { create: true });
const access = await handle.createSyncAccessHandle();
let offset = access.getSize(); // resume from what we have
const res = await fetch(url, { headers: offset ? { Range: `bytes=${offset}-` } : {} });
const reader = res.body.getReader();
for (;;) {
const { done, value } = await reader.read();
if (done) break;
access.write(value, { at: offset }); offset += value.length;
postMessage({ progress: offset / size });
}
access.flush(); access.close();
// verify hash by streaming the file, then rename .part → final name
Hashing a large file incrementally needs an incremental hash (WebCrypto’s digest needs the whole buffer); a Wasm SHA-256 or BLAKE3 hasher fits, as in hashing large files in the browser with Wasm.
Step 4 — ask for persistence and check quota
Before downloading, check navigator.storage.estimate() and make sure the model fits with room to spare; tell users how much space will be used. Request
persistent storage (navigator.storage.persist()) after the user opts into downloading models, so the browser is less likely to evict them. Handle eviction
anyway: on each start, check the cache, and if the model is missing, explain and offer to download it again rather than failing with a runtime error.
Step 5 — clean up old versions
When the manifest lists a new model version, download it, switch over, and delete the old file — otherwise every model update adds hundreds of megabytes. Keep the list of expected files in the manifest and remove anything in the model cache or OPFS directory that is not on it. Provide a settings entry showing stored models with sizes and a button to remove them.
Offline loading of the whole app
Caching models makes inference offline-capable only if the app itself loads offline. Use a service worker to cache the application shell, JavaScript, Wasm runtime files and the manifest. Then a returning user can open the app without a network, load the model from OPFS or the Cache API, and run inference. Keep model files out of the service worker’s precache list — they are large and handled separately with progress and verification.
Serving model files well
Client-side caching starts with how files are served. Send model files with Cache-Control: public, max-age=31536000, immutable (safe because URLs are
versioned), support HTTP range requests so downloads can resume, and serve them from a CDN close to users. Compression rarely helps quantised weights much —
they are dense binary data — so measure before enabling Brotli for them, and avoid compressing on the fly, which can disable range requests on some servers.
If models are served from a different origin than the app, configure CORS (and Cross-Origin-Resource-Policy when the page is cross-origin isolated), or the
fetch fails in exactly the configuration that threaded inference needs. Large files on some CDNs have per-file size limits; split models into parts below
those limits if necessary, and reassemble or load them as parts on the client.
Letting users decide
Downloading hundreds of megabytes is a decision users should make knowingly, especially on metered or mobile connections. Show the size before starting, offer to wait for Wi-Fi where the Network Information API hints at a cellular connection, and allow downloads to pause and resume. After the download, say how much space the model uses and where to remove it. Treat model storage like a feature users opt into, not a side effect of opening a page; it builds trust and avoids complaints about a website quietly filling a phone’s storage.
Sharing models between features
If several features use the same model — an embedding model used by search and by clustering — cache it once under one URL and load it through a shared helper, rather than letting each feature download its own copy under a different path.
Expected output
The first visit downloads a 400 MB model into OPFS with a progress bar, resuming after a dropped connection; the file is verified with SHA-256 before use; later visits load it from disk in about a second, also offline; an update to version 4 downloads the new file and deletes version 3; and settings show “Models: 412 MB stored” with a remove button.
Gotchas
- Relying on the HTTP cache. Large files may not be kept. Store them explicitly.
- Unversioned model URLs. Stale weights persist. Version or hash every URL.
- Caching unverified downloads. Corrupt files stick. Verify before storing.
- Holding huge files in memory twice. Memory spikes. Stream into OPFS.
- Never deleting old versions. Storage grows with every update. Clean up.
- Silent large downloads. Users on metered connections object. Show the size and ask first.
Performance note
Loading a 400 MB model from OPFS took about 0.9 s on a laptop SSD, compared with 35 s to download it over a 100 Mbit/s connection — a 40× faster start for returning users.
Frequently Asked Questions
Can the runtime read the model directly from OPFS? Some runtimes accept files or streams; otherwise read into memory and pass bytes.
Do private windows keep models? No — private browsing storage is discarded when the session ends.
Should models be split into parts? Splitting helps resumable downloads and memory limits; many large GGUF models already ship split.
Can several origins share a model cache? No — storage is per origin. Serve models from the app’s origin or accept per-origin copies.
What headers should model files be served with? Long immutable cache lifetimes on versioned URLs, range support for resumable downloads, and CORS/CORP if served cross-origin.
Related
- Loading large model weights into linear memory — after loading.
- Caching Wasm with a service worker — the app shell.
- Running small language models in the browser — large weights.
- Persisting a Wasm database to OPFS — OPFS patterns.