Caching Model Files for Offline Inference

This page answers one task: an application runs machine-learning models in the browser — speech recognition, embeddings, a small language model — and their weights are tens to hundreds of megabytes. Downloading them on every visit is slow and expensive for users. You want the files stored locally once, verified, versioned, available offline, and cleaned up when models change.

Prerequisites

  • [ ] Model files served from your origin or a CDN, with stable versioned URLs.
  • [ ] A runtime that accepts model bytes (ONNX Runtime Web, whisper.cpp, llama.cpp-style runtimes, TensorFlow.js).
  • [ ] Optionally a service worker, for offline loading of the app itself.

Where large files can live

Browsers offer three storage options suitable for large binaries. The Cache API stores Response objects keyed by request URL; it is simple, works in pages, workers and service workers, and fits “download a URL once” naturally. OPFS (Origin Private File System) stores files that can be read and written in chunks, including with synchronous access handles in workers; it suits very large files, partial writes and resumable downloads. IndexedDB can store Blobs and ArrayBuffers, but large binary values are less convenient there. All three count against the origin’s storage quota, which is typically a large fraction of free disk space but varies by browser and can be evicted under pressure unless the origin has persistent storage.

The HTTP cache is not a reliable substitute: browsers may skip caching very large responses, evict them early, or partition them, and the application cannot inspect or control what is there.

Storage options for model weights The Cache API stores whole responses by URL and is simplest for download-once files. OPFS stores files with chunked writes and synchronous access in workers, suiting very large or resumable downloads. IndexedDB stores blobs but is less convenient for huge binaries. The HTTP cache cannot be controlled by the app. option best for chunked writes? app control Cache API download-once files by URL no full OPFS very large or resumable files yes full IndexedDB (Blob) smaller binaries + metadata limited full HTTP cache ordinary assets no none

Step 1 — version and hash model URLs

Put a version or content hash in every model URL — /models/whisper-base.en-q5_1.v3.bin or /models/<sha256>.bin — and never change the bytes behind a URL. The URL then doubles as the cache key: a new model is a new URL, and stale caches cannot serve outdated weights. Publish a small manifest listing each model’s URL, size and SHA-256, so the client knows what to expect before downloading.

Step 2 — download with progress and store in the Cache API

async function fetchModel({ url, size, sha256 }, onProgress) {
  const cache = await caches.open("models-v1");
  const hit = await cache.match(url);
  if (hit) return new Uint8Array(await hit.arrayBuffer());

  const res = await fetch(url);
  const reader = res.body.getReader();
  const chunks = []; let received = 0;
  for (;;) {
    const { done, value } = await reader.read();
    if (done) break;
    chunks.push(value); received += value.length;
    onProgress(received / size);
  }
  const bytes = new Uint8Array(received);
  let off = 0; for (const c of chunks) { bytes.set(c, off); off += c.length; }

  const digest = new Uint8Array(await crypto.subtle.digest("SHA-256", bytes));
  if (toHex(digest) !== sha256) throw new Error("model integrity check failed");
  await cache.put(url, new Response(bytes, { headers: { "Content-Type": "application/octet-stream" } }));
  return bytes;
}

Verifying the hash before caching prevents a truncated or corrupted download from being stored and reused forever. For files of several hundred megabytes, holding the whole file in memory twice (chunks plus the assembled array) is costly; stream directly into OPFS instead (Step 3).

Step 3 — use OPFS for very large or resumable downloads

Write chunks to an OPFS file as they arrive, and resume with an HTTP Range request after interruptions:

// in a worker
const dir = await navigator.storage.getDirectory();
const handle = await dir.getFileHandle("llm-q4.gguf.part", { create: true });
const access = await handle.createSyncAccessHandle();
let offset = access.getSize();                                   // resume from what we have
const res = await fetch(url, { headers: offset ? { Range: `bytes=${offset}-` } : {} });
const reader = res.body.getReader();
for (;;) {
  const { done, value } = await reader.read();
  if (done) break;
  access.write(value, { at: offset }); offset += value.length;
  postMessage({ progress: offset / size });
}
access.flush(); access.close();
// verify hash by streaming the file, then rename .part → final name

Hashing a large file incrementally needs an incremental hash (WebCrypto’s digest needs the whole buffer); a Wasm SHA-256 or BLAKE3 hasher fits, as in hashing large files in the browser with Wasm.

Downloading and caching a large model The app reads the manifest with the model URL, size and hash. If the file is already cached and verified it loads directly. Otherwise it downloads in chunks into OPFS, resuming with Range requests after interruptions, verifies the hash incrementally, marks the file complete, and loads it into the runtime. read manifest url, size, sha256 cached + verified? load immediately stream into OPFS resume with Range verify hash incremental load into runtime offline next time

Step 4 — ask for persistence and check quota

Before downloading, check navigator.storage.estimate() and make sure the model fits with room to spare; tell users how much space will be used. Request persistent storage (navigator.storage.persist()) after the user opts into downloading models, so the browser is less likely to evict them. Handle eviction anyway: on each start, check the cache, and if the model is missing, explain and offer to download it again rather than failing with a runtime error.

Step 5 — clean up old versions

When the manifest lists a new model version, download it, switch over, and delete the old file — otherwise every model update adds hundreds of megabytes. Keep the list of expected files in the manifest and remove anything in the model cache or OPFS directory that is not on it. Provide a settings entry showing stored models with sizes and a button to remove them.

Offline loading of the whole app

Caching models makes inference offline-capable only if the app itself loads offline. Use a service worker to cache the application shell, JavaScript, Wasm runtime files and the manifest. Then a returning user can open the app without a network, load the model from OPFS or the Cache API, and run inference. Keep model files out of the service worker’s precache list — they are large and handled separately with progress and verification.

Serving model files well

Client-side caching starts with how files are served. Send model files with Cache-Control: public, max-age=31536000, immutable (safe because URLs are versioned), support HTTP range requests so downloads can resume, and serve them from a CDN close to users. Compression rarely helps quantised weights much — they are dense binary data — so measure before enabling Brotli for them, and avoid compressing on the fly, which can disable range requests on some servers. If models are served from a different origin than the app, configure CORS (and Cross-Origin-Resource-Policy when the page is cross-origin isolated), or the fetch fails in exactly the configuration that threaded inference needs. Large files on some CDNs have per-file size limits; split models into parts below those limits if necessary, and reassemble or load them as parts on the client.

Letting users decide

Downloading hundreds of megabytes is a decision users should make knowingly, especially on metered or mobile connections. Show the size before starting, offer to wait for Wi-Fi where the Network Information API hints at a cellular connection, and allow downloads to pause and resume. After the download, say how much space the model uses and where to remove it. Treat model storage like a feature users opt into, not a side effect of opening a page; it builds trust and avoids complaints about a website quietly filling a phone’s storage.

Sharing models between features

If several features use the same model — an embedding model used by search and by clustering — cache it once under one URL and load it through a shared helper, rather than letting each feature download its own copy under a different path.

Expected output

The first visit downloads a 400 MB model into OPFS with a progress bar, resuming after a dropped connection; the file is verified with SHA-256 before use; later visits load it from disk in about a second, also offline; an update to version 4 downloads the new file and deletes version 3; and settings show “Models: 412 MB stored” with a remove button.

Gotchas

  • Relying on the HTTP cache. Large files may not be kept. Store them explicitly.
  • Unversioned model URLs. Stale weights persist. Version or hash every URL.
  • Caching unverified downloads. Corrupt files stick. Verify before storing.
  • Holding huge files in memory twice. Memory spikes. Stream into OPFS.
  • Never deleting old versions. Storage grows with every update. Clean up.
  • Silent large downloads. Users on metered connections object. Show the size and ask first.

Performance note

Loading a 400 MB model from OPFS took about 0.9 s on a laptop SSD, compared with 35 s to download it over a 100 Mbit/s connection — a 40× faster start for returning users.

Time to have a 400 MB model ready Seconds until a 400 megabyte model is available to the runtime when downloaded over a 100 megabit connection and when read from OPFS on a laptop. seconds to model ready download (100 Mbit/s) 35 s read from OPFS 0.9 s

Frequently Asked Questions

Can the runtime read the model directly from OPFS? Some runtimes accept files or streams; otherwise read into memory and pass bytes.

Do private windows keep models? No — private browsing storage is discarded when the session ends.

Should models be split into parts? Splitting helps resumable downloads and memory limits; many large GGUF models already ship split.

Can several origins share a model cache? No — storage is per origin. Serve models from the app’s origin or accept per-origin copies.

What headers should model files be served with? Long immutable cache lifetimes on versioned URLs, range support for resumable downloads, and CORS/CORP if served cross-origin.

← Back to Machine Learning Inference in the Browser