Hashing Large Files in the Browser with Wasm
This page answers one task: users select large files — videos, disk images, datasets of several gigabytes — and the application must compute a cryptographic hash in the browser, to deduplicate uploads, verify integrity or sign a manifest, without reading the whole file into memory or freezing the page.
Prerequisites
- [ ] A file input or drag-and-drop source providing
Fileobjects. - [ ] A Wasm hashing module (Rust
sha2orblake3crates, or a C implementation) with an incremental API. - [ ] A dedicated worker to run the hashing.
Why WebCrypto alone is not enough
crypto.subtle.digest("SHA-256", data) is fast — it runs native, often hardware-accelerated code — but it takes the whole input at once. There is no
incremental update/finish API in WebCrypto. To hash a 4 GB file with it, you must have 4 GB in one ArrayBuffer, which exceeds what most browsers will
allocate and certainly what phones can hold. WebCrypto also has no BLAKE3, which many systems use for its speed and tree structure.
An incremental hash in WebAssembly solves both: the file is read in chunks with File.stream(), each chunk is fed to the hasher’s update, and only the
small hash state persists between chunks. Memory use stays at the size of one chunk regardless of file size. For files that fit comfortably in memory
(tens of megabytes), WebCrypto remains simpler and faster for SHA-256; see
WebCrypto vs Wasm for hashing.
Step 1 — expose an incremental hasher from Wasm
use wasm_bindgen::prelude::*;
#[wasm_bindgen]
pub struct Blake3Hasher { inner: blake3::Hasher, buf: Vec<u8> }
#[wasm_bindgen]
impl Blake3Hasher {
#[wasm_bindgen(constructor)]
pub fn new(chunk_capacity: usize) -> Blake3Hasher {
Blake3Hasher { inner: blake3::Hasher::new(), buf: vec![0; chunk_capacity] }
}
/// Address of the reusable input buffer; JavaScript copies each chunk here.
pub fn buf_ptr(&mut self) -> *mut u8 { self.buf.as_mut_ptr() }
pub fn update(&mut self, len: usize) { self.inner.update(&self.buf[..len]); }
pub fn finalize_hex(&self) -> String { self.inner.finalize().to_hex().to_string() }
}
The reusable buffer means chunks are copied once — from the file reader’s buffer into linear memory — and no per-chunk allocation happens on either side.
A Sha256Hasher built on the sha2 crate follows the same shape.
Step 2 — stream the file in a worker
// hash-worker.js
import init, { Blake3Hasher } from "./pkg/hash.js";
const { memory } = await init();
const CHUNK = 4 * 1024 * 1024;
self.onmessage = async ({ data: { file } }) => {
const h = new Blake3Hasher(CHUNK);
const reader = file.stream().getReader({ mode: "byob" });
let buf = new ArrayBuffer(CHUNK), done = 0;
try {
for (;;) {
const { value, done: end } = await reader.read(new Uint8Array(buf));
if (end) break;
new Uint8Array(memory.buffer, h.buf_ptr(), value.length).set(value); // fresh view each time
h.update(value.length);
done += value.length;
buf = value.buffer; // BYOB returns the buffer
self.postMessage({ progress: done / file.size });
}
self.postMessage({ hash: h.finalize_hex() });
} finally {
h.free();
}
};
File objects can be posted to workers cheaply (they are references to data, not copies). Reading with a BYOB (“bring your own buffer”) reader reuses one
ArrayBuffer for all reads, so the JavaScript side allocates nothing per chunk either. Where BYOB readers are unavailable, the default reader works with
slightly more garbage.
Step 3 — report progress without flooding
Posting a message per 4 MB chunk is about 250 messages per gigabyte — acceptable, but throttle to a few per second if chunks are smaller, so the main thread does not spend time on progress updates. Display throughput (MB/s) as well as percent; users understand “1.2 GB of 4 GB, 380 MB/s” better than a bar alone. Support cancellation with an abort flag the worker checks between chunks, so closing the dialog stops reading the file.
Step 4 — choose the algorithm and enable SIMD
BLAKE3 is designed for speed and parallelism; its portable implementation runs well in Wasm, and with Wasm SIMD enabled (-C target-feature=+simd128) it
benefits further. SHA-256 in Wasm cannot use the SHA hardware instructions native code uses, so it is several times slower than WebCrypto’s native
implementation per byte — still fast enough to keep up with many disks, but measure. If the server side already uses SHA-256 for integrity, keep it; if you
control both sides and speed matters, BLAKE3 is the better choice.
Step 5 — verify against the server
A hash is only useful if it matches what another party computes. Test with known vectors (empty input, “abc”, a multi-chunk file) against a native tool
(sha256sum, b3sum) and assert equality in CI. Chunk boundaries must not affect the result — feed the same file with different chunk sizes in tests and
confirm identical digests.
Parallel hashing for BLAKE3
BLAKE3’s tree structure allows hashing different parts of a file in parallel and combining the results. In a browser, that means several workers each
hashing a slice of the file, with a final combination step — but combining requires access to BLAKE3’s internal chaining values, which the high-level API
hides. The blake3 crate exposes lower-level hazmat APIs for this in recent versions; alternatively, use the crate’s Rayon support in a threaded Wasm build.
For most cases, one worker already saturates the disk or the file-reading pipeline, and parallelism adds complexity for little gain; measure before
building it.
Reading speed is often the limit
The browser’s file reading pipeline — and the underlying storage — frequently limits throughput below what the hasher can do, particularly for files on
network drives or slow USB devices. If profiling shows the worker mostly waiting for reader.read, a faster hash will not help. Larger chunk sizes (4–16 MB)
reduce per-read overhead; beyond that, the disk sets the pace.
Hashing for deduplicated or resumable uploads
A common reason to hash in the browser is to avoid uploading what the server already has. The client hashes the file, asks the server whether that hash exists, and uploads only if it does not. For very large files, hash in fixed-size blocks as well as overall — a list of per-block hashes lets the server report which blocks it already has, so an interrupted upload resumes from the first missing block, and edits to a large file re-upload only changed blocks. The incremental hasher supports this naturally: run one hasher per block (resetting between blocks) alongside one for the whole file, both fed from the same chunk stream so the file is read once. BLAKE3’s tree mode is designed for this kind of verified streaming and partial verification, which is another reason to prefer it when you control both ends of the protocol.
Security notes
A hash computed in the browser proves nothing to the server by itself: a malicious client can send any hash. Use client-side hashes for deduplication, progress and user-facing integrity checks, and have the server recompute or verify the hash of what it actually received before trusting it. For integrity guarantees in the other direction — users verifying that a download matches a published hash — computing the hash locally in the browser is exactly right, since the computation happens on the user’s machine with code they loaded.
Expected output
Hashing a 4 GB video with BLAKE3 in a worker completes at about 900 MB/s on a laptop SSD with memory flat at around 10 MB; the UI shows progress and
throughput and stays responsive; cancelling stops reading within one chunk; and digests match b3sum for every test file and chunk size.
Gotchas
- Using
file.arrayBuffer()for large files. It loads everything into memory. Stream instead. - Hashing on the main thread. The page freezes. Use a worker.
- Allocating per chunk. Garbage and copies. Reuse buffers on both sides.
- Untested chunk boundaries. Different chunk sizes must give the same digest. Test them.
- Expecting Wasm SHA-256 to beat WebCrypto. It does not for in-memory data. Use Wasm for streaming and BLAKE3.
Performance note
On a laptop with an NVMe SSD, streaming hashes of a 4 GB file ran at about 900 MB/s for BLAKE3 with Wasm SIMD and about 350 MB/s for SHA-256 in Wasm; WebCrypto SHA-256 over a 512 MB in-memory buffer reached about 1.5 GB/s but could not handle the 4 GB file at all.
Frequently Asked Questions
Can I hash files from the File System Access API?
Yes — FileSystemFileHandle.getFile() returns a File you can stream the same way.
Is MD5 available? Not in WebCrypto; Wasm implementations exist, but use MD5 only for non-security checksums.
Does this work on phones? Yes; throughput is lower, and memory stays flat, which is the point on phones.
Can I hash in a service worker? Possible, but a dedicated worker tied to the page is simpler to control and cancel.
Can the server trust a hash computed in the browser? No — use it for deduplication and progress, and have the server verify the hash of the bytes it actually received.
Related
- WebCrypto vs Wasm for hashing — choosing per file size.
- Streaming file uploads into Wasm memory — the reading pattern.
- Encrypting files client-side with Wasm — the next step after hashing.
- Reporting progress from Wasm to the UI — progress patterns.
← Back to Cryptography & Untrusted Code