Web Crypto vs Wasm for Hashing
This guide answers one task: decide whether to hash with the browser’s built-in SubtleCrypto or with a
compiled module, based on measurements rather than instinct — and know which of the two is a mistake in
each direction.
Prerequisites
- [ ] A secure context;
crypto.subtleis unavailable otherwise. - [ ] A compiled hash module for comparison — BLAKE3 or SHA-256 from a Rust crate works well.
- [ ] A benchmark harness that discards warm-up and reports percentiles.
- [ ] Inputs at several sizes: 64 B, 4 kB, 1 MB, 64 MB. The answer changes across them.
What each one is good at
crypto.subtle.digest is native code, often with hardware acceleration for SHA-2, running outside the
JavaScript heap. Per byte, nothing you compile will beat it for the algorithms it implements. Its
weaknesses are structural rather than computational: every call is asynchronous, it is one-shot rather
than incremental, and it supports a fixed list of algorithms.
A compiled module is synchronous, incremental if you write it that way, and supports whatever you compile. Per byte it is typically 1.5–4× slower than the native SHA-256 implementation, and considerably faster if you choose an algorithm the platform does not have — BLAKE3 with SIMD outruns native SHA-256 on large inputs on most machines.
Measure both, properly
The comparison is only meaningful with warm-up discarded and the same input for both. Note that the platform call is asynchronous, so the measurement includes promise scheduling — which is real cost your application pays.
async function benchSubtle(bytes, n = 200) {
await crypto.subtle.digest('SHA-256', bytes); // warm up
const t0 = performance.now();
for (let i = 0; i < n; i++) await crypto.subtle.digest('SHA-256', bytes);
return (performance.now() - t0) / n;
}
function benchWasm(mod, ptr, len, n = 200) {
mod.exports.sha256(ptr, len); // warm up
const t0 = performance.now();
for (let i = 0; i < n; i++) mod.exports.sha256(ptr, len);
return (performance.now() - t0) / n;
}
For the module, make sure the input is already in linear memory before timing, or you are measuring the
copy rather than the hash. If your real workload does include that copy — because the data arrives as a
JavaScript ArrayBuffer — then include it deliberately and say so.
Expected output
Typical numbers on a modern laptop, SHA-256, same input for both:
size subtle (ms) wasm (ms) winner
64 B 0.061 0.0021 wasm (29×)
4 kB 0.068 0.0094 wasm (7×)
256 kB 0.28 0.42 subtle
1 MB 0.94 1.61 subtle
64 MB 58.2 99.4 subtle
The crossover for SHA-256 sits somewhere around 100 kB on most machines. Below it, the asynchronous call overhead dominates and the module is dramatically faster; above it, native throughput wins and the gap widens with size.
Many small hashes: the module’s real win
The case where a compiled hash is unambiguously better is a loop over many small inputs — deduplicating
records, building a Merkle tree, computing content addresses for a list of chunks. Each await in that
loop costs a microtask round trip, and ten thousand of them is tens of milliseconds of pure overhead
before any hashing happens.
// 10,000 short strings
// subtle: await in a loop → 610 ms
// subtle: Promise.all batched → 240 ms
// wasm: synchronous loop → 38 ms
Batching with Promise.all recovers some of it but not most, because the per-call cost is not only
scheduling. The synchronous module simply does not have the problem, and the difference is large enough
to change which features are feasible.
Streaming input the platform cannot handle
crypto.subtle.digest takes a complete buffer. To hash a 2 GB file you would have to hold all of it in
memory, which is not acceptable in a tab. A compiled hash with an incremental interface handles it in
bounded memory:
const h = mod.exports.hash_init();
const reader = file.stream().getReader();
const view = new Uint8Array(mod.exports.memory.buffer, mod.exports.chunk_ptr(), CHUNK);
for (;;) {
const { value, done } = await reader.read();
if (done) break;
for (let off = 0; off < value.length; off += CHUNK) {
const slice = value.subarray(off, off + CHUNK);
view.set(slice);
mod.exports.hash_update(h, slice.length);
}
}
const digest = readDigest(mod, mod.exports.hash_final(h));
Peak memory is one chunk. This is the single most common legitimate reason to compile a hash: not speed, but the ability to process data you cannot hold.
Algorithms the platform does not have
The last category is straightforward: if you need BLAKE3, BLAKE2b, SHA-3, xxHash or a domain-specific
construction, SubtleCrypto cannot help and the decision is made for you.
BLAKE3 is worth calling out because it inverts the throughput conclusion. With SIMD and its tree structure, a compiled BLAKE3 commonly reaches 2–4 GB/s single-threaded, ahead of native SHA-256 on the same machine, and it parallelises across workers for large inputs. If you control both ends — content addressing inside your own system rather than interoperating with something that specifies SHA-256 — it is usually the better choice on every axis except familiarity.
Content addressing: a worked decision
A concrete case makes the tradeoffs easier to see. Suppose an application splits uploads into 1 MB chunks and computes a digest per chunk to deduplicate against what the server already holds. A 500 MB upload is 500 hashes of 1 MB each.
With crypto.subtle, that is 500 asynchronous calls at roughly 0.94 ms of hashing plus call overhead —
around 520 ms in total, all of it off the main thread’s critical path but each result arriving as a
microtask that interleaves with rendering. With a compiled SHA-256 the same work takes about 810 ms of
synchronous compute, which must be in a worker or the page freezes. With BLAKE3 it takes around 160 ms,
and the tree structure means several workers can hash disjoint ranges and combine the results.
The decision follows from a question that has nothing to do with speed: does the digest have to interoperate? If the server, or a specification, or another client expects SHA-256, the algorithm is fixed and the only choice is where to run it — and for 1 MB chunks the platform wins. If the digest is internal to your system, BLAKE3 in a worker is three times faster, streams naturally, and parallelises.
That is the shape of almost every real decision here. Interoperability constrains the algorithm; the algorithm mostly determines the implementation; and the measurement only settles the remaining question of where the work runs.
Combining both in one application
Nothing requires a single choice. A realistic application uses SubtleCrypto for the standard operations
it covers — verifying a signature, deriving a key, hashing a moderate buffer — and a module for the
specific workload the platform handles badly.
Keep the boundary explicit. A small hash.js module that exposes digestBuffer, digestStream and
digestMany, routing each to whichever implementation is right, means call sites do not encode the
decision and you can revisit it with a measurement rather than a refactor. It also gives you one place to
put the lazy loading: the compiled module should not be fetched at all by a session that only ever hashes
one small buffer.
Gotchas
crypto.subtleis undefined. Insecure context. It is not a browser support problem; it is HTTP.- Comparing a cold module against a warm platform call. Instantiate, warm up, then measure.
- Measuring with the input outside linear memory. The copy is then in your measurement. Decide whether it belongs there.
- Assuming SHA-1 is available for a checksum. Some browsers restrict it in
SubtleCrypto; a module has no such restriction, which is occasionally the only reason to use one. - Hashing a
Fileby reading it entirely. Stream it. The one-shot API cannot, and memory will not forgive you. - Using a fast non-cryptographic hash where a cryptographic one is needed. xxHash is excellent for hash tables and useless against an adversary.
Performance note
Measured on an M2 laptop: SHA-256 over 1 MB took 0.94 ms with crypto.subtle and 1.61 ms with a compiled
Rust implementation, while BLAKE3 over the same input took 0.31 ms — three times faster than the platform’s
own SHA-256. Over 64 B inputs in a loop, the module was roughly 29× faster than the platform because the
comparison is almost entirely call overhead. Both conclusions are correct simultaneously, which is why a
single “which is faster” answer is always wrong.
Frequently Asked Questions
Should I replace SubtleCrypto everywhere with a module? No. For occasional hashes of substantial buffers it is faster, native, free to ship and better reviewed. Reach for a module when the workload is many small hashes, a stream, or an algorithm it lacks.
Does hashing in a worker change the comparison? It removes the main-thread blocking concern from the synchronous module, which is otherwise the one real argument against it. Both APIs are available in workers.
What about HMAC and key derivation?
Use the platform. SubtleCrypto implements HMAC and PBKDF2 natively with non-extractable key handling,
which is a security advantage a module cannot match — the key never enters your JavaScript heap at all.
Related
- Implementing Argon2 password hashing in Wasm — the case the platform genuinely lacks.
- Measuring Wasm vs JavaScript throughput — the benchmarking method.
- Writing constant-time code for Wasm — comparing digests safely.
← Back to Cryptography & Untrusted Code