Running Whisper Speech Recognition in the Browser
This page answers one task: an application needs speech-to-text — meeting notes, voice commands, subtitles for uploaded videos — and the audio should not leave the user’s device. You want OpenAI’s Whisper model running in the browser through WebAssembly, fast enough to be useful, with model files that do not have to download on every visit.
Prerequisites
- [ ] whisper.cpp built for the web with Emscripten (the project includes a WebAssembly example), or a packaged build.
- [ ] A cross-origin isolated page (COOP and COEP headers) to enable threads.
- [ ] A model file in whisper.cpp’s format (GGML/GGUF), ideally quantised.
How whisper.cpp runs in Wasm
whisper.cpp is a C/C++ implementation of Whisper inference with its own tensor library. Compiled with Emscripten, it runs the encoder and decoder on the CPU using Wasm SIMD for the matrix operations and pthreads (Web Workers sharing memory) for parallelism. The model’s weights load into linear memory; audio is converted to a log-mel spectrogram in Wasm; the encoder processes 30-second windows; the decoder produces tokens, which become text.
Performance depends on three things: model size (tiny, base, small, medium — from tens of megabytes to over a gigabyte), quantisation (8-bit or 5-bit weights are smaller and faster with a small accuracy cost), and thread count. On a modern laptop, tiny and base models transcribe faster than real time; small models approach real time; medium and larger are impractical for most browsers because of memory limits and speed.
Step 1 — build or obtain the Wasm module
whisper.cpp’s repository contains Emscripten build configuration for a browser example. Build with SIMD and threads enabled:
emcmake cmake -B build-wasm -DWHISPER_WASM_SINGLE_FILE=OFF
cmake --build build-wasm -j --target libmain
# produces a JS loader and a .wasm module using SIMD and pthreads
Exact targets and flags change between versions; follow the example’s current README. Keep the build reproducible (pinned Emscripten and whisper.cpp versions), since performance and model format support evolve quickly.
Step 2 — serve with cross-origin isolation
Threads need SharedArrayBuffer, which needs the page to be cross-origin isolated:
Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp
Without isolation, the module must run single-threaded, which is several times slower. Check crossOriginIsolated at startup and warn or fall back to a
smaller model when it is false. Any third-party resources on the page then need CORS or Cross-Origin-Resource-Policy headers.
Step 3 — load and cache the model
Download the model once and store it in the Cache API or OPFS, keyed by name and hash, then pass the bytes to the module (which copies them into its file system or memory):
async function loadModel(url) {
const cache = await caches.open("whisper-models-v1");
let res = await cache.match(url);
if (!res) { res = await fetch(url); await cache.put(url, res.clone()); }
return new Uint8Array(await res.arrayBuffer());
}
const model = await loadModel("/models/ggml-base.en-q5_1.bin");
Module.FS_createDataFile("/", "model.bin", model, true, false);
const ctx = Module.init("model.bin");
Show download progress for first-time users — a 55 MB model takes a while on mobile networks — and request persistent storage so the browser is less likely to evict it. See caching model files for offline inference.
Step 4 — prepare audio at 16 kHz mono
Whisper expects 16 kHz mono floating-point samples. For files, decode with Web Audio and resample with an OfflineAudioContext at 16 kHz:
async function toWhisperPcm(arrayBuffer) {
const decoded = await new AudioContext().decodeAudioData(arrayBuffer);
const offline = new OfflineAudioContext(1, Math.ceil(decoded.duration * 16000), 16000);
const src = offline.createBufferSource();
src.buffer = decoded; src.connect(offline.destination); src.start();
const rendered = await offline.startRendering();
return rendered.getChannelData(0); // Float32Array at 16 kHz
}
For the microphone, capture with getUserMedia and an AudioWorklet, downsample to 16 kHz, and accumulate chunks.
Step 5 — run transcription off the main thread
Inference takes seconds to minutes. Run it in a worker (Emscripten’s pthreads already use workers for parallel work, but the call itself should not run on the main thread either), report progress per segment, and stream segments to the UI as they are decoded so users see text appear rather than waiting for the whole file. Allow cancellation between segments.
Real-time and streaming transcription
Whisper processes 30-second windows, so “live” transcription works by repeatedly transcribing a sliding window of recent audio (for example the last 10 seconds, every 2 seconds) and merging overlapping results. That costs much more compute than file transcription per second of audio, so it is practical only with tiny or base models on fast machines. Voice activity detection (only sending speech, skipping silence) reduces the load substantially and avoids hallucinated text during silence, a known Whisper behaviour.
Accuracy, languages and hallucinations
English-only models (.en) are more accurate for English at the same size; multilingual models detect and transcribe many languages. Quantisation slightly
reduces accuracy; measure on your audio. Whisper can produce repeated or invented text on silence, music or noise — trim silence, set a no-speech threshold,
and treat outputs for low-confidence segments carefully, especially for captions shown to others.
Choosing a model per device
One model rarely fits every user. Decide at runtime from what the browser reports and a quick measurement: crossOriginIsolated (threads or not),
navigator.hardwareConcurrency, navigator.deviceMemory where available, and a short benchmark that transcribes a few seconds of bundled audio with the tiny
model. From those, pick the largest model expected to run at the speed your feature needs — faster than real time for live captions, a few times slower for
batch transcription where users can wait — and let users override the choice in settings. Remember the download cost: offering the small model to a phone
user on a metered connection may be a worse experience than running the base model a little less accurately. Store the choice so later visits skip the
benchmark, and re-evaluate when the app updates its models.
Memory limits and long recordings
Model weights, the encoder’s working buffers and audio samples all live in linear memory. A small model plus an hour of 16 kHz float audio (about 230 MB of samples) can exceed what a phone tab tolerates. Process long recordings in chunks instead of loading everything: decode and resample the file in segments, transcribe each 30-second window, and discard audio already processed. Keep only the text and timestamps. That bounds memory to the model plus a few windows of audio regardless of recording length, and it also allows progress reporting and cancellation at natural boundaries.
Privacy expectations
Local transcription is attractive precisely because audio stays on the device. Make that true and visible: no analytics on audio or transcripts by default, a strict Content Security Policy that prevents unexpected uploads, and clear wording about what is processed locally.
Expected output
A 10-minute recording transcribes in about 2 minutes with the base English q5 model and 8 threads on a laptop; the model downloads once (55 MB) and loads from cache in under a second afterwards; text segments with timestamps stream into the UI; audio never leaves the device; and without cross-origin isolation the app warns and switches to the tiny model.
Gotchas
- No cross-origin isolation. Threads are unavailable and speed drops sharply. Set COOP and COEP.
- Wrong sample rate. Whisper needs 16 kHz mono. Resample.
- Large models on phones. Memory limits and speed. Choose tiny or base for mobile.
- Downloading the model every visit. Cache it.
- Transcribing silence. Hallucinated text. Use voice activity detection.
- Loading entire long recordings into memory. Phones run out. Process in chunks.
Performance note
On an 8-core laptop, the base English q5 model transcribed audio at about 5× real time with 8 threads and about 1.4× with 1 thread; on a mid-range phone with 4 threads, the tiny model reached about 3× real time.
Frequently Asked Questions
Can WebGPU make it faster? GPU-backed runtimes exist and can be much faster where WebGPU is available; whisper.cpp’s Wasm build runs on the CPU.
Does it work offline? Yes, once the page, module and model are cached.
How accurate is the tiny model? Usable for clear speech and commands; noticeably worse than small or larger models on difficult audio.
Can I get word-level timestamps? whisper.cpp supports token-level timestamps with some configuration; segment timestamps are the default.
How do I transcribe hour-long recordings on a phone? Decode, resample and transcribe in chunks, discarding processed audio, so memory stays bounded by the model and a few windows.
How should the app pick a model size? From isolation status, core count, device memory and a short benchmark, with a user override stored in settings.
Does the audio ever leave the device? Not with this setup — decoding, resampling and inference all run locally; enforce it with a strict Content Security Policy.
Related
- Multi-threaded inference with Wasm threads — threads for inference.
- Caching model files for offline inference — model storage.
- Quantizing models for Wasm inference — smaller weights.
- Encoding audio in the browser with Wasm — audio handling.