Transcoding Video in the Browser with ffmpeg.wasm

This guide answers one task: take a video file the user picked, transcode it entirely in the browser with ffmpeg.wasm, and hand back a downloadable result — without blocking the page, without shipping 20 MB during first paint, and with an honest view of when this is the wrong architecture.

Prerequisites

  • [ ] @ffmpeg/ffmpeg 0.12+ and @ffmpeg/util, installed from npm or loaded from a self-hosted path.
  • [ ] A dev server that can set Cross-Origin-Opener-Policy: same-origin and Cross-Origin-Embedder-Policy: require-corp for the multi-threaded core.
  • [ ] Somewhere to host the core files (ffmpeg-core.js, ffmpeg-core.wasm) — a CDN or your own origin.
  • [ ] A test clip under 20 MB. Debugging with a 2 GB file wastes an afternoon per iteration.

Load the core lazily, never on first paint

The single most important decision is when the core arrives. ffmpeg-core.wasm is roughly 10 MB for the single-threaded build and over 25 MB for the multi-threaded one, so fetching it during page load is indefensible on any page that has other work to do. Load it when the user has committed to the task — after they pick a file, or when they click a button that plainly means “convert this”.

import { FFmpeg } from '@ffmpeg/ffmpeg';
import { toBlobURL, fetchFile } from '@ffmpeg/util';

let ffmpeg;                       // created on demand, kept for the session
async function getFFmpeg(onLog) {
  if (ffmpeg) return ffmpeg;
  ffmpeg = new FFmpeg();
  ffmpeg.on('log', ({ message }) => onLog?.(message));
  const base = '/vendor/ffmpeg';  // self-hosted so the headers are yours to set
  await ffmpeg.load({
    coreURL: await toBlobURL(`${base}/ffmpeg-core.js`, 'text/javascript'),
    wasmURL: await toBlobURL(`${base}/ffmpeg-core.wasm`, 'application/wasm'),
  });
  return ffmpeg;
}

Self-hosting matters more than it looks. A cross-origin core has to satisfy require-corp, which means the CDN must send Cross-Origin-Resource-Policy: cross-origin; when it does not, the load fails with an error that mentions neither headers nor the actual file. Hosting the two files yourself removes that whole class of problem and gives you a cache lifetime you control.

When the 10 MB core should arrive Loading the core during first paint delays every other resource and is wasted for users who never transcode. Loading it after the file picker overlaps the download with the user choosing a file, so the perceived wait is near zero. eager: core competes with first paint html + css + js ffmpeg-core.wasm — 10 MB nobody asked for page usable lazy: core overlaps the file picker html + css + js page usable user browses for a file core downloads underneath, in parallel The lazy row finishes later on the clock and feels immediate, because the only waiting the user does is waiting they were already doing. Cache the core after the first use and the second visit skips the download entirely.

Write the input, run the command, read the output

ffmpeg.wasm exposes an in-memory filesystem. You write the user’s file into it, run an ordinary ffmpeg argument list against those paths, and read the result back out. The arguments are exactly the ones you would type in a terminal, which makes the whole API easy to reason about.

async function transcodeToMp4(file, onProgress) {
  const ff = await getFFmpeg();
  ff.on('progress', ({ progress }) => onProgress(Math.min(1, progress)));

  await ff.writeFile('input', await fetchFile(file));
  await ff.exec([
    '-i', 'input',
    '-c:v', 'libx264', '-preset', 'veryfast', '-crf', '28',
    '-c:a', 'aac', '-b:a', '128k',
    '-movflags', '+faststart',
    'output.mp4',
  ]);
  const data = await ff.readFile('output.mp4');
  await ff.deleteFile('input');
  await ff.deleteFile('output.mp4');
  return new Blob([data.buffer], { type: 'video/mp4' });
}

Three details in that argument list are doing real work. -preset veryfast trades compression ratio for speed, which is the right trade in a browser where the user is watching a progress bar. -crf 28 targets a quality level rather than a bitrate, so the output size adapts to the content. -movflags +faststart moves the MP4 index to the front of the file, which is what makes the result playable while still downloading — free to add and easy to forget.

Expected output

The log stream tells you the job actually ran, and the final line reports the muxing summary:

Input #0, mov,mp4,m4a, from 'input':
  Duration: 00:00:31.55, start: 0.000000, bitrate: 8412 kb/s
  Stream #0:0: Video: h264 (High), yuv420p, 1920x1080, 8210 kb/s, 29.97 fps
[libx264 @ ...] frame I:16 Avg QP:23.11 size: 38122
video:2914kB audio:487kB subtitle:0kB other streams:0kB global headers:0kB

A useful sanity check is the ratio of output to input duration in wall-clock terms. On a modern laptop, a 1080p30 clip transcodes at roughly 0.4–1.2× real time with the single-threaded core and 1.5–4× with the threaded one. If you are seeing 0.05×, you are almost certainly on the single-threaded core with a -preset that is far too slow for a browser.

The path a file takes through ffmpeg.wasm The picked file is copied into the module's virtual filesystem, which lives inside linear memory. ffmpeg reads and writes those paths exactly as it would on disk, and the result is copied back out as a Blob the page can download. File from input ArrayBuffer writeFile linear memory /input /output.mp4 decode · filter · encode no boundary crossings during the run readFile Blob · download or upload to your API Peak memory is roughly input + output + working set, all inside one buffer — which is why a 1 GB source fails on a phone long before it fails on a desktop.

Choosing arguments that suit a browser

The argument list you would use on a build server is the wrong one here. A server transcode optimises for bytes, because it runs once and the output is served thousands of times; a browser transcode optimises for wall-clock time, because a human is watching and the output is usually consumed once.

That inverts several defaults. -preset should sit at veryfast or superfast rather than slow; the file will be perhaps 15% larger and will finish three to five times sooner. -crf should be a little higher than you would pick on a server — 26 to 30 for H.264 — because the extra quality is invisible on the device the result will be viewed on and expensive to produce. Scaling down is the single biggest lever available: -vf scale=-2:720 on a 4K source cuts the pixel count by nine and the encode time by something close to it, and for a clip destined for a web player it is not a quality regression at all.

// a sensible browser profile: bounded resolution, fast preset, quality-targeted
const args = [
  '-i', 'input',
  '-vf', 'scale=-2:min(720\,ih)',   // never upscale, cap at 720p
  '-c:v', 'libx264', '-preset', 'superfast', '-crf', '28',
  '-c:a', 'aac', '-b:a', '96k', '-ac', '2',
  '-movflags', '+faststart',
  'output.mp4',
];

Two arguments are worth adding whenever they apply. -t 60 caps the duration, which turns “the user picked a two-hour recording” from an out-of-memory crash into a predictable clip. -threads 0 lets the threaded core use every available worker, which it does not always do by default depending on how the core was built. And if the job is only a container change — extracting audio, remuxing an MP4 into a fragmented MP4 — use -c copy and skip the encode entirely; that path runs at hundreds of times real time because no pixels are touched at all.

Handling files that are too big

The failure mode for an oversized input is abrupt: an Aborted() in the log and an unusable instance. Since linear memory holds the input, the output and the encoder’s working set simultaneously, the practical ceiling on a desktop is a few hundred megabytes of source and considerably less on a phone. Deciding what to do about that is part of the feature, not an afterthought.

The cheapest mitigation is to refuse early and explain. Read file.size before loading the core at all, and if it exceeds your threshold, offer the server path or ask the user to trim. A two-line check saves a 10 MB download followed by a crash. The next mitigation is segmenting: run the transcode in chunks using -ss and -t to select a window, write each segment out, and concatenate the results at the end. This keeps peak memory bounded by the segment length rather than the file length, at the cost of a concat step and a small quality discontinuity at each boundary if you are not careful to align cuts with keyframes.

if (file.size > 250 * 1024 * 1024) {
  return { tooLarge: true, suggestion: 'server' };   // decide before loading 10 MB of core
}

Finally, be explicit about what you free. Delete input and output files from the virtual filesystem after each job; the module’s memory does not shrink, but the space becomes reusable for the next run, which is the difference between a session that handles ten files and one that dies on the third. If your page handles a long queue, terminating and reloading the instance between jobs is a legitimate strategy — instantiation is fast once the compiled module is cached, and it guarantees a clean heap.

What actually costs the time Loading the module and the input is a small fraction. Decoding and encoding dominate, and encoding dominates that, which is where every settings decision pays off. module load 2.1 s one-off, and cacheable across runs decode 6.4 s encode 15.2 s — preset and CRF decide almost all of this A faster preset is the single biggest lever; resolution is the second, and both beat any loading tweak. Run it in a worker: a transcode is tens of seconds, and no page should be frozen for that long.

Gotchas

  • SharedArrayBuffer is not defined. The multi-threaded core needs cross-origin isolation. Verify with console.log(crossOriginIsolated) — if it is false, fix the headers as described in configuring COOP/COEP headers or load the single-threaded core deliberately.
  • Aborted(). Build with -sASSERTIONS for more info partway through a large file. This is almost always an out-of-memory condition inside linear memory. Cap the input size you accept, or scale the output down so the working set fits.
  • Progress jumps straight to 1. The progress event needs a known duration; for a stream without one, ffmpeg cannot compute a fraction. Parse the time= field out of the log line instead and divide by the duration you read from the input.
  • The page freezes anyway. ffmpeg.wasm 0.12 runs the core in a worker by default, but a synchronous fetchFile on a huge File on the main thread will still block. Read the file in chunks, or hand the File to a worker and do everything there.
  • Output is silent. The audio codec was not compiled into the core build you loaded. Check the log’s configuration line before assuming your arguments are wrong.

Performance note

The dominant cost is the encode, not the boundary. On a four-core laptop, a 30-second 1080p clip re-encoded at -preset veryfast -crf 28 takes about 25 seconds single-threaded and about 8 seconds with the threaded core — roughly 3.2× for four cores, with the shortfall going to the serial muxing stage. The two writeFile/readFile copies together account for well under 2% of that, so optimising them is wasted effort. If you need faster than the threaded core, the answer is not tuning: it is a server, or WebCodecs with hardware acceleration for the codecs the platform exposes.

Frequently Asked Questions

Should this run in the browser at all? It should when privacy matters (the file never leaves the device), when you want to avoid egress and compute costs, or when the clip is short and the user is already waiting. It should not when files are large, when the result is needed by your backend anyway, or when your users are on phones — a 20-minute 4K source will exhaust memory long before it finishes.

How do I cancel a running job? Call ffmpeg.terminate(). There is no way to interrupt a running module from the same thread, so cancellation means destroying the worker; afterwards you must load() again before the next job, which is another reason to keep the core cached.

Can I transcode several files at once? Sequentially, yes, on one instance. In parallel you would need several instances, each with its own multi-megabyte memory, which on typical hardware is slower overall than doing them one at a time — the encode is already using every core it can.

← Back to Media Processing & Codecs in Wasm