Running Image Classification on Webcam Frames

This page answers one task: an application analyses a live camera feed — recognising gestures, objects, document types or product categories — entirely on the device. You want frames captured, preprocessed and classified by a small model running in WebAssembly, with a steady frame rate and a UI that never stutters.

Prerequisites

  • [ ] A small image-classification model in ONNX or TFLite format (MobileNet- or EfficientNet-lite-class, ideally quantised).
  • [ ] ONNX Runtime Web or TensorFlow.js with the Wasm backend.
  • [ ] Camera permission via getUserMedia over HTTPS.

The pipeline and where time goes

Each analysed frame passes through capture (getting pixels from the video element), preprocessing (crop, resize to the model’s input size such as 224×224, convert RGBA bytes to normalised float tensors in the right channel order), inference, and postprocessing (softmax, top-k, mapping indices to labels). On a laptop, inference for a small quantised model on the Wasm backend with SIMD takes on the order of 10–30 ms; preprocessing in JavaScript can take nearly as long if written naively; and capturing pixels through a canvas costs a few milliseconds.

The camera delivers 30 or 60 frames per second, faster than most devices can classify. A good pipeline therefore never queues frames: it processes the latest frame when the previous result is done and drops the rest. Users perceive smooth video plus labels that update several times per second as responsive; what they notice is a frozen preview or labels lagging seconds behind.

A webcam classification loop that drops frames The video element shows the camera at full rate. When the worker is idle, the latest frame is captured as an ImageBitmap and transferred to the worker. Wasm resizes and normalises it into the model's input tensor, the model runs, and the top labels return to the main thread, which smooths and displays them. Frames arriving while the worker is busy are skipped. camera → video 30–60 fps preview capture latest frame only if worker idle Wasm preprocess resize + normalise inference in worker Wasm backend, SIMD smooth + display labels ~10 updates/s

Step 1 — capture frames without blocking

const stream = await navigator.mediaDevices.getUserMedia({ video: { width: 640, height: 480, facingMode: "environment" } });
video.srcObject = stream; await video.play();

let busy = false;
function onFrame() {
  if (!busy) {
    busy = true;
    createImageBitmap(video).then((bitmap) => worker.postMessage({ bitmap }, [bitmap]));
  }
  video.requestVideoFrameCallback(onFrame);
}
video.requestVideoFrameCallback(onFrame);
worker.onmessage = ({ data }) => { busy = false; showLabels(data.top); };

requestVideoFrameCallback fires once per new video frame. createImageBitmap grabs the frame efficiently, and ImageBitmap is transferable, so the worker receives it without a copy. The busy flag implements frame dropping.

Step 2 — preprocess in the worker with Wasm

In the worker, draw the bitmap into an OffscreenCanvas at a manageable size, read the pixels, and let a Wasm function crop, resize and normalise them into the model’s input layout:

const canvas = new OffscreenCanvas(256, 256);
const ctx = canvas.getContext("2d", { willReadFrequently: true });
self.onmessage = async ({ data: { bitmap } }) => {
  ctx.drawImage(bitmap, 0, 0, 256, 256); bitmap.close();
  const rgba = ctx.getImageData(16, 16, 224, 224).data;          // centre crop
  pre.rgba_to_nchw(rgba, inputTensorView, MEAN, STD);             // Wasm: uint8 RGBA → float32 NCHW
  const out = await session.run({ input: inputTensor });
  self.postMessage({ top: topK(out.logits.data, 3) });
};

A small Rust or C function converting RGBA bytes to normalised planar floats with SIMD is several times faster than an equivalent JavaScript loop, and it writes directly into the buffer the runtime reads. Check the model’s expected layout (NCHW or NHWC), channel order (RGB or BGR) and normalisation constants; a mismatch gives confidently wrong predictions.

Step 3 — run inference on the Wasm backend

import * as ort from "onnxruntime-web";
ort.env.wasm.numThreads = crossOriginIsolated ? 4 : 1;
const session = await ort.InferenceSession.create("/models/mobilenetv3-small-q8.onnx", { executionProviders: ["wasm"] });
const inputTensor = new ort.Tensor("float32", new Float32Array(3 * 224 * 224), [1, 3, 224, 224]);

Create the session once. Reuse the input tensor’s buffer for every frame to avoid allocations. Quantised (int8) models run faster on the Wasm backend; if WebGPU is available, its backend may be faster still, but for small models the Wasm backend is often competitive once data transfer costs are included.

Per-frame time budget on a laptop For a small quantised classifier on the Wasm backend, capture takes about 2 milliseconds, Wasm preprocessing about 1 millisecond, inference about 15 milliseconds and postprocessing well under a millisecond, allowing roughly 50 classifications per second; JavaScript preprocessing would add several milliseconds. stage time per frame notes capture (ImageBitmap + draw) ~2 ms transferred to worker preprocess in Wasm ~1 ms JS version ~6 ms inference (q8, Wasm, 4 threads) ~15 ms dominant cost postprocess (softmax, top-k) < 0.2 ms negligible

Step 4 — smooth predictions

Frame-by-frame predictions flicker as the camera moves. Average class probabilities over the last few results (an exponential moving average works well), require a label to stay on top for a short time before announcing it, and hide predictions below a confidence threshold. This stabilises the UI without adding noticeable lag.

Step 5 — measure end-to-end latency and rate

Measure classifications per second and the delay from a frame’s capture time (requestVideoFrameCallback provides metadata including mediaTime) to the displayed label. On phones, reduce input resolution, use the smallest model that meets accuracy needs, and cap the analysis rate (for example 5 per second) to save battery and avoid thermal throttling, which otherwise slows everything after a minute or two.

Privacy and permissions

Camera access requires a permission prompt and HTTPS. Processing frames locally means images never leave the device; say so, and avoid sending frames or thumbnails to analytics. Stop the camera tracks (stream.getTracks().forEach(t => t.stop())) when the feature is closed, so the camera indicator turns off.

Regions of interest instead of whole frames

Classifying the whole frame works when the subject fills the view; often it does not. A document scanner cares about the paper, a product recogniser about the item in the user’s hand. Cropping to a region of interest before classification improves accuracy far more than a larger model would. The region can be fixed (a guide rectangle drawn over the preview, which users align the subject with), derived from a cheap detector run every few frames, or tracked between detections. The crop happens in the same Wasm preprocessing step, so it costs nothing extra: pass the rectangle along with the frame, and resize only that region to the model’s input size. Showing the guide rectangle in the UI also teaches users how to hold the camera, which improves results more than any amount of post-processing.

Adapting to the device at runtime

Devices vary by an order of magnitude in inference speed. Measure the first few inference times after start-up and adapt: if results take longer than the target interval, lower the analysis rate, switch to a smaller input resolution, or load a smaller model variant; if the device is fast, raise the rate up to a cap. Watch for slowdowns over time as well — phones throttle when warm — and back off before the UI starts to lag. An adaptive loop delivers a consistent experience across a cheap phone and a fast laptop without separate builds.

Testing without a camera

Camera-dependent features are awkward to test. Feed recorded video instead: a <video> element playing a test clip exposes the same requestVideoFrameCallback API, so the whole pipeline runs against known footage. Browsers’ automation tools can also supply fake camera streams (Chromium’s --use-file-for-fake-video-capture flag), which lets end-to-end tests exercise permission handling and capture together. Assert on the labels produced for key moments in the clip, with tolerance for timing.

Expected output

The preview runs at 30 fps while classification runs at about 40 results per second on a laptop and 8 per second on a mid-range phone; labels update smoothly with moving-average filtering; preprocessing in Wasm takes about 1 ms per frame; frames are never queued; and closing the feature stops the camera.

Gotchas

  • Queuing every frame. Latency grows without bound. Drop frames while busy.
  • Wrong tensor layout or normalisation. Confident wrong answers. Match the model’s preprocessing.
  • Inference on the main thread. The preview stutters. Use a worker.
  • Allocating tensors per frame. GC pauses. Reuse buffers.
  • Leaving the camera on. Users notice the indicator. Stop tracks on close.
  • Fixed analysis rate on every device. Slow phones lag and heat up. Adapt the rate to measured inference time.

Performance note

Moving preprocessing from a JavaScript loop to a SIMD Wasm function cut it from about 6 ms to 1 ms per frame, raising classifications per second on a laptop from about 30 to about 40.

Classifications per second on a laptop Classification results per second for a quantised MobileNet-class model on the Wasm backend with JavaScript preprocessing and with SIMD Wasm preprocessing, both with frame dropping. results per second JS preprocessing 30 /s Wasm SIMD preprocessing 40 /s

Frequently Asked Questions

Can I use MediaPipe or TensorFlow.js instead? Yes — both offer browser runtimes with Wasm backends; the frame-handling patterns are the same.

Does VideoFrame from WebCodecs help? MediaStreamTrackProcessor (where supported) gives VideoFrames directly, avoiding the video element; copyTo can write pixels into a buffer.

How do I handle front versus back cameras? Use facingMode constraints and mirror the preview for the front camera only.

Should models run on the GPU? For larger models, WebGPU often wins; for tiny classifiers, Wasm is competitive and more widely available.

How can I improve accuracy without a bigger model? Crop to a region of interest — a guide rectangle or a detector’s box — before classification.

How do I test the pipeline without a camera? Play a recorded clip in a video element, or use the browser’s fake camera capture in automated tests.

What should happen when the phone gets warm? Detect rising inference times and lower the analysis rate or resolution before the preview starts to stutter.

Should the guide rectangle be shown to users? Yes — it tells users where to place the subject and makes the crop match what the model expects.

← Back to Machine Learning Inference in the Browser