Running Image Classification on Webcam Frames
This page answers one task: an application analyses a live camera feed — recognising gestures, objects, document types or product categories — entirely on the device. You want frames captured, preprocessed and classified by a small model running in WebAssembly, with a steady frame rate and a UI that never stutters.
Prerequisites
- [ ] A small image-classification model in ONNX or TFLite format (MobileNet- or EfficientNet-lite-class, ideally quantised).
- [ ] ONNX Runtime Web or TensorFlow.js with the Wasm backend.
- [ ] Camera permission via
getUserMediaover HTTPS.
The pipeline and where time goes
Each analysed frame passes through capture (getting pixels from the video element), preprocessing (crop, resize to the model’s input size such as 224×224, convert RGBA bytes to normalised float tensors in the right channel order), inference, and postprocessing (softmax, top-k, mapping indices to labels). On a laptop, inference for a small quantised model on the Wasm backend with SIMD takes on the order of 10–30 ms; preprocessing in JavaScript can take nearly as long if written naively; and capturing pixels through a canvas costs a few milliseconds.
The camera delivers 30 or 60 frames per second, faster than most devices can classify. A good pipeline therefore never queues frames: it processes the latest frame when the previous result is done and drops the rest. Users perceive smooth video plus labels that update several times per second as responsive; what they notice is a frozen preview or labels lagging seconds behind.
Step 1 — capture frames without blocking
const stream = await navigator.mediaDevices.getUserMedia({ video: { width: 640, height: 480, facingMode: "environment" } });
video.srcObject = stream; await video.play();
let busy = false;
function onFrame() {
if (!busy) {
busy = true;
createImageBitmap(video).then((bitmap) => worker.postMessage({ bitmap }, [bitmap]));
}
video.requestVideoFrameCallback(onFrame);
}
video.requestVideoFrameCallback(onFrame);
worker.onmessage = ({ data }) => { busy = false; showLabels(data.top); };
requestVideoFrameCallback fires once per new video frame. createImageBitmap grabs the frame efficiently, and ImageBitmap is transferable, so the worker
receives it without a copy. The busy flag implements frame dropping.
Step 2 — preprocess in the worker with Wasm
In the worker, draw the bitmap into an OffscreenCanvas at a manageable size, read the pixels, and let a Wasm function crop, resize and normalise them into
the model’s input layout:
const canvas = new OffscreenCanvas(256, 256);
const ctx = canvas.getContext("2d", { willReadFrequently: true });
self.onmessage = async ({ data: { bitmap } }) => {
ctx.drawImage(bitmap, 0, 0, 256, 256); bitmap.close();
const rgba = ctx.getImageData(16, 16, 224, 224).data; // centre crop
pre.rgba_to_nchw(rgba, inputTensorView, MEAN, STD); // Wasm: uint8 RGBA → float32 NCHW
const out = await session.run({ input: inputTensor });
self.postMessage({ top: topK(out.logits.data, 3) });
};
A small Rust or C function converting RGBA bytes to normalised planar floats with SIMD is several times faster than an equivalent JavaScript loop, and it writes directly into the buffer the runtime reads. Check the model’s expected layout (NCHW or NHWC), channel order (RGB or BGR) and normalisation constants; a mismatch gives confidently wrong predictions.
Step 3 — run inference on the Wasm backend
import * as ort from "onnxruntime-web";
ort.env.wasm.numThreads = crossOriginIsolated ? 4 : 1;
const session = await ort.InferenceSession.create("/models/mobilenetv3-small-q8.onnx", { executionProviders: ["wasm"] });
const inputTensor = new ort.Tensor("float32", new Float32Array(3 * 224 * 224), [1, 3, 224, 224]);
Create the session once. Reuse the input tensor’s buffer for every frame to avoid allocations. Quantised (int8) models run faster on the Wasm backend; if WebGPU is available, its backend may be faster still, but for small models the Wasm backend is often competitive once data transfer costs are included.
Step 4 — smooth predictions
Frame-by-frame predictions flicker as the camera moves. Average class probabilities over the last few results (an exponential moving average works well), require a label to stay on top for a short time before announcing it, and hide predictions below a confidence threshold. This stabilises the UI without adding noticeable lag.
Step 5 — measure end-to-end latency and rate
Measure classifications per second and the delay from a frame’s capture time (requestVideoFrameCallback provides metadata including mediaTime) to the
displayed label. On phones, reduce input resolution, use the smallest model that meets accuracy needs, and cap the analysis rate (for example 5 per second)
to save battery and avoid thermal throttling, which otherwise slows everything after a minute or two.
Privacy and permissions
Camera access requires a permission prompt and HTTPS. Processing frames locally means images never leave the device; say so, and avoid sending frames or
thumbnails to analytics. Stop the camera tracks (stream.getTracks().forEach(t => t.stop())) when the feature is closed, so the camera indicator turns off.
Regions of interest instead of whole frames
Classifying the whole frame works when the subject fills the view; often it does not. A document scanner cares about the paper, a product recogniser about the item in the user’s hand. Cropping to a region of interest before classification improves accuracy far more than a larger model would. The region can be fixed (a guide rectangle drawn over the preview, which users align the subject with), derived from a cheap detector run every few frames, or tracked between detections. The crop happens in the same Wasm preprocessing step, so it costs nothing extra: pass the rectangle along with the frame, and resize only that region to the model’s input size. Showing the guide rectangle in the UI also teaches users how to hold the camera, which improves results more than any amount of post-processing.
Adapting to the device at runtime
Devices vary by an order of magnitude in inference speed. Measure the first few inference times after start-up and adapt: if results take longer than the target interval, lower the analysis rate, switch to a smaller input resolution, or load a smaller model variant; if the device is fast, raise the rate up to a cap. Watch for slowdowns over time as well — phones throttle when warm — and back off before the UI starts to lag. An adaptive loop delivers a consistent experience across a cheap phone and a fast laptop without separate builds.
Testing without a camera
Camera-dependent features are awkward to test. Feed recorded video instead: a <video> element playing a test clip exposes the same
requestVideoFrameCallback API, so the whole pipeline runs against known footage. Browsers’ automation tools can also supply fake camera streams (Chromium’s
--use-file-for-fake-video-capture flag), which lets end-to-end tests exercise permission handling and capture together. Assert on the labels produced for
key moments in the clip, with tolerance for timing.
Expected output
The preview runs at 30 fps while classification runs at about 40 results per second on a laptop and 8 per second on a mid-range phone; labels update smoothly with moving-average filtering; preprocessing in Wasm takes about 1 ms per frame; frames are never queued; and closing the feature stops the camera.
Gotchas
- Queuing every frame. Latency grows without bound. Drop frames while busy.
- Wrong tensor layout or normalisation. Confident wrong answers. Match the model’s preprocessing.
- Inference on the main thread. The preview stutters. Use a worker.
- Allocating tensors per frame. GC pauses. Reuse buffers.
- Leaving the camera on. Users notice the indicator. Stop tracks on close.
- Fixed analysis rate on every device. Slow phones lag and heat up. Adapt the rate to measured inference time.
Performance note
Moving preprocessing from a JavaScript loop to a SIMD Wasm function cut it from about 6 ms to 1 ms per frame, raising classifications per second on a laptop from about 30 to about 40.
Frequently Asked Questions
Can I use MediaPipe or TensorFlow.js instead? Yes — both offer browser runtimes with Wasm backends; the frame-handling patterns are the same.
Does VideoFrame from WebCodecs help?
MediaStreamTrackProcessor (where supported) gives VideoFrames directly, avoiding the video element; copyTo can write pixels into a buffer.
How do I handle front versus back cameras?
Use facingMode constraints and mirror the preview for the front camera only.
Should models run on the GPU? For larger models, WebGPU often wins; for tiny classifiers, Wasm is competitive and more widely available.
How can I improve accuracy without a bigger model? Crop to a region of interest — a guide rectangle or a detector’s box — before classification.
How do I test the pipeline without a camera? Play a recorded clip in a video element, or use the browser’s fake camera capture in automated tests.
What should happen when the phone gets warm? Detect rising inference times and lower the analysis rate or resolution before the preview starts to stutter.
Should the guide rectangle be shown to users? Yes — it tells users where to place the subject and makes the crop match what the model expects.
Related
- Running ONNX models with onnxruntime-web — the runtime.
- Feeding WebCodecs frames into Wasm — frame access.
- Measuring inference latency in the browser — measurement.
- Passing canvas pixels to Wasm without extra copies — pixel handling.