Removing Image Backgrounds in the Browser
This page answers one task: users upload product photos, profile pictures or graphics and want the background removed — and you want to do it on their device, for privacy and to avoid paying for a server-side API, with edges clean enough for real use.
Prerequisites
- [ ] A salient-object or person segmentation model in ONNX format (U²-Net-, MODNet- or similar class), sized for the browser.
- [ ] ONNX Runtime Web (Wasm backend, optionally WebGPU).
- [ ] A small Wasm module (Rust or C) for pixel post-processing.
How background removal works
Background removal has two stages. A segmentation model looks at a downscaled version of the image (commonly 320×320 or 512×512) and outputs a soft mask: for each pixel, the probability that it belongs to the foreground. Post-processing turns that low-resolution mask into a full-resolution alpha channel: upsample it to the original size, refine edges using the full-resolution image (so hair and fine detail are not blocky), feather the boundary, and remove the background colour that bleeds into edge pixels. The model decides what is foreground; post-processing decides how good the edges look — and much of the perceived quality comes from the second stage.
Running the model in ONNX Runtime Web and the pixel work in a custom Wasm module keeps both on the device and off the main thread.
Step 1 — choose a model for the subject type
General salient-object models handle products and arbitrary subjects; portrait models (person segmentation, matting) produce better hair and edges for people. Model sizes range from a few megabytes (fast, coarser) to over 150 MB (slow in the browser). Quantised variants reduce size and speed up the Wasm backend; check quality on your own images, since edges degrade first. Licences vary — confirm commercial use is allowed.
Step 2 — preprocess and run the model
const session = await ort.InferenceSession.create("/models/seg-small-q8.onnx", { executionProviders: ["wasm"] });
async function predictMask(bitmap) {
const S = 320;
const ctx = new OffscreenCanvas(S, S).getContext("2d", { willReadFrequently: true });
ctx.drawImage(bitmap, 0, 0, S, S);
const rgba = ctx.getImageData(0, 0, S, S).data;
const input = post.to_nchw_normalized(rgba, S, S, MEAN, STD); // Wasm: Float32Array [1,3,S,S]
const out = await session.run({ input: new ort.Tensor("float32", input, [1, 3, S, S]) });
return out[session.outputNames[0]].data; // Float32Array S×S in [0,1]
}
Match the model’s expected input size, channel order and normalisation exactly; mismatches produce masks that look plausible but cut off parts of the subject.
Step 3 — upsample and refine the mask in Wasm
Bilinear upsampling of a 320×320 mask to a 4000×3000 photo produces soft, blurry edges. Edge-aware upsampling — a guided filter using the full-resolution image as the guide — snaps mask edges to real image edges and recovers fine detail. Implemented in Rust or C over the image in linear memory, a guided filter is a handful of box filters, which run quickly with SIMD:
#[wasm_bindgen]
pub fn refine_alpha(rgba: &mut [u8], w: usize, h: usize, mask: &[f32], mw: usize, mh: usize, radius: usize, eps: f32) {
let coarse = upsample_bilinear(mask, mw, mh, w, h); // initial full-size alpha
let guide = luminance(rgba, w, h); // single-channel guide
let alpha = guided_filter(&guide, &coarse, w, h, radius, eps);
for (px, a) in rgba.chunks_exact_mut(4).zip(alpha) {
px[3] = (a.clamp(0.0, 1.0) * 255.0) as u8;
}
}
Operate on the image in place in linear memory (copied in once with a typed-array set), so there are no extra copies of a 12-megapixel image.
Step 4 — clean up edge colours
Pixels along the boundary mix foreground and background colours; with the background removed, they show a fringe of the old background (a green halo from a green wall). Colour decontamination estimates the foreground colour of partially transparent pixels — for example by extending colours from confidently opaque neighbours — and replaces the mixed colour. Light feathering (a small blur of the alpha at the edge) avoids jagged boundaries. Both are cheap per-pixel operations in Wasm.
Step 5 — export and offer adjustments
Export as PNG or WebP with alpha (OffscreenCanvas.convertToBlob({ type: "image/webp" }) or a Wasm encoder for consistent results). Let users adjust the
result: a slider for edge softness, a brush to add or remove areas (which edits the mask and re-runs refinement locally), and a preview on checkerboard and on
a chosen background colour. Automatic segmentation gets most images right; manual touch-up handles the rest.
Performance and memory
A small quantised model runs in tens to a few hundred milliseconds on the Wasm backend depending on device; refinement of a 12-megapixel image takes tens of
milliseconds. Memory is dominated by the full-resolution image and its float intermediates (a single f32 channel at 12 MP is 48 MB); process very large images
at a capped working resolution (for example 4 MP) and upscale the final alpha, or process in tiles.
Batch processing product catalogues
E-commerce users often need hundreds of product photos cut out at once. Treat it as a queue in a worker: load the model once, process images one at a time (each needs its full-resolution buffer, so parallelism is limited by memory more than CPU), and stream results — thumbnail previews first, full-resolution exports when done. Let users review results in a grid and flag the few that need manual touch-up, rather than reviewing each image in a modal. Consistency matters in catalogues: apply the same edge softness and the same output size and padding to every image, and optionally place each cut-out on a uniform background or shadow, all in the same Wasm post-processing step. A pause and resume control is worth adding, because batch jobs compete with everything else the user does in the browser, and a laptop on battery may need to stop halfway.
Evaluating quality
Judging background removal by eye on a few examples is misleading; models vary by subject type. Assemble a small evaluation set from your users’ typical images — products on cluttered backgrounds, people with hair, transparent objects like glasses, white objects on white — with hand-made reference masks, and measure mask quality (for example intersection over union, and an edge-focused error measure) for each candidate model and post-processing setting. Pay particular attention to failure modes users notice most: missing parts of the subject, leftover background patches and fringes. Re-run the set whenever you change the model, quantisation or refinement parameters.
When to fall back to a server
Some images defeat small browser models — complex scenes, unusual subjects, very high resolutions on weak devices. Detect low-confidence masks (large areas with probabilities near 0.5) and offer an optional server-side pass with a larger model, clearly labelled, for users who accept uploading that image.
Expected output
A 12-megapixel product photo loses its background in about 600 ms on a laptop (380 ms inference, 150 ms refinement, the rest I/O), with clean edges around fine details, no colour fringe, and an exported transparent WebP; users can touch up the mask with a brush; nothing leaves the device.
Gotchas
- Bilinear-only upsampling. Blurry halos. Refine edges with the full-resolution image.
- Wrong normalisation or channel order. Masks miss parts of the subject. Match the model.
- Ignoring colour fringes. Halos of the old background. Decontaminate edge colours.
- Float intermediates at full resolution. Memory spikes. Cap working size or tile.
- Main-thread processing. The UI freezes. Use a worker.
- Judging quality on a few demo images. Models vary by subject. Evaluate on representative images with reference masks.
Performance note
On a laptop, inference with a quantised small segmentation model took 380 ms; guided-filter refinement of a 12-megapixel image took 150 ms in Wasm with SIMD versus about 900 ms for an equivalent JavaScript implementation.
Frequently Asked Questions
Is WebGPU faster for the model? Often, for larger models; small quantised models run acceptably on the Wasm backend everywhere.
Can it handle hair well? Portrait matting models plus edge-aware refinement do; general models struggle with fine hair.
What about video? Per-frame segmentation is possible with small models and temporal smoothing, at reduced resolution.
Do I need a separate Wasm module for post-processing? It is the fastest option; canvas filters cannot express guided filtering efficiently.
How should hundreds of images be processed? As a queue in one worker with the model loaded once, one image at a time, streaming previews and allowing pause and resume.
How can low-quality results be detected automatically? Look for large mask areas with probabilities near 0.5 and offer touch-up or an optional server pass for those images.
Can the cut-out be placed on a new background automatically? Yes — compose it onto a colour, image or soft shadow in the same Wasm post-processing step before export.
Related
- Running ONNX models with onnxruntime-web — the runtime.
- Building a Wasm image filter pipeline — pixel processing.
- Compressing images before upload with Wasm — export.
- Running image classification on webcam frames — preprocessing patterns.
← Back to Media Processing & Codecs in Wasm