Removing Image Backgrounds in the Browser

This page answers one task: users upload product photos, profile pictures or graphics and want the background removed — and you want to do it on their device, for privacy and to avoid paying for a server-side API, with edges clean enough for real use.

Prerequisites

  • [ ] A salient-object or person segmentation model in ONNX format (U²-Net-, MODNet- or similar class), sized for the browser.
  • [ ] ONNX Runtime Web (Wasm backend, optionally WebGPU).
  • [ ] A small Wasm module (Rust or C) for pixel post-processing.

How background removal works

Background removal has two stages. A segmentation model looks at a downscaled version of the image (commonly 320×320 or 512×512) and outputs a soft mask: for each pixel, the probability that it belongs to the foreground. Post-processing turns that low-resolution mask into a full-resolution alpha channel: upsample it to the original size, refine edges using the full-resolution image (so hair and fine detail are not blocky), feather the boundary, and remove the background colour that bleeds into edge pixels. The model decides what is foreground; post-processing decides how good the edges look — and much of the perceived quality comes from the second stage.

Running the model in ONNX Runtime Web and the pixel work in a custom Wasm module keeps both on the device and off the main thread.

Background removal pipeline The full-size image is downscaled and normalised into the model's input tensor. The segmentation model produces a low-resolution soft mask. Wasm upsamples the mask to full size with edge-aware refinement guided by the original image, feathers the boundary, decontaminates edge colours, and writes the alpha channel, producing a transparent PNG or WebP. downscale + normalise e.g. 320×320 tensor segmentation model soft mask upsample + refine (Wasm) edge-aware feather + decontaminate clean edges RGBA output PNG / WebP

Step 1 — choose a model for the subject type

General salient-object models handle products and arbitrary subjects; portrait models (person segmentation, matting) produce better hair and edges for people. Model sizes range from a few megabytes (fast, coarser) to over 150 MB (slow in the browser). Quantised variants reduce size and speed up the Wasm backend; check quality on your own images, since edges degrade first. Licences vary — confirm commercial use is allowed.

Step 2 — preprocess and run the model

const session = await ort.InferenceSession.create("/models/seg-small-q8.onnx", { executionProviders: ["wasm"] });

async function predictMask(bitmap) {
  const S = 320;
  const ctx = new OffscreenCanvas(S, S).getContext("2d", { willReadFrequently: true });
  ctx.drawImage(bitmap, 0, 0, S, S);
  const rgba = ctx.getImageData(0, 0, S, S).data;
  const input = post.to_nchw_normalized(rgba, S, S, MEAN, STD);   // Wasm: Float32Array [1,3,S,S]
  const out = await session.run({ input: new ort.Tensor("float32", input, [1, 3, S, S]) });
  return out[session.outputNames[0]].data;                         // Float32Array S×S in [0,1]
}

Match the model’s expected input size, channel order and normalisation exactly; mismatches produce masks that look plausible but cut off parts of the subject.

Step 3 — upsample and refine the mask in Wasm

Bilinear upsampling of a 320×320 mask to a 4000×3000 photo produces soft, blurry edges. Edge-aware upsampling — a guided filter using the full-resolution image as the guide — snaps mask edges to real image edges and recovers fine detail. Implemented in Rust or C over the image in linear memory, a guided filter is a handful of box filters, which run quickly with SIMD:

#[wasm_bindgen]
pub fn refine_alpha(rgba: &mut [u8], w: usize, h: usize, mask: &[f32], mw: usize, mh: usize, radius: usize, eps: f32) {
    let coarse = upsample_bilinear(mask, mw, mh, w, h);       // initial full-size alpha
    let guide = luminance(rgba, w, h);                          // single-channel guide
    let alpha = guided_filter(&guide, &coarse, w, h, radius, eps);
    for (px, a) in rgba.chunks_exact_mut(4).zip(alpha) {
        px[3] = (a.clamp(0.0, 1.0) * 255.0) as u8;
    }
}

Operate on the image in place in linear memory (copied in once with a typed-array set), so there are no extra copies of a 12-megapixel image.

Plain upsampling versus edge-aware refinement Bilinear upsampling of the low-resolution mask produces blurry, haloed edges that ignore image detail. Edge-aware refinement with a guided filter uses the full-resolution image to align mask edges with real edges and recover detail such as hair, at a modest extra cost in Wasm. bilinear upsample soft, blocky edges halos of background very cheap previews guided-filter refinement edges follow image detail better hair and fur tens of ms in Wasm final output

Step 4 — clean up edge colours

Pixels along the boundary mix foreground and background colours; with the background removed, they show a fringe of the old background (a green halo from a green wall). Colour decontamination estimates the foreground colour of partially transparent pixels — for example by extending colours from confidently opaque neighbours — and replaces the mixed colour. Light feathering (a small blur of the alpha at the edge) avoids jagged boundaries. Both are cheap per-pixel operations in Wasm.

Step 5 — export and offer adjustments

Export as PNG or WebP with alpha (OffscreenCanvas.convertToBlob({ type: "image/webp" }) or a Wasm encoder for consistent results). Let users adjust the result: a slider for edge softness, a brush to add or remove areas (which edits the mask and re-runs refinement locally), and a preview on checkerboard and on a chosen background colour. Automatic segmentation gets most images right; manual touch-up handles the rest.

Performance and memory

A small quantised model runs in tens to a few hundred milliseconds on the Wasm backend depending on device; refinement of a 12-megapixel image takes tens of milliseconds. Memory is dominated by the full-resolution image and its float intermediates (a single f32 channel at 12 MP is 48 MB); process very large images at a capped working resolution (for example 4 MP) and upscale the final alpha, or process in tiles.

Batch processing product catalogues

E-commerce users often need hundreds of product photos cut out at once. Treat it as a queue in a worker: load the model once, process images one at a time (each needs its full-resolution buffer, so parallelism is limited by memory more than CPU), and stream results — thumbnail previews first, full-resolution exports when done. Let users review results in a grid and flag the few that need manual touch-up, rather than reviewing each image in a modal. Consistency matters in catalogues: apply the same edge softness and the same output size and padding to every image, and optionally place each cut-out on a uniform background or shadow, all in the same Wasm post-processing step. A pause and resume control is worth adding, because batch jobs compete with everything else the user does in the browser, and a laptop on battery may need to stop halfway.

Evaluating quality

Judging background removal by eye on a few examples is misleading; models vary by subject type. Assemble a small evaluation set from your users’ typical images — products on cluttered backgrounds, people with hair, transparent objects like glasses, white objects on white — with hand-made reference masks, and measure mask quality (for example intersection over union, and an edge-focused error measure) for each candidate model and post-processing setting. Pay particular attention to failure modes users notice most: missing parts of the subject, leftover background patches and fringes. Re-run the set whenever you change the model, quantisation or refinement parameters.

When to fall back to a server

Some images defeat small browser models — complex scenes, unusual subjects, very high resolutions on weak devices. Detect low-confidence masks (large areas with probabilities near 0.5) and offer an optional server-side pass with a larger model, clearly labelled, for users who accept uploading that image.

Expected output

A 12-megapixel product photo loses its background in about 600 ms on a laptop (380 ms inference, 150 ms refinement, the rest I/O), with clean edges around fine details, no colour fringe, and an exported transparent WebP; users can touch up the mask with a brush; nothing leaves the device.

Gotchas

  • Bilinear-only upsampling. Blurry halos. Refine edges with the full-resolution image.
  • Wrong normalisation or channel order. Masks miss parts of the subject. Match the model.
  • Ignoring colour fringes. Halos of the old background. Decontaminate edge colours.
  • Float intermediates at full resolution. Memory spikes. Cap working size or tile.
  • Main-thread processing. The UI freezes. Use a worker.
  • Judging quality on a few demo images. Models vary by subject. Evaluate on representative images with reference masks.

Performance note

On a laptop, inference with a quantised small segmentation model took 380 ms; guided-filter refinement of a 12-megapixel image took 150 ms in Wasm with SIMD versus about 900 ms for an equivalent JavaScript implementation.

Mask refinement time for a 12-megapixel photo Milliseconds to refine a low-resolution mask into a full-resolution alpha channel with a guided filter in JavaScript and in Wasm with SIMD. ms per image JavaScript guided filter 900 ms Wasm + SIMD guided filter 150 ms

Frequently Asked Questions

Is WebGPU faster for the model? Often, for larger models; small quantised models run acceptably on the Wasm backend everywhere.

Can it handle hair well? Portrait matting models plus edge-aware refinement do; general models struggle with fine hair.

What about video? Per-frame segmentation is possible with small models and temporal smoothing, at reduced resolution.

Do I need a separate Wasm module for post-processing? It is the fastest option; canvas filters cannot express guided filtering efficiently.

How should hundreds of images be processed? As a queue in one worker with the model loaded once, one image at a time, streaming previews and allowing pause and resume.

How can low-quality results be detected automatically? Look for large mask areas with probabilities near 0.5 and offer touch-up or an optional server pass for those images.

Can the cut-out be placed on a new background automatically? Yes — compose it onto a colour, image or soft shadow in the same Wasm post-processing step before export.

← Back to Media Processing & Codecs in Wasm