Running ONNX Models with onnxruntime-web
This guide answers one task: load an ONNX model in a browser with onnxruntime-web, feed it a correctly
shaped tensor, and read the result — with the runtime files served from your own origin so threads and
SIMD actually engage.
Prerequisites
- [ ]
onnxruntime-web1.19 or later from npm. - [ ] An
.onnxmodel with known input and output names — check them withonnx.checkeror Netron. - [ ] A build step that can copy files into your static output directory.
- [ ] Optional but recommended: COOP/COEP headers so the threaded backend is available.
Install and self-host the runtime files
onnxruntime-web ships its execution engine as separate .wasm and .mjs files. The library fetches
them at runtime from a path you control, and getting that path wrong is the single most common setup
failure. Copy them into your public directory as part of the build rather than relying on a CDN.
npm i onnxruntime-web
# copy the runtime artifacts next to your app's static assets
mkdir -p public/vendor/ort
cp node_modules/onnxruntime-web/dist/*.wasm public/vendor/ort/
cp node_modules/onnxruntime-web/dist/*.mjs public/vendor/ort/
Then tell the library where they live, before creating any session:
import * as ort from 'onnxruntime-web';
ort.env.wasm.wasmPaths = '/vendor/ort/';
ort.env.wasm.simd = true;
ort.env.wasm.numThreads = crossOriginIsolated ? Math.min(4, navigator.hardwareConcurrency || 1) : 1;
Self-hosting also removes the cross-origin resource policy problem: with require-corp in force, a
cross-origin .wasm without the matching Cross-Origin-Resource-Policy header simply fails to load, and
the error points at the wrong thing.
Create a session
Session creation parses the graph, allocates tensors and applies graph optimisations. Do it once, keep the result, and never create a session per inference.
const session = await ort.InferenceSession.create('/models/mobilenet-int8.onnx', {
executionProviders: ['wasm'],
graphOptimizationLevel: 'all',
});
console.log(session.inputNames, session.outputNames);
// [ 'pixel_values' ] [ 'logits' ]
Log the input and output names the first time you integrate a model. Guessing them is how you end up
with Error: invalid input 'input', and the names are model-specific: exporters variously produce
input, input.1, images or pixel_values for what is conceptually the same tensor.
Build the input tensor
The tensor is a flat typed array plus a shape. For a vision model the layout is almost always NCHW —
batch, channel, height, width — which means the three colour channels are stored as three contiguous
planes rather than interleaved per pixel as they are in ImageData.
function toNCHW(imageData, mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225]) {
const { data, width: w, height: h } = imageData;
const out = new Float32Array(3 * w * h);
const plane = w * h;
for (let i = 0, p = 0; i < data.length; i += 4, p++) {
out[p] = (data[i] / 255 - mean[0]) / std[0]; // R plane
out[p + plane] = (data[i + 1] / 255 - mean[1]) / std[1]; // G plane
out[p + 2 * plane] = (data[i + 2] / 255 - mean[2]) / std[2]; // B plane
}
return new ort.Tensor('float32', out, [1, 3, h, w]);
}
Two things are easy to get wrong here and hard to notice. The mean and standard deviation must match the ones used in training — the values above are the ImageNet defaults, and a model trained differently will produce confident nonsense with them. And the alpha channel is skipped entirely; including it shifts every subsequent plane and produces an image the model sees as scrambled.
Run and read the output
run() takes an object keyed by input name and resolves to an object keyed by output name. The result’s
data is a typed array in the output’s declared dtype.
const feeds = { [session.inputNames[0]]: tensor };
const results = await session.run(feeds);
const logits = results[session.outputNames[0]].data; // Float32Array(1000)
const top = [...logits]
.map((score, index) => ({ index, score }))
.sort((a, b) => b.score - a.score)
.slice(0, 5);
If the model emits logits rather than probabilities — most do — apply softmax yourself before showing numbers to a user. Reporting raw logits as confidence produces values above one and below zero, which looks like a bug to everyone who sees it.
Expected output
A successful first run prints the shapes you expect and a plausible ranking:
inputs [ 'pixel_values' ]
outputs [ 'logits' ]
shape [ 1, 1000 ]
top5 [ { index: 281, score: 9.41 }, // tabby cat
{ index: 285, score: 8.77 },
{ index: 282, score: 7.12 }, ... ]
The sanity check that catches preprocessing mistakes is running a fixture image whose expected class you know, and comparing against the same model outside the browser. If the browser’s top class differs from Python’s, the problem is upstream of the runtime — resize filter, normalisation or channel order — essentially every time.
Loading the model file yourself
InferenceSession.create() accepts a URL, but it also accepts an ArrayBuffer or a Uint8Array. Taking
the fetch into your own hands is worth it as soon as the model is more than a few megabytes, because it
gives you three things the URL form cannot: a progress indicator, control over caching, and the ability
to load from somewhere other than the network.
async function loadModelBytes(url, onProgress) {
const cache = await caches.open('models-v1');
let res = await cache.match(url);
if (!res) {
res = await fetch(url);
await cache.put(url, res.clone()); // keep it for next time
}
const total = Number(res.headers.get('content-length')) || 0;
const reader = res.body.getReader();
const chunks = [];
let seen = 0;
for (;;) {
const { value, done } = await reader.read();
if (done) break;
chunks.push(value);
seen += value.length;
if (total) onProgress(seen / total);
}
const bytes = new Uint8Array(seen);
let off = 0;
for (const c of chunks) { bytes.set(c, off); off += c.length; }
return bytes;
}
const session = await ort.InferenceSession.create(await loadModelBytes(MODEL_URL, setPct), OPTIONS);
The Cache Storage API is the right place for model files rather than plain HTTP caching, because you
decide when an entry is evicted and you can version the cache name when the model changes. Note that the
assembled Uint8Array is a full copy of the model in memory, which is briefly doubled while the session
parses it — the reason very large models need the streaming approach described in
loading large model weights.
Handling models with more than one input
Vision classifiers take one tensor; almost everything else does not. Text models want input_ids and
attention_mask; encoder-decoder models add decoder_input_ids; models with past-key-value caching want
a tensor per layer. The feeds object simply grows, and the shapes must be exact.
const feeds = {
input_ids: new ort.Tensor('int64', BigInt64Array.from(ids.map(BigInt)), [1, ids.length]),
attention_mask: new ort.Tensor('int64', BigInt64Array.from(ids.map(() => 1n)), [1, ids.length]),
};
const { last_hidden_state } = await session.run(feeds);
Two details trip people up here. Token identifiers are int64 in most exported models, which means
BigInt64Array and BigInt values — passing a plain Int32Array throws a type error that names the
tensor but not the reason. And the sequence length appears in several shapes at once; if the mask and the
identifiers disagree by even one element, the runtime reports a broadcast failure deep inside an operator
rather than at the boundary where you can see it.
Warm up before you measure anything
The first run() after session creation allocates arenas, selects kernels and touches cold pages. It is
routinely five to ten times slower than the steady state, and benchmarking without discarding it produces
numbers that are simply wrong.
const dummy = new ort.Tensor('float32', new Float32Array(3 * 224 * 224), [1, 3, 224, 224]);
await session.run({ [session.inputNames[0]]: dummy }); // discard
Do the warm-up during a moment the user is not waiting — right after session creation, while they are still choosing an input. It also surfaces shape and name errors immediately rather than on the first real interaction.
Gotchas
no available backend found. ERR: [wasm] ...—wasmPathsis wrong, or the variant the runtime wants was not copied. Check the network panel for a 404 on a.wasmor.mjsunder your vendor path.Error: invalid input 'x'— the feed key does not matchsession.inputNames[0]. Never hard-code the name; read it from the session.- Threads silently disabled.
crossOriginIsolatedisfalse. The runtime falls back to one thread and says nothing, and inference is three times slower than your local measurements. Cannot read properties of undefined (reading 'data')— you indexed the results object with the wrong output name. Logsession.outputNamesonce.- Bundler rewrites the worker URL. Some bundlers try to process the runtime’s
.mjsworker file. Mark the vendor directory as an external static asset rather than letting the bundler touch it — the same class of problem described in bundling Wasm ESM with Vite.
Performance note
A quantized MobileNet-class model at 224 × 224 runs in roughly 12–25 ms on a laptop with SIMD and four threads, and 45–120 ms on a mid-range phone. Session creation for the same model is 150–400 ms, which is why it must not be repeated. The preprocessing loop above costs about 1.5 ms for a 224 × 224 image in JavaScript — worth moving into the module only if you are running it per video frame, in which case the copy and the conversion fuse into the same pass.
Frequently Asked Questions
Does this work in a Web Worker?
Yes, and it is the recommended deployment. Everything in this page runs unchanged in a worker; only the
image acquisition has to move, using createImageBitmap and OffscreenCanvas rather than a DOM canvas.
How do I run several inputs at once? Batch them into one tensor with a first dimension greater than one, if the model’s graph allows it. Batching four images typically costs far less than four separate runs, because the GEMM kernels get larger matrices to work with.
Can I use a model that expects dynamic shapes? Yes — pass the actual shape in the tensor and the runtime resolves it per run. Expect a slower first run for each new shape, since kernel selection happens per resolved shape, so a page that feeds arbitrary sizes will warm up repeatedly. Padding to a small set of fixed sizes is usually faster overall.
Related
- Measuring inference latency in the browser — how to get numbers worth quoting.
- Quantizing models for Wasm inference — making the file smaller before you ship it.
- Multi-threaded inference with Wasm threads — turning on the biggest speedup available.