Computing Text Embeddings Locally with Wasm

This page answers one task: an application needs semantic search, deduplication or clustering over the user’s own text — notes, messages, bookmarks — and the text should not be sent to an embedding API. You want a small embedding model running in the browser through WebAssembly, producing vectors fast enough to index thousands of items.

Prerequisites

  • [ ] A sentence-embedding model exported to ONNX (for example a MiniLM- or BGE-small-class model), ideally quantised.
  • [ ] ONNX Runtime Web or Transformers.js (which uses ONNX Runtime Web) in your build.
  • [ ] A worker in which to run the model.

What an embedding model does

A sentence-embedding model maps text to a fixed-length vector — commonly 384 or 768 dimensions — such that texts with similar meaning have vectors with high cosine similarity. Computing one involves tokenising the text into subword IDs, running a transformer encoder over them, pooling the per-token outputs into one vector (usually the mean over tokens, weighted by the attention mask), and normalising it to unit length. The model is far smaller than a generative language model — small embedding models are 20–130 MB at full precision and a quarter of that quantised — and a forward pass over a sentence takes milliseconds on a laptop CPU with Wasm SIMD.

ONNX Runtime Web runs ONNX models in the browser with a WebAssembly backend (SIMD, optionally threads) and a WebGPU backend. Transformers.js wraps it with tokenisers and pipelines that handle tokenisation, pooling and normalisation for you.

From text to a normalised embedding The text is tokenised into subword IDs with an attention mask. The transformer encoder runs on the Wasm backend and produces one vector per token. Mean pooling weighted by the mask yields one vector for the text, which is normalised to unit length so cosine similarity becomes a dot product. text "local-first notes" tokenise IDs + attention mask encoder (ONNX, Wasm) per-token vectors mean pooling mask-weighted L2 normalise 384-d unit vector

Step 1 — choose a model

Pick a small model trained for sentence embeddings, with good scores on retrieval benchmarks for your language. 384-dimensional models are a good default: smaller vectors mean smaller indexes and faster search. Multilingual models are larger; use them only if you need them. Check the model card for the expected pooling and whether queries need a prefix (some models expect query: and passage: prefixes, and results degrade without them).

Step 2 — run it with Transformers.js on the Wasm backend

// embed-worker.js
import { pipeline, env } from "@huggingface/transformers";

env.allowRemoteModels = false;                 // serve model files from your own origin
env.localModelPath = "/models/";
env.backends.onnx.wasm.numThreads = navigator.hardwareConcurrency > 4 ? 4 : 1;

const extractor = await pipeline("feature-extraction", "all-MiniLM-L6-v2", { dtype: "q8", device: "wasm" });

self.onmessage = async ({ data: { id, texts } }) => {
  const output = await extractor(texts, { pooling: "mean", normalize: true });
  self.postMessage({ id, vectors: output.data, dims: output.dims }, [output.data.buffer]);
};

The pipeline tokenises, runs the model, pools and normalises. Serving model files from your own origin avoids third-party dependencies at runtime and lets you cache them with your headers. numThreads above 1 requires cross-origin isolation.

Step 3 — batch for throughput

Embedding one text at a time wastes work: each call has overhead, and the encoder runs more efficiently on batches. Send batches of 16–64 texts of similar length (padding is per batch, so mixing very short and very long texts wastes compute on padding). For indexing thousands of notes, process batches in the worker continuously and report progress; for interactive queries, embed the single query immediately.

Step 4 — check results against a reference

Verify that browser embeddings match a reference implementation — the same model in Python with sentence-transformers — on a handful of sentences. Cosine similarity between the browser vector and the reference vector should be above 0.99 for the full-precision model and slightly lower for quantised variants. A large mismatch usually means wrong pooling (CLS instead of mean, or no mask weighting), missing normalisation, a missing query prefix, or a tokeniser version mismatch.

Wasm backend versus WebGPU backend for embeddings The Wasm backend runs everywhere with SIMD and optional threads and is fast enough for single queries and modest indexing. The WebGPU backend is much faster for large batches but needs WebGPU support and has a higher startup cost, so it pays off when indexing many thousands of texts. Wasm backend works in every browser ms per short text fine for queries default WebGPU backend needs WebGPU fast large batches higher startup cost bulk indexing

Step 5 — cache the model and the vectors

Cache model files (Cache API or OPFS) so they download once. Cache computed vectors too — store them with the item they belong to, together with the model name and version, so you only embed new or changed items. When you change models, all vectors must be recomputed, because vectors from different models are not comparable.

Long texts and chunking

Embedding models have a maximum input length (often 256 or 512 tokens); longer input is truncated. For long documents, split into chunks of a few hundred tokens with some overlap, embed each chunk, and either index chunks separately (better for finding specific passages) or average them into a document vector (simpler, coarser). Search then returns the best-matching chunks, which also makes it easy to highlight why a document matched.

Performance and memory

A quantised MiniLM-class model uses tens of megabytes of memory plus activations proportional to batch size and sequence length. On a laptop, single short texts embed in a few milliseconds; batches of 32 sentences take tens of milliseconds. On phones, expect several times slower — index in the background, in small batches, and pause when the page is hidden to save battery.

Incremental indexing as users edit

Text changes constantly in a notes app, and re-embedding everything after each edit is wasteful. Track a content hash per item (or per chunk) alongside its vector; when an item is saved, compare the new hash with the stored one and re-embed only if it changed. Debounce re-embedding during active editing — wait until the user pauses for a few seconds — so a long typing session produces one embedding rather than dozens. For chunked documents, re-embed only the chunks whose text changed, which keeps the cost proportional to the edit rather than the document. Queue embedding jobs in the worker with priorities: the user’s search query first, recently edited items next, background backfill last. That keeps search responsive even while a large initial index is still being built.

Evaluating search quality

Fast embeddings are only useful if search returns what users expect. Build a small evaluation set from real usage — twenty or thirty queries with the items a person would consider correct answers — and measure how often the right item appears in the top results (recall at 5 or 10). Use it to compare models, quantisation levels, chunk sizes and query prefixes before shipping a change; differences that are invisible in a demo often show up clearly in such a set. Re-run it whenever you change the model, since a “better” model on public benchmarks is not always better on a particular user’s notes.

Privacy and model loading

Serving model files from your own origin, as in step 2, also keeps the privacy story simple: nothing about what users embed or search for reaches a third party, and a Content Security Policy that blocks other origins makes that verifiable.

Expected output

A notes app embeds 5,000 notes in the background in about 40 seconds on a laptop with a quantised 384-dimensional model in a worker; each search query embeds in 5 ms; vectors match the Python reference with cosine similarity above 0.99; model files and vectors are cached; and no text is sent to any server.

Gotchas

  • Wrong pooling or no normalisation. Similarities are off. Match the model card.
  • Missing query/passage prefixes. Retrieval quality drops. Apply them where required.
  • Embedding on the main thread. The UI stalls during indexing. Use a worker.
  • Mixing vectors from different models. Meaningless similarities. Store the model version.
  • Truncated long documents. Content is lost silently. Chunk long texts.
  • Re-embedding on every keystroke. Wasted work. Hash content and debounce.

Performance note

Embedding batches of 32 short sentences took about 60 ms per batch with 4 threads on a laptop and about 190 ms single-threaded; a single query took about 5 ms.

Time to embed a batch of 32 sentences Milliseconds to embed 32 short sentences with a quantised MiniLM-class model on the ONNX Runtime Web Wasm backend with one thread and with four threads on a laptop. ms per batch of 32 Wasm, 1 thread 190 ms Wasm, 4 threads 60 ms

Frequently Asked Questions

Can I use the embeddings for clustering? Yes — normalised vectors work with k-means or hierarchical clustering on cosine similarity.

How large is the model download? Quantised small models are typically 20–35 MB.

Do I need Transformers.js? No — ONNX Runtime Web plus a tokeniser library works; Transformers.js saves glue code.

Is quantisation harmful? 8-bit quantisation usually changes results very little; verify against the reference.

How do I avoid re-embedding unchanged notes? Store a content hash with each vector and re-embed only when the hash changes, debounced while the user is typing.

How do I know whether a different model would search better? Keep a small evaluation set of real queries and expected results, and compare recall at 5 or 10 across models.

Should search queries wait for background indexing? No — give query embedding the highest priority in the worker’s queue so search stays responsive.

What vector size should I choose? 384 dimensions is a good default: compact indexes and fast search with strong quality for most retrieval tasks.

← Back to Machine Learning Inference in the Browser