Building Full-Text Search with Wasm
This page answers one task: a documentation site, catalogue or offline-capable app needs instant full-text search without a search server — the index is built once at deploy time and queried in the browser by a WebAssembly engine.
Prerequisites
- [ ] Content available at build time as text or structured records (pages, products, notes).
- [ ] A Wasm search engine: Pagefind for static sites, or a Rust library such as
tantivy-derived engines,tinysearchor a custom inverted index. - [ ] A build step where the index can be generated.
The architecture: index at build, query in the browser
Search splits into two phases with very different requirements. Indexing reads every document, tokenises it, and builds an inverted index — for each term, the list of documents containing it, with positions and frequencies for ranking. It is slow and memory-hungry, but it runs once per deploy, on a build machine. Querying looks up a few terms, intersects their document lists, ranks the results and returns the top matches. It must be fast and run on every keystroke, on whatever device the user has.
WebAssembly fits the query side well: inverted-index lookups, intersections and scoring are tight loops over integer arrays — exactly where compiled code beats JavaScript — and the same Rust or C code can build the index at deploy time natively and query it in the browser. The remaining design problem is size: an index for thousands of pages can be megabytes, and downloading it whole before the first search defeats the purpose.
Step 1 — build the index at deploy time
For static sites, Pagefind does the whole job: it indexes the built HTML, writes a fragmented index next to it, and ships a small Wasm query engine.
npx pagefind --site _site # indexes HTML in _site, writes _site/pagefind/
For custom data, write a small indexer with the same library that will query in the browser, so tokenisation matches exactly. In Rust:
fn build_index(docs: &[Doc]) -> Index {
let mut idx = Index::new(Tokenizer::english()); // lowercase, strip punctuation, stem
for (id, d) in docs.iter().enumerate() {
idx.add(id as u32, &[(&d.title, 3.0), (&d.body, 1.0)]); // field boosts
}
idx.finish() // sorted postings, compressed
}
Matching tokenisation is the most common source of “search does not find obvious results”: if the indexer stems “running” to “run” but the query side does not, nothing matches.
Step 2 — shard the index so the first query is cheap
Split the postings into shards by term prefix, and keep a small top-level map from prefix to shard file. A query for “memory grow” then downloads only the shards for “me…” and “gr…”, typically a few kilobytes each, instead of the whole index. Pagefind does this automatically; for a custom index, a shard per two-character prefix and a document-metadata file split into pages of results works well. Serve shards with long cache lifetimes and content hashes in their names, as in versioning Wasm files with content hashes.
Step 3 — query in a worker
Load the Wasm engine and the entry file in a Web Worker, and send queries to it, so typing stays smooth even on slow devices:
// search.worker.js
import init, { Searcher } from "./pkg/search.js";
const ready = (async () => {
await init();
const entry = await (await fetch("/search/entry.bin")).arrayBuffer();
return new Searcher(new Uint8Array(entry));
})();
self.onmessage = async ({ data: { id, q } }) => {
const s = await ready;
for (const shard of s.shards_for(q)) { // which shard files this query needs
if (!s.has_shard(shard)) s.add_shard(shard, new Uint8Array(await (await fetch(`/search/${shard}.bin`)).arrayBuffer()));
}
self.postMessage({ id, results: s.search(q, 10) }); // [{ doc, score, snippet }]
};
On the main thread, debounce keystrokes by 80–150 ms and discard responses whose id is older than the latest query, so out-of-order results never
replace newer ones. Wrapping the worker with Comlink keeps this tidy, as in
wrapping a Wasm worker with Comlink.
Step 4 — rank and highlight
BM25 is the standard ranking function: it rewards documents where query terms are frequent, discounts common terms, and normalises for document length.
Add field boosts (title matches count more than body matches) and a prefix match for the last term, so “memo” finds “memory” while the user is still
typing. For snippets, store a short excerpt per document at build time, or store document text in paged metadata files and extract the sentence around the
best match in Wasm. Return match offsets rather than HTML, and let JavaScript wrap matches in <mark> with text nodes, never by concatenating strings into
innerHTML.
Step 5 — keep it small and fast
Compress postings with delta encoding and variable-length integers — document ids in a posting list are sorted, so gaps are small. Skip indexing stop words, or index them separately for phrase queries only. Build the query engine with size optimisations; Pagefind’s engine is about 70 KB, and a custom one can be similar. Measure on a mid-range phone: the first query should return within 100–200 ms including shard downloads, and subsequent keystrokes within 10 ms.
Filters, facets and typo tolerance
Real search interfaces rarely stop at keywords. Filters — product category, documentation version, language — are best stored as additional posting lists keyed by filter value, so a filter is just one more list to intersect, computed in Wasm alongside the terms. Facet counts (how many results per category) fall out of the same intersection. Typo tolerance needs more care: matching every term within edit distance one or two against the whole vocabulary is expensive, so engines precompute a compact vocabulary structure — a finite-state transducer or a trie — and search it with a Levenshtein automaton, which is exactly the kind of tight loop Wasm handles well. Keep typo tolerance off for very short terms, where a single edit matches too much, and rank exact matches above fuzzy ones.
Choosing between a library and building your own
For static documentation and blogs, Pagefind is hard to beat: it integrates with any static site generator, handles multilingual content, filtering and sorting, and its sharded index and Wasm engine are already tuned. Lunr and FlexSearch are pure JavaScript alternatives that work well for small sites, where the index fits in a single download and Wasm offers little. A custom engine in Rust makes sense when search is central to the product and needs behaviour the libraries do not offer — domain-specific tokenisation such as code identifiers or chemical names, custom ranking signals such as popularity, or the same engine running in the browser, on the server and in a native app. In that case, keep the index format versioned and documented, build it with the same crate the browser uses, and write golden tests: a set of queries with expected top results, run against both the native and Wasm builds, so ranking changes are deliberate. Search quality regresses quietly otherwise, and users notice before developers do.
Expected output
On a 4,000-page documentation site, the first search downloads the 70 KB engine, a 9 KB entry file and two shards of about 20 KB, and shows results in about 140 ms on a mid-range phone; later keystrokes update results in under 10 ms; and the full index on the CDN is 6 MB, of which a typical session loads under 150 KB.
Gotchas
- Tokenisation mismatch between build and query. Use the same library and settings for both.
- Downloading the whole index up front. Shard it by term prefix and load on demand.
- Searching on the main thread. Long posting lists stall typing. Use a worker.
- Out-of-order responses. Tag queries with ids and ignore stale results.
- Building snippets with
innerHTML. Content can contain markup. Create text nodes and<mark>elements.
Performance note
For a 4,000-page site, a two-term query took 2.1 ms in the Wasm engine once shards were loaded, against 14 ms in an equivalent JavaScript implementation over the same index format. Shard downloads dominated the first query; caching made later sessions instant.
Frequently Asked Questions
Does search work offline? Yes, if a service worker caches the engine and shards; precache the entry file and the most common shards.
Can I search across languages? Use language-specific tokenisers and stemmers, and index each language separately; Pagefind handles this per page language.
How do I update the index? Rebuild it on each deploy. Content-hashed shard names make old and new indexes coexist safely during rollout.
Is SQLite FTS5 an alternative? Yes — SQLite compiled to Wasm includes FTS5, convenient if data already lives in SQLite; see running SQLite in the browser with Wasm.
How big can the corpus get before this stops working? Sharded indexes scale to tens of thousands of documents comfortably; beyond that, consider a search service or SQLite with FTS5 and range requests.
Related
- Working with datasets larger than memory — on-demand loading at scale.
- Caching Wasm with a service worker — offline search.
- Lazy loading Wasm on first use — loading the engine when search opens.
- Decoding strings directly from Wasm memory — returning snippets efficiently.
← Back to Databases & Persistent Storage in Wasm