Building Full-Text Search with Wasm

This page answers one task: a documentation site, catalogue or offline-capable app needs instant full-text search without a search server — the index is built once at deploy time and queried in the browser by a WebAssembly engine.

Prerequisites

  • [ ] Content available at build time as text or structured records (pages, products, notes).
  • [ ] A Wasm search engine: Pagefind for static sites, or a Rust library such as tantivy-derived engines, tinysearch or a custom inverted index.
  • [ ] A build step where the index can be generated.

The architecture: index at build, query in the browser

Search splits into two phases with very different requirements. Indexing reads every document, tokenises it, and builds an inverted index — for each term, the list of documents containing it, with positions and frequencies for ranking. It is slow and memory-hungry, but it runs once per deploy, on a build machine. Querying looks up a few terms, intersects their document lists, ranks the results and returns the top matches. It must be fast and run on every keystroke, on whatever device the user has.

WebAssembly fits the query side well: inverted-index lookups, intersections and scoring are tight loops over integer arrays — exactly where compiled code beats JavaScript — and the same Rust or C code can build the index at deploy time natively and query it in the browser. The remaining design problem is size: an index for thousands of pages can be megabytes, and downloading it whole before the first search defeats the purpose.

Full-text search split between build and browser At deploy time a native build of the search engine indexes all documents and writes a sharded index. In the browser, a small Wasm query engine loads the shards needed for the typed terms, ranks matching documents and returns results with snippets. content at build pages, records native indexer tokenise, invert, shard static index shards served from the CDN Wasm query engine in a worker ranked results with snippets

Step 1 — build the index at deploy time

For static sites, Pagefind does the whole job: it indexes the built HTML, writes a fragmented index next to it, and ships a small Wasm query engine.

npx pagefind --site _site            # indexes HTML in _site, writes _site/pagefind/

For custom data, write a small indexer with the same library that will query in the browser, so tokenisation matches exactly. In Rust:

fn build_index(docs: &[Doc]) -> Index {
    let mut idx = Index::new(Tokenizer::english());       // lowercase, strip punctuation, stem
    for (id, d) in docs.iter().enumerate() {
        idx.add(id as u32, &[(&d.title, 3.0), (&d.body, 1.0)]);   // field boosts
    }
    idx.finish()                                            // sorted postings, compressed
}

Matching tokenisation is the most common source of “search does not find obvious results”: if the indexer stems “running” to “run” but the query side does not, nothing matches.

Step 2 — shard the index so the first query is cheap

Split the postings into shards by term prefix, and keep a small top-level map from prefix to shard file. A query for “memory grow” then downloads only the shards for “me…” and “gr…”, typically a few kilobytes each, instead of the whole index. Pagefind does this automatically; for a custom index, a shard per two-character prefix and a document-metadata file split into pages of results works well. Serve shards with long cache lifetimes and content hashes in their names, as in versioning Wasm files with content hashes.

A sharded search index on the CDN The index consists of a small entry file mapping term prefixes to shards, many small posting shards by prefix, and paged document metadata for titles and snippets. A query downloads the entry file once and then only the shards its terms need. index files: entry map, prefix shards, paged metadata entry shard aa–al shard am–az … 400 shards … metadata pages 1 request only what a query needs

Step 3 — query in a worker

Load the Wasm engine and the entry file in a Web Worker, and send queries to it, so typing stays smooth even on slow devices:

// search.worker.js
import init, { Searcher } from "./pkg/search.js";
const ready = (async () => {
  await init();
  const entry = await (await fetch("/search/entry.bin")).arrayBuffer();
  return new Searcher(new Uint8Array(entry));
})();

self.onmessage = async ({ data: { id, q } }) => {
  const s = await ready;
  for (const shard of s.shards_for(q)) {                    // which shard files this query needs
    if (!s.has_shard(shard)) s.add_shard(shard, new Uint8Array(await (await fetch(`/search/${shard}.bin`)).arrayBuffer()));
  }
  self.postMessage({ id, results: s.search(q, 10) });       // [{ doc, score, snippet }]
};

On the main thread, debounce keystrokes by 80–150 ms and discard responses whose id is older than the latest query, so out-of-order results never replace newer ones. Wrapping the worker with Comlink keeps this tidy, as in wrapping a Wasm worker with Comlink.

Step 4 — rank and highlight

BM25 is the standard ranking function: it rewards documents where query terms are frequent, discounts common terms, and normalises for document length. Add field boosts (title matches count more than body matches) and a prefix match for the last term, so “memo” finds “memory” while the user is still typing. For snippets, store a short excerpt per document at build time, or store document text in paged metadata files and extract the sentence around the best match in Wasm. Return match offsets rather than HTML, and let JavaScript wrap matches in <mark> with text nodes, never by concatenating strings into innerHTML.

Step 5 — keep it small and fast

Compress postings with delta encoding and variable-length integers — document ids in a posting list are sorted, so gaps are small. Skip indexing stop words, or index them separately for phrase queries only. Build the query engine with size optimisations; Pagefind’s engine is about 70 KB, and a custom one can be similar. Measure on a mid-range phone: the first query should return within 100–200 ms including shard downloads, and subsequent keystrokes within 10 ms.

Filters, facets and typo tolerance

Real search interfaces rarely stop at keywords. Filters — product category, documentation version, language — are best stored as additional posting lists keyed by filter value, so a filter is just one more list to intersect, computed in Wasm alongside the terms. Facet counts (how many results per category) fall out of the same intersection. Typo tolerance needs more care: matching every term within edit distance one or two against the whole vocabulary is expensive, so engines precompute a compact vocabulary structure — a finite-state transducer or a trie — and search it with a Levenshtein automaton, which is exactly the kind of tight loop Wasm handles well. Keep typo tolerance off for very short terms, where a single edit matches too much, and rank exact matches above fuzzy ones.

Choosing between a library and building your own

For static documentation and blogs, Pagefind is hard to beat: it integrates with any static site generator, handles multilingual content, filtering and sorting, and its sharded index and Wasm engine are already tuned. Lunr and FlexSearch are pure JavaScript alternatives that work well for small sites, where the index fits in a single download and Wasm offers little. A custom engine in Rust makes sense when search is central to the product and needs behaviour the libraries do not offer — domain-specific tokenisation such as code identifiers or chemical names, custom ranking signals such as popularity, or the same engine running in the browser, on the server and in a native app. In that case, keep the index format versioned and documented, build it with the same crate the browser uses, and write golden tests: a set of queries with expected top results, run against both the native and Wasm builds, so ranking changes are deliberate. Search quality regresses quietly otherwise, and users notice before developers do.

Expected output

On a 4,000-page documentation site, the first search downloads the 70 KB engine, a 9 KB entry file and two shards of about 20 KB, and shows results in about 140 ms on a mid-range phone; later keystrokes update results in under 10 ms; and the full index on the CDN is 6 MB, of which a typical session loads under 150 KB.

Gotchas

  • Tokenisation mismatch between build and query. Use the same library and settings for both.
  • Downloading the whole index up front. Shard it by term prefix and load on demand.
  • Searching on the main thread. Long posting lists stall typing. Use a worker.
  • Out-of-order responses. Tag queries with ids and ignore stale results.
  • Building snippets with innerHTML. Content can contain markup. Create text nodes and <mark> elements.

Performance note

For a 4,000-page site, a two-term query took 2.1 ms in the Wasm engine once shards were loaded, against 14 ms in an equivalent JavaScript implementation over the same index format. Shard downloads dominated the first query; caching made later sessions instant.

Time to show results for the first query Milliseconds from typing the first query to showing results on a mid-range phone, loading a single full JavaScript index, loading a sharded index with a Wasm engine, and a later query with shards cached. ms to first results full JS index (6 MB) 2,900 ms sharded index + Wasm engine 140 ms later query, cached shards 8 ms

Frequently Asked Questions

Does search work offline? Yes, if a service worker caches the engine and shards; precache the entry file and the most common shards.

Can I search across languages? Use language-specific tokenisers and stemmers, and index each language separately; Pagefind handles this per page language.

How do I update the index? Rebuild it on each deploy. Content-hashed shard names make old and new indexes coexist safely during rollout.

Is SQLite FTS5 an alternative? Yes — SQLite compiled to Wasm includes FTS5, convenient if data already lives in SQLite; see running SQLite in the browser with Wasm.

How big can the corpus get before this stops working? Sharded indexes scale to tens of thousands of documents comfortably; beyond that, consider a search service or SQLite with FTS5 and range requests.

← Back to Databases & Persistent Storage in Wasm