Tokenizing Text in Wasm for Model Inference

This page answers one task: a text model runs in the browser — embeddings for search, a classifier, a small language model — and its input must be tokenised exactly as it was during training and on the server, so you run the real tokenizer, compiled to WebAssembly, instead of a JavaScript reimplementation.

Prerequisites

  • [ ] A model with a Hugging Face tokenizer.json (BPE, WordPiece, Unigram or SentencePiece-derived).
  • [ ] Rust with wasm-bindgen, or the prebuilt Wasm builds that some JavaScript libraries include.
  • [ ] A model runtime in the browser, such as ONNX Runtime Web or TensorFlow.js.

Why tokenisation must match exactly

Models do not read text; they read sequences of integer token ids. The tokenizer that produced the training data defines that mapping: how text is normalised (Unicode NFC or NFKC, lowercasing, accent stripping), how it is pre-split (whitespace, punctuation, bytes), how sub-words are merged, and which special tokens are added. If the browser’s tokenizer differs in any detail — a different normalisation of a curly apostrophe, a missing special token, an off-by-one in merges — the model receives ids it never saw in that arrangement. Nothing crashes; accuracy silently drops, embeddings drift away from the server’s, and generated text degrades.

Hand-written JavaScript tokenizers are a common source of such drift. Compiling the same Rust tokenizers library the Python ecosystem uses — it powers Hugging Face’s fast tokenizers — to WebAssembly gives byte-for-byte identical behaviour in the browser, from the same tokenizer.json.

The tokenisation pipeline the browser must reproduce Input text is normalised, pre-tokenised into words or bytes, split into sub-word tokens by the model's algorithm, wrapped with special tokens, and converted to ids with an attention mask. Every step is defined by tokenizer.json and must match training exactly. normalise NFC/NFKC, lowercase … pre-tokenise whitespace, bytes model BPE / WordPiece / Unigram post-process [CLS] … [SEP] ids + mask into the model

Step 1 — build a tokenizer module

The tokenizers crate compiles to Wasm with its default features turned off (no multithreading, no HTTP downloads) and the unstable_wasm feature:

[dependencies]
tokenizers = { version = "0.21", default-features = false, features = ["unstable_wasm"] }
wasm-bindgen = "0.2"
serde-wasm-bindgen = "0.6"
use tokenizers::Tokenizer;
use wasm_bindgen::prelude::*;

#[wasm_bindgen]
pub struct Tok { inner: Tokenizer }

#[wasm_bindgen]
impl Tok {
    #[wasm_bindgen(constructor)]
    pub fn new(json: &str) -> Result<Tok, JsError> {
        Ok(Tok { inner: Tokenizer::from_bytes(json.as_bytes()).map_err(|e| JsError::new(&e.to_string()))? })
    }

    pub fn encode(&self, text: &str, add_special: bool) -> Result<Vec<u32>, JsError> {
        Ok(self.inner.encode(text, add_special).map_err(|e| JsError::new(&e.to_string()))?.get_ids().to_vec())
    }

    pub fn decode(&self, ids: &[u32], skip_special: bool) -> Result<String, JsError> {
        self.inner.decode(ids, skip_special).map_err(|e| JsError::new(&e.to_string()))
    }
}

Build with size optimisations; the module is typically 1–2 MB uncompressed, 400–700 KB with Brotli, largely from Unicode normalisation tables and regex support. Libraries such as Transformers.js ship JavaScript tokenizers instead; compare their output with the Rust implementation for your model before choosing.

Step 2 — load tokenizer.json and encode

import init, { Tok } from "./pkg/tok.js";
await init();
const json = await (await fetch("/models/minilm/tokenizer.json")).text();
const tok = new Tok(json);

const ids = tok.encode("WebAssembly memory can't shrink.", true);
// Uint32Array [101, 4773, 27241, 3638, 2064, 1005, 1056, 22802, 1012, 102]

Construct the tokenizer once and reuse it; parsing tokenizer.json (often several megabytes of vocabulary and merges) is the expensive part. Run tokenizer and model in the same worker so ids go straight into the input tensor without crossing threads.

Step 3 — return attention masks, offsets and type ids

Models need more than ids. Return everything the model and the UI need in one call to avoid repeated boundary crossings:

#[wasm_bindgen]
pub fn encode_full(&self, text: &str) -> Result<JsValue, JsError> {
    let enc = self.inner.encode(text, true).map_err(|e| JsError::new(&e.to_string()))?;
    let out = serde_json::json!({
        "ids": enc.get_ids(), "mask": enc.get_attention_mask(), "types": enc.get_type_ids(),
        "offsets": enc.get_offsets(),                       // [start, end] character offsets per token
    });
    Ok(serde_wasm_bindgen::to_value(&out)?)
}

Offsets map tokens back to character positions in the original text, which is what you need to highlight the words a classifier relied on or the span a question-answering model selected. Note that offsets are in Unicode scalar values or bytes depending on the tokenizer settings; convert carefully to JavaScript’s UTF-16 indices for display.

One encoded sentence, field by field Encoding a sentence returns token ids, an attention mask marking real tokens, token type ids for segment pairs, and character offsets mapping each token back to the input text. Special tokens have empty offsets. ids: [101, 4773, 27241, …, 102] vocabulary indices mask: [1, 1, 1, …, 1] 1 = real token, 0 = padding types: [0, 0, 0, …, 0] segment A / B for pairs offsets: [[0,0], [0,3], [3,11], …] chars in the input [CLS] and [SEP] added by the post-processor

If the UI highlights tokens, compute the highlight ranges in Wasm from the offsets and pass back only the ranges, rather than every token’s offset pair, when the input is long.

Step 4 — batch and pad efficiently

For embeddings over many passages, encode in batches and pad to the longest sequence in the batch rather than to the model’s maximum length — attention cost grows with sequence length, so padding every passage to 512 tokens wastes most of the compute. The tokenizers crate’s encode_batch with a padding configuration does this; return a flat Uint32Array of ids plus the batch shape, which maps directly onto the input tensor without per-element conversion. Truncate long inputs with the same strategy the model used in training — usually keeping the start, sometimes a sliding window.

Step 5 — verify parity with the server

Build a golden set: a few hundred strings covering the hard cases — Unicode punctuation, emoji, combining accents, CJK text, URLs, numbers, leading and trailing spaces, very long words — and their expected ids from the Python tokenizers or transformers library. Assert in CI that the Wasm build reproduces every one exactly, and rerun the check whenever the tokenizer crate or tokenizer.json changes. Exact parity is the whole point; a test that allows “close enough” defeats it. Store the golden ids in the repository so failures show the exact string and position that differ.

Loading tokenizer files efficiently

A tokenizer.json can be larger than the tokenizer module itself — a 50,000-entry BPE vocabulary with its merge list often reaches several megabytes of JSON. It compresses extremely well, so serve it with Brotli; a 4 MB file typically transfers as 600–800 KB. Cache it next to the model weights, with the same versioning, so a model update always comes with its matching tokenizer — a model paired with the wrong tokenizer version is one of the hardest bugs to spot, because everything runs and nothing errors. Parse it in the worker that will use it, not on the main thread, since parsing several megabytes of JSON and building the vocabulary maps takes tens of milliseconds. For apps that use several models sharing a tokenizer, construct one tokenizer instance and pass it to each model’s pipeline rather than loading the file repeatedly.

Decoding for generation

Text-generation models produce ids one at a time, and decoding them for display has a subtlety: a single token is often not a complete character. Byte-level BPE tokenizers split multi-byte UTF-8 characters across tokens, so decoding each new token on its own produces replacement characters for emoji and many non-Latin scripts. Decode incrementally instead: keep the ids generated so far, decode the full sequence (or a suffix of it), and emit only the newly completed text — holding back trailing bytes that do not yet form a valid character. The tokenizers crate’s decoders handle the byte-level and sentencepiece-style spacing rules (▁ markers, leading spaces) that differ between models, so using the same decoder as the server avoids output that looks subtly wrong, such as missing spaces between words. Streaming the decoded text to the UI in small batches every few tokens keeps rendering cheap and the output smooth.

Expected output

The Wasm tokenizer reproduces the server’s ids for all golden strings; encoding a 200-word paragraph takes well under a millisecond; and embeddings computed in the browser match the server’s to within floating-point tolerance, because the inputs are now identical.

Gotchas

  • Reimplemented tokenizers. Small differences silently reduce accuracy. Use the same library and tokenizer.json.
  • Re-parsing tokenizer.json per call. It is expensive. Construct once and reuse.
  • Offsets in the wrong units. Byte or scalar offsets differ from JavaScript string indices. Convert before highlighting.
  • Padding to the maximum length. Wastes compute. Pad to the longest in the batch.
  • Mismatched tokenizer and model versions. Everything runs but outputs degrade. Version and cache them together.
  • Decoding tokens one at a time. Splits multi-byte characters. Decode incrementally.

Performance note

Encoding 1,000 short passages took 38 ms with the Wasm tokenizer in Chrome against 95 ms with a JavaScript implementation, and parsing a 2.1 MB tokenizer.json took 70 ms once at startup. Tokenisation was under 5% of total embedding time; model inference dominated.

Tokenizing 1,000 short passages Milliseconds to tokenize one thousand short passages with the Hugging Face tokenizers crate compiled to Wasm and with a JavaScript tokenizer implementation, plus the one-time cost of parsing tokenizer.json. ms Wasm tokenizers crate 38 ms JavaScript implementation 95 ms parse tokenizer.json (once) 70 ms

Frequently Asked Questions

Does this work for SentencePiece models? Yes, for models whose SentencePiece tokenizer has been converted to tokenizer.json, which most Hugging Face models provide.

Can I tokenise in the main thread? For short inputs, yes. Put it in the inference worker anyway, so ids go straight into tensors.

How big is the module? About 0.5 MB compressed; cache it with the model files.

What about tiktoken-style BPE? Rust implementations compile to Wasm the same way; the parity testing advice applies equally.

Can I add custom special tokens? Yes — but only if the model was trained or fine-tuned with them. Adding tokens the model does not know produces meaningless ids.

← Back to Machine Learning Inference in the Browser