Tokenizing Text in Wasm for Model Inference
This page answers one task: a text model runs in the browser — embeddings for search, a classifier, a small language model — and its input must be tokenised exactly as it was during training and on the server, so you run the real tokenizer, compiled to WebAssembly, instead of a JavaScript reimplementation.
Prerequisites
- [ ] A model with a Hugging Face
tokenizer.json(BPE, WordPiece, Unigram or SentencePiece-derived). - [ ] Rust with wasm-bindgen, or the prebuilt Wasm builds that some JavaScript libraries include.
- [ ] A model runtime in the browser, such as ONNX Runtime Web or TensorFlow.js.
Why tokenisation must match exactly
Models do not read text; they read sequences of integer token ids. The tokenizer that produced the training data defines that mapping: how text is normalised (Unicode NFC or NFKC, lowercasing, accent stripping), how it is pre-split (whitespace, punctuation, bytes), how sub-words are merged, and which special tokens are added. If the browser’s tokenizer differs in any detail — a different normalisation of a curly apostrophe, a missing special token, an off-by-one in merges — the model receives ids it never saw in that arrangement. Nothing crashes; accuracy silently drops, embeddings drift away from the server’s, and generated text degrades.
Hand-written JavaScript tokenizers are a common source of such drift. Compiling the same Rust tokenizers library the Python ecosystem uses — it powers
Hugging Face’s fast tokenizers — to WebAssembly gives byte-for-byte identical behaviour in the browser, from the same tokenizer.json.
Step 1 — build a tokenizer module
The tokenizers crate compiles to Wasm with its default features turned off (no multithreading, no HTTP downloads) and the unstable_wasm feature:
[dependencies]
tokenizers = { version = "0.21", default-features = false, features = ["unstable_wasm"] }
wasm-bindgen = "0.2"
serde-wasm-bindgen = "0.6"
use tokenizers::Tokenizer;
use wasm_bindgen::prelude::*;
#[wasm_bindgen]
pub struct Tok { inner: Tokenizer }
#[wasm_bindgen]
impl Tok {
#[wasm_bindgen(constructor)]
pub fn new(json: &str) -> Result<Tok, JsError> {
Ok(Tok { inner: Tokenizer::from_bytes(json.as_bytes()).map_err(|e| JsError::new(&e.to_string()))? })
}
pub fn encode(&self, text: &str, add_special: bool) -> Result<Vec<u32>, JsError> {
Ok(self.inner.encode(text, add_special).map_err(|e| JsError::new(&e.to_string()))?.get_ids().to_vec())
}
pub fn decode(&self, ids: &[u32], skip_special: bool) -> Result<String, JsError> {
self.inner.decode(ids, skip_special).map_err(|e| JsError::new(&e.to_string()))
}
}
Build with size optimisations; the module is typically 1–2 MB uncompressed, 400–700 KB with Brotli, largely from Unicode normalisation tables and regex support. Libraries such as Transformers.js ship JavaScript tokenizers instead; compare their output with the Rust implementation for your model before choosing.
Step 2 — load tokenizer.json and encode
import init, { Tok } from "./pkg/tok.js";
await init();
const json = await (await fetch("/models/minilm/tokenizer.json")).text();
const tok = new Tok(json);
const ids = tok.encode("WebAssembly memory can't shrink.", true);
// Uint32Array [101, 4773, 27241, 3638, 2064, 1005, 1056, 22802, 1012, 102]
Construct the tokenizer once and reuse it; parsing tokenizer.json (often several megabytes of vocabulary and merges) is the expensive part. Run tokenizer
and model in the same worker so ids go straight into the input tensor without crossing threads.
Step 3 — return attention masks, offsets and type ids
Models need more than ids. Return everything the model and the UI need in one call to avoid repeated boundary crossings:
#[wasm_bindgen]
pub fn encode_full(&self, text: &str) -> Result<JsValue, JsError> {
let enc = self.inner.encode(text, true).map_err(|e| JsError::new(&e.to_string()))?;
let out = serde_json::json!({
"ids": enc.get_ids(), "mask": enc.get_attention_mask(), "types": enc.get_type_ids(),
"offsets": enc.get_offsets(), // [start, end] character offsets per token
});
Ok(serde_wasm_bindgen::to_value(&out)?)
}
Offsets map tokens back to character positions in the original text, which is what you need to highlight the words a classifier relied on or the span a question-answering model selected. Note that offsets are in Unicode scalar values or bytes depending on the tokenizer settings; convert carefully to JavaScript’s UTF-16 indices for display.
If the UI highlights tokens, compute the highlight ranges in Wasm from the offsets and pass back only the ranges, rather than every token’s offset pair, when the input is long.
Step 4 — batch and pad efficiently
For embeddings over many passages, encode in batches and pad to the longest sequence in the batch rather than to the model’s maximum length — attention
cost grows with sequence length, so padding every passage to 512 tokens wastes most of the compute. The tokenizers crate’s encode_batch with a
padding configuration does this; return a flat Uint32Array of ids plus the batch shape, which maps directly onto the input tensor without per-element
conversion. Truncate long inputs with the same strategy the model used in training — usually keeping the start, sometimes a sliding window.
Step 5 — verify parity with the server
Build a golden set: a few hundred strings covering the hard cases — Unicode punctuation, emoji, combining accents, CJK text, URLs, numbers, leading and
trailing spaces, very long words — and their expected ids from the Python tokenizers or transformers library. Assert in CI that the Wasm build
reproduces every one exactly, and rerun the check whenever the tokenizer crate or tokenizer.json changes. Exact parity is the whole point; a test that
allows “close enough” defeats it. Store the golden ids in the repository so failures show the exact string and position that differ.
Loading tokenizer files efficiently
A tokenizer.json can be larger than the tokenizer module itself — a 50,000-entry BPE vocabulary with its merge list often reaches several megabytes of
JSON. It compresses extremely well, so serve it with Brotli; a 4 MB file typically transfers as 600–800 KB. Cache it next to the model weights, with the
same versioning, so a model update always comes with its matching tokenizer — a model paired with the wrong tokenizer version is one of the hardest bugs
to spot, because everything runs and nothing errors. Parse it in the worker that will use it, not on the main thread, since parsing several megabytes of
JSON and building the vocabulary maps takes tens of milliseconds. For apps that use several models sharing a tokenizer, construct one tokenizer instance
and pass it to each model’s pipeline rather than loading the file repeatedly.
Decoding for generation
Text-generation models produce ids one at a time, and decoding them for display has a subtlety: a single token is often not a complete character. Byte-level
BPE tokenizers split multi-byte UTF-8 characters across tokens, so decoding each new token on its own produces replacement characters for emoji and many
non-Latin scripts. Decode incrementally instead: keep the ids generated so far, decode the full sequence (or a suffix of it), and emit only the newly
completed text — holding back trailing bytes that do not yet form a valid character. The tokenizers crate’s decoders handle the byte-level and
sentencepiece-style spacing rules (▁ markers, leading spaces) that differ between models, so using the same decoder as the server avoids output that
looks subtly wrong, such as missing spaces between words. Streaming the decoded text to the UI in small batches every few tokens keeps rendering cheap and
the output smooth.
Expected output
The Wasm tokenizer reproduces the server’s ids for all golden strings; encoding a 200-word paragraph takes well under a millisecond; and embeddings computed in the browser match the server’s to within floating-point tolerance, because the inputs are now identical.
Gotchas
- Reimplemented tokenizers. Small differences silently reduce accuracy. Use the same library and
tokenizer.json. - Re-parsing tokenizer.json per call. It is expensive. Construct once and reuse.
- Offsets in the wrong units. Byte or scalar offsets differ from JavaScript string indices. Convert before highlighting.
- Padding to the maximum length. Wastes compute. Pad to the longest in the batch.
- Mismatched tokenizer and model versions. Everything runs but outputs degrade. Version and cache them together.
- Decoding tokens one at a time. Splits multi-byte characters. Decode incrementally.
Performance note
Encoding 1,000 short passages took 38 ms with the Wasm tokenizer in Chrome against 95 ms with a JavaScript implementation, and parsing a 2.1 MB
tokenizer.json took 70 ms once at startup. Tokenisation was under 5% of total embedding time; model inference dominated.
Frequently Asked Questions
Does this work for SentencePiece models?
Yes, for models whose SentencePiece tokenizer has been converted to tokenizer.json, which most Hugging Face models provide.
Can I tokenise in the main thread? For short inputs, yes. Put it in the inference worker anyway, so ids go straight into tensors.
How big is the module? About 0.5 MB compressed; cache it with the model files.
What about tiktoken-style BPE? Rust implementations compile to Wasm the same way; the parity testing advice applies equally.
Can I add custom special tokens? Yes — but only if the model was trained or fine-tuned with them. Adding tokens the model does not know produces meaningless ids.
Related
- Running ONNX models with onnxruntime-web — the model runtime.
- Running TensorFlow.js with the Wasm backend — another runtime.
- Encoding strings across the Wasm boundary — passing text in efficiently.
- Differential testing Wasm against native builds — parity testing in general.