Running Small Language Models in the Browser

This page answers one task: an application wants text generation, summarisation, classification or chat without sending user data to a server — and you want to know whether a small language model can run in the browser through WebAssembly, which models fit, how fast they are, and how to keep the page usable while they run.

Prerequisites

  • [ ] A WebAssembly build of a llama.cpp-compatible runtime (for example wllama or a similar Emscripten build), or a framework that packages one.
  • [ ] A small model in GGUF format, quantised (0.5–3 billion parameters is the practical range).
  • [ ] A cross-origin isolated page for multi-threaded inference.

What fits in a browser

Language-model inference is dominated by reading weights: every generated token touches every weight once. Two constraints follow. Memory: weights must fit in the Wasm module’s linear memory, which for 32-bit Wasm is at most 4 GB and in practice less (browsers and devices may cap it lower). A 1-billion-parameter model at 4-bit quantisation is about 600–700 MB; a 3B model at 4 bits about 1.8–2 GB — near the limit. Some runtimes split weights into chunks and use Memory64 where available to go further. Speed: token generation is memory-bandwidth-bound; on a laptop CPU with SIMD and several threads, small models produce tokens at conversational speeds, while larger ones crawl.

That puts browser LLMs in a particular niche: small, quantised models for focused tasks — classifying text, extracting fields, rewriting short passages, autocomplete, private chat with modest expectations. For larger models, WebGPU-based runtimes are typically much faster where available, with Wasm as the fallback.

Model sizes and browser practicality Models around half a billion parameters at 4-bit quantisation are a few hundred megabytes and run quickly on most laptops. One to two billion parameters fit and run at usable speeds on laptops. Three billion approaches 32-bit memory limits. Seven billion and above exceed practical CPU Wasm limits and need WebGPU or a server. model size (4-bit) weights laptop CPU (Wasm, threads) notes 0.5B ~350 MB fast phones feasible 1–2B 0.7–1.3 GB usable laptops comfortable 3B ~2 GB slow near 32-bit memory limits 7B+ 4 GB+ impractical WebGPU or server

Step 1 — pick a model for the task

Choose the smallest model that does the job. Instruction-tuned models in the 0.5–2B range handle classification, extraction and short rewriting surprisingly well; open-ended chat quality improves noticeably with size. Prefer models published in GGUF with 4-bit quantisation (Q4_K_M or similar). Test candidate models on your actual prompts natively first (with llama.cpp on a laptop) — it is faster to iterate there, and the browser will produce the same outputs more slowly.

Step 2 — load the runtime and the weights

import { Wllama } from "@wllama/wllama";

const wllama = new Wllama({
  "single-thread/wllama.wasm": "/wasm/single-thread/wllama.wasm",
  "multi-thread/wllama.wasm": "/wasm/multi-thread/wllama.wasm",
});
await wllama.loadModelFromUrl("/models/qwen2.5-0.5b-instruct-q4_k_m.gguf", {
  n_ctx: 2048,
  progressCallback: ({ loaded, total }) => showProgress(loaded / total),
});

Library APIs differ and change; the shape is the same: load the runtime (choosing the multi-threaded build when the page is cross-origin isolated), download or read cached weights, and create a context with a size that fits memory. Large GGUF files are often split into parts so each download and allocation stays manageable.

Step 3 — stream tokens to the UI

Generation produces one token at a time. Stream them to the page as they arrive, so users see text appear rather than waiting for the whole response:

let text = "";
await wllama.createCompletion(prompt, {
  nPredict: 256,
  sampling: { temp: 0.7, top_p: 0.9 },
  onNewToken: (_token, _piece, currentText) => { text = currentText; render(text); },
});

Run the runtime in a worker (most browser LLM libraries do this internally) so token generation does not block input. Provide a stop button that aborts generation, and limit nPredict so a runaway response cannot run for minutes.

Generating a response locally The page loads the runtime and checks for cross-origin isolation to choose the threaded build. Weights load from cache or the network into Wasm memory. The prompt is tokenised and processed, then tokens are generated one by one with SIMD and threads, each streamed to the UI until the limit or a stop token. choose threaded build crossOriginIsolated? load weights cache or network process prompt prefill generate token by token SIMD + threads stream to UI stop button, limits

Step 4 — cache the weights

Hundreds of megabytes must not download on every visit. Store weights in the Cache API or OPFS after the first download, verify their hash, and request persistent storage. Check available quota with navigator.storage.estimate() before downloading, and explain to users what will be stored and how to remove it. Some runtimes include caching; otherwise follow caching model files for offline inference.

Step 5 — measure tokens per second and time to first token

Two numbers describe the experience: time to first token (prompt processing, grows with prompt length) and tokens per second during generation. Measure both on target devices with realistic prompts. Keep prompts short — long system prompts and retrieved context cost prompt-processing time on every request — and reuse the KV cache across turns where the runtime supports it, so a chat does not re-process the whole conversation for each reply.

WebGPU versus Wasm

Runtimes that use WebGPU (for example WebLLM, or ONNX Runtime Web’s WebGPU backend for supported models) run matrix operations on the GPU and are typically several times faster for the same model, and can handle larger models. They need a WebGPU-capable browser and a GPU with enough memory. A robust application detects WebGPU and uses it, falling back to the Wasm CPU runtime elsewhere, possibly with a smaller model for the fallback. See choosing between the Wasm and WebGPU backends.

Setting expectations

Small local models make mistakes more often than large hosted ones. Use them where errors are cheap or reviewable (suggestions the user accepts or edits, classification with a fallback), constrain outputs (structured output grammars, short answers), and avoid presenting generated text as authoritative. Being explicit in the UI that processing is local — and therefore private — helps users understand both the benefit and the limits.

Constraining output for reliable features

Free-form generation is hard to build features on: a classification that sometimes answers “Positive!” and sometimes “I think this is positive” breaks parsing. llama.cpp-based runtimes support grammars (GBNF) and JSON-schema-constrained sampling, which restrict generated tokens to a format you define — one of a fixed set of labels, a JSON object with given fields, a number. Constrained sampling makes small models far more dependable for structured tasks, and it also shortens responses, which matters when every token costs tens of milliseconds. Combine it with short, few-shot prompts that show the expected output, and validate results anyway; a well-formed answer can still be wrong.

Retrieval with local data

Many useful local-LLM features answer questions about the user’s own data — notes, documents, emails — without uploading it. The pattern is retrieval: embed the user’s documents locally, find the passages most similar to the question with a vector index, and include only those passages in the prompt. Small models handle “answer from this context” far better than open-ended knowledge questions, and the prompt stays short enough for reasonable time to first token. Keep retrieved context to a few hundred tokens, cite which passage an answer came from, and let users open the source to check it. See indexing vectors for similarity search in the browser.

Licensing and model provenance

Model weights come with licences that vary — some permit commercial use freely, others restrict it or require attribution. Record each model’s licence and source with the files you serve, and pin exact files by hash so the model users download is the one you evaluated.

Expected output

A 0.5B instruct model at Q4_K_M (about 400 MB) loads from cache in 2 seconds and generates about 20 tokens per second on a laptop with 8 threads; time to first token for a 200-token prompt is under a second; responses stream into the UI with a working stop button; WebGPU-capable browsers use a GPU runtime instead; and the page stays responsive throughout.

Gotchas

  • Models too large for 32-bit memory. Allocation fails. Stay within practical limits or use WebGPU.
  • No cross-origin isolation. Single-threaded inference is slow. Set COOP/COEP.
  • Long prompts. Time to first token grows. Keep prompts tight; reuse KV cache.
  • Downloading weights every visit. Cache them and verify hashes.
  • Generation on the main thread. The UI freezes. Use a worker.
  • Unconstrained output for structured tasks. Parsing breaks. Use grammars or JSON-schema sampling.

Performance note

On an 8-core laptop, a 0.5B Q4 model generated about 20 tokens/s with 8 threads and about 6 tokens/s with 1 thread; a 1.5B Q4 model reached about 8 tokens/s with 8 threads. A WebGPU runtime on the same machine ran the 1.5B model at over 30 tokens/s.

Generation speed on an 8-core laptop Tokens per second for a 0.5 billion parameter model with one and eight threads in Wasm, a 1.5 billion parameter model with eight threads in Wasm, and the 1.5 billion model on a WebGPU runtime. tokens per second 0.5B Q4, Wasm, 1 thread 6 tok/s 0.5B Q4, Wasm, 8 threads 20 tok/s 1.5B Q4, Wasm, 8 threads 8 tok/s 1.5B Q4, WebGPU 32 tok/s

Frequently Asked Questions

Can phones run these models? The smallest models, slowly; memory limits and thermal throttling make larger ones impractical.

Does Memory64 remove the 4 GB limit? It raises the address limit where supported, but device memory and browser policies still apply.

Can I fine-tune in the browser? Training is far heavier than inference; do it offline and ship the result.

Are outputs identical to native llama.cpp? With the same model, settings and seed, they should match closely; minor numeric differences can occur.

How do I get reliably structured output from a small model? Use grammar or JSON-schema constrained sampling with short few-shot prompts, and still validate the result.

Can a local model answer questions about the user’s documents? Yes, with local retrieval: embed documents, find relevant passages, and include only those in a short prompt.

← Back to Machine Learning Inference in the Browser