Processing PDFs in the Browser with Wasm
This page answers one task: an application must work with PDF files on the client — show page previews, extract text for search or AI features, merge uploaded documents, split or reorder pages — without sending the files to a server, and with the fidelity of a mature PDF engine.
Prerequisites
- [ ] A WebAssembly build of a PDF engine: PDFium (for example via
@hyzyla/pdfiumor similar packages) or MuPDF (mupdf.js), or PDF.js as a JavaScript alternative. - [ ] A worker in which the engine runs.
- [ ] A clear view of licensing: PDFium is BSD-style; MuPDF is AGPL or commercial.
Engines and how they differ
Three engines dominate browser PDF work. PDF.js (Mozilla) is pure JavaScript, powers Firefox’s viewer, and renders and extracts text well; it is the default choice for viewing. PDFium (Google, used in Chrome) compiled to Wasm brings Chrome’s rendering engine to any browser, with strong fidelity for difficult documents and form support. MuPDF compiled to Wasm is fast, compact and offers rich editing: merging, splitting, page manipulation, annotation and conversion, with an AGPL licence (or a commercial one) that matters for closed-source products.
All three need fonts, image decoders and a parser for a complex, often malformed format — which is exactly what makes a mature C/C++ engine compiled to Wasm attractive: decades of handling broken PDFs come with it.
Step 1 — load the engine in a worker
PDF operations are CPU-heavy and engines allocate large buffers; keep them off the main thread:
// pdf-worker.js
import * as mupdf from "mupdf"; // mupdf.js: Wasm engine with a JavaScript API
self.onmessage = async ({ data }) => {
if (data.type === "open") {
const doc = mupdf.Document.openDocument(new Uint8Array(data.bytes), "application/pdf");
self.postMessage({ type: "opened", pages: doc.countPages() });
self.doc = doc;
}
};
The main thread transfers the file’s ArrayBuffer to the worker and receives results as messages. Keep one document open per worker and free it when done; PDF
engines hold parsed objects, fonts and caches in Wasm memory.
Step 2 — render pages to images
Render the visible pages at the needed scale and return bitmaps:
function renderPage(doc, index, scale) {
const page = doc.loadPage(index);
const pixmap = page.toPixmap(mupdf.Matrix.scale(scale, scale), mupdf.ColorSpace.DeviceRGB, false);
const png = pixmap.asPNG(); // or read raw pixels into an ImageData
pixmap.destroy(); page.destroy();
return png;
}
For a viewer, render only visible pages plus one or two ahead, at the device pixel ratio, and re-render when zoom changes. Returning raw RGBA pixels and putting
them into an ImageBitmap or canvas avoids PNG encoding and decoding. For thumbnails, render at a small scale — far cheaper than full-size rendering.
Step 3 — extract text with positions
Text extraction is the input for search indexes, accessibility layers and AI features. Engines return text in reading order with bounding boxes, which lets you highlight matches on rendered pages:
const page = doc.loadPage(0);
const st = page.toStructuredText("preserve-whitespace");
const json = JSON.parse(st.asJSON()); // blocks → lines → characters with boxes
PDFs without text (scanned images) yield nothing; detect that (no text on pages with images) and offer OCR, which is a separate, heavier pipeline.
Step 4 — merge, split and reorder
Page-level operations do not require rendering. With MuPDF, create a new document and graft pages from sources:
const out = new mupdf.PDFDocument();
for (const src of sources) {
const n = src.countPages();
for (let i = 0; i < n; i++) out.graftPage(-1, src, i); // append page i of src
}
const bytes = out.saveToBuffer("compress").asUint8Array();
Splitting is the same in reverse. These operations copy page objects and their resources (fonts, images) and are fast even for large documents. PDFium and other libraries offer equivalent APIs; pure-JavaScript libraries such as pdf-lib also handle merging well without Wasm.
Step 5 — respect licensing
MuPDF’s AGPL licence requires that applications using it — including web applications that users interact with — make their source available under the AGPL, unless you buy a commercial licence. PDFium’s BSD-style licence and PDF.js’s Apache licence have no such requirement. Decide before building: for a closed product, prefer PDFium or PDF.js, or license MuPDF commercially.
Large and malicious files
PDFs can be enormous (thousands of pages, embedded high-resolution images) or deliberately malformed. Engines in Wasm are memory-safe with respect to the page — a parser bug stays inside the module — but they can still exhaust memory or spin. Process in a worker, enforce page and size limits, render lazily, and terminate and recreate the worker if a document causes a trap or exceeds a time budget. Never execute JavaScript embedded in PDFs (engines used for rendering ignore it by default).
Building a responsive viewer
A PDF viewer built on a Wasm engine needs the same care as any virtualised list. Lay out all pages using their sizes (cheap to read from the document) so the scrollbar is correct immediately, render only pages near the viewport, and cancel renders for pages that scroll out of view before they finish. Render at a low resolution first for instant feedback, then at full resolution when scrolling stops. Keep a small cache of rendered bitmaps and evict by distance from the viewport, because a 300-page document rendered at device resolution would need gigabytes. Text selection and accessibility come from an invisible text layer positioned over each page using the extracted boxes, which also enables browser find-in-page. Zoom should re-render visible pages at the new scale rather than stretching bitmaps, with the old bitmap shown scaled until the new one arrives.
Redaction is not drawing a box
Applications that “redact” PDFs by drawing black rectangles over text leave the text in the file, recoverable by anyone who copies it or opens the file in another tool — a well-known source of leaked information. True redaction removes the underlying text and image content within the area, then rewrites the page. Engines such as MuPDF and PDFium provide redaction or content-editing functions for this; verify the result by extracting text from the redacted page and checking the removed content is gone. If your feature cannot guarantee removal, do not call it redaction.
Memory per document
Engines hold parsed objects and decoded images in Wasm memory, which never shrinks. Viewing many large documents in one session can therefore grow the worker’s memory steadily. Close documents explicitly, and recycle the worker after a number of documents or when memory passes a threshold.
Expected output
The app opens a 300-page PDF in a worker in under a second, renders visible pages at device resolution as the user scrolls, extracts searchable text with boxes for highlighting, merges three uploaded PDFs into one download in about 400 ms, flags scanned pages for OCR, and uses PDFium for a closed-source deployment after a licence review.
Gotchas
- Rendering every page up front. Slow and memory-hungry. Render visible pages lazily.
- Running engines on the main thread. The UI freezes. Use a worker.
- Ignoring licences. AGPL obligations apply to MuPDF. Review before shipping.
- Assuming all PDFs have text. Scans need OCR. Detect and handle.
- Not freeing engine objects. Wasm memory grows. Destroy pages and pixmaps.
- Visual-only redaction. Text under black boxes remains extractable. Remove content and verify.
Performance note
Rendering a typical text-heavy page at 2× scale took about 25 ms with a Wasm engine on a laptop; thumbnails at 0.2× took about 4 ms; extracting structured text took about 6 ms per page.
Frequently Asked Questions
Should I just use PDF.js? For viewing, often yes; use a Wasm engine for fidelity issues, forms or editing operations.
Can I fill and flatten forms? PDFium and MuPDF support form fields; flattening bakes values into page content.
Do fonts need to be shipped? Engines include base fonts; documents normally embed theirs. Missing fonts fall back to substitutes.
Can it convert PDFs to images for upload? Yes — render pages and encode them as PNG or JPEG.
Is drawing a black box over text enough to redact it? No — the text remains in the file; use the engine’s redaction functions and verify by extracting text afterwards.
How many pages should a viewer keep rendered? Only those near the viewport, with a small bitmap cache evicted by distance; full-document rendering needs far too much memory.
Why does memory keep growing as users open documents? Engines keep parsed data in Wasm memory, which never shrinks; close documents and recycle the worker periodically.
Can pdf-lib replace a Wasm engine for merging? For merging, splitting and simple edits, yes; rendering and text extraction still need PDF.js or a Wasm engine.
Does extracted text preserve reading order? Mostly; multi-column layouts and tables can come out in unexpected order, so check structured output for complex documents.
Related
- Compressing images before upload with Wasm — encoding rendered pages.
- Building full-text search with Wasm — indexing extracted text.
- Timing out slow Wasm calls — bounding work.
- Sandboxing untrusted code with Wasm — untrusted input.
← Back to Media Processing & Codecs in Wasm