Parsing a Wasm Module Header in JavaScript
This page answers one task: read the structure of a .wasm file from JavaScript — check that it is a WebAssembly module, list its sections
and their sizes, and find custom sections — with a parser small enough to paste into a page or a script.
Prerequisites
- [ ] Node 18+ or any browser.
- [ ] A
.wasmfile to test with. - [ ] Familiarity with LEB128, as in understanding LEB128 encoding.
Why parse the bytes yourself
The JavaScript API compiles modules and reflects on their imports and exports, but it does not tell you the size of each section, whether a
debug section is present, or which proposals a module uses — and it compiles the whole module first, which is wasteful if you only want to
check something. A tiny parser that reads the section table answers those questions instantly. Typical uses: rejecting an upload that is not
a WebAssembly module before passing it to a compiler, building a size report in CI, checking that a release build has no .debug_* sections,
or showing module details in a developer tool.
The format makes this easy. Every module starts with an eight-byte header, followed by sections, and every section starts with a one-byte id and a LEB128 length. A parser can walk from section to section by reading ids and lengths and skipping contents, without understanding any of them.
Step 1 — check the header
export function checkHeader(bytes) {
if (bytes.length < 8) return { ok: false, reason: "too short" };
const view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength);
const magic = view.getUint32(0, true); // little-endian
const version = view.getUint32(4, true);
if (magic !== 0x6d736100) return { ok: false, reason: "not a WebAssembly module" };
if (version === 0x1000d) return { ok: true, kind: "component" }; // component-model layer
if (version !== 1) return { ok: false, reason: `unknown version ${version}` };
return { ok: true, kind: "module" };
}
The magic number is the bytes 00 61 73 6d — \0asm — read as a little-endian 32-bit integer, 0x6d736100. Core modules have version 1.
Components produced by component-model tooling use a different version/layer value, which is worth distinguishing because they cannot be
instantiated with WebAssembly.instantiate.
Step 2 — read LEB128 values
Section sizes and many other fields are unsigned LEB128. A reader that returns the value and the number of bytes consumed is all you need:
export function readULEB(bytes, offset) {
let result = 0, shift = 0, pos = offset, byte;
do {
if (pos >= bytes.length) throw new Error("truncated LEB128");
byte = bytes[pos++];
result += (byte & 0x7f) * 2 ** shift; // avoid 32-bit overflow in bit ops
shift += 7;
} while (byte & 0x80);
return { value: result, next: pos };
}
Multiplying by 2 ** shift instead of shifting keeps values above 2³¹ correct, since JavaScript’s bitwise operators work on 32-bit signed integers.
Step 3 — walk the sections
const NAMES = ["custom", "type", "import", "function", "table", "memory", "global", "export",
"start", "element", "code", "data", "datacount", "tag"];
export function listSections(bytes) {
const header = checkHeader(bytes);
if (!header.ok) throw new Error(header.reason);
const sections = [];
let pos = 8;
while (pos < bytes.length) {
const id = bytes[pos++];
const { value: size, next } = readULEB(bytes, pos);
const start = next, end = next + size;
if (end > bytes.length) throw new Error(`section ${id} runs past end of file`);
let name = NAMES[id] ?? `unknown(${id})`;
if (id === 0) { // custom: read its name
const { value: len, next: n } = readULEB(bytes, start);
name = "custom:" + new TextDecoder().decode(bytes.subarray(n, n + len));
}
sections.push({ id, name, offset: start, size });
pos = end;
}
return sections;
}
The walk never interprets section contents except to read custom-section names. That makes it robust to new proposals: a module using features the parser has never heard of still has the same section framing.
Step 4 — use it
In Node, a size report for any module:
import { readFileSync } from "node:fs";
import { listSections } from "./wasm-sections.mjs";
const bytes = new Uint8Array(readFileSync(process.argv[2]));
const rows = listSections(bytes).map((s) => ({ section: s.name, bytes: s.size, pct: (100 * s.size / bytes.length).toFixed(1) + "%" }));
console.table(rows);
┌─────────┬─────────────────┬────────┬────────┐
│ (index) │ section │ bytes │ pct │
├─────────┼─────────────────┼────────┼────────┤
│ 0 │ type │ 1204 │ 0.2% │
│ 1 │ import │ 8872 │ 1.4% │
│ 2 │ function │ 1630 │ 0.3% │
│ 6 │ code │ 512407 │ 81.5% │
│ 7 │ data │ 47288 │ 7.5% │
│ 8 │ custom:name │ 46639 │ 7.4% │
│ 9 │ custom:producers│ 89 │ 0.0% │
└─────────┴─────────────────┴────────┴────────┘
In a browser, validate an uploaded file before compiling it:
input.addEventListener("change", async () => {
const bytes = new Uint8Array(await input.files[0].arrayBuffer());
const h = checkHeader(bytes);
if (!h.ok) return showError(`Not a WebAssembly file: ${h.reason}`);
const debug = listSections(bytes).filter((s) => s.name.startsWith("custom:.debug"));
if (debug.length) showWarning("This module contains debug information and will be larger than necessary.");
});
Step 5 — add CI checks with it
The same function makes cheap CI assertions: no debug sections in release modules, a size budget for the code section, the presence of a required custom section such as build metadata from adding custom sections to a Wasm binary. Because it reads only the framing, it runs in milliseconds even on large modules and has no dependencies to install in the CI image.
Hardening the parser for untrusted input
If the parser will see files from users — an upload form, a plugin directory — treat every length as hostile. The walk above already checks that each section ends within the file, which prevents reading past the buffer. Two more checks make it robust. Bound the LEB128 reader: a valid 32-bit LEB128 value is at most five bytes, and a reader that accepts more can be made to loop over a long run of continuation bytes, so stop after five and treat anything longer as malformed. And bound the work: a crafted file can contain millions of tiny custom sections, each of which the walker visits; cap the number of sections, or the total time spent, at a generous limit such as ten thousand.
Decode names defensively too. Custom-section names are supposed to be UTF-8, but a hostile file can put any bytes there. Decoding with
new TextDecoder("utf-8", { fatal: true }) throws on invalid sequences instead of substituting replacement characters, which lets you reject the
file cleanly rather than displaying garbled names. None of these checks costs measurable time on normal modules, and together they make the
parser safe to run on anything a user hands you — which is exactly when a quick format check is most useful.
Where a header parser stops
A section walker is deliberately shallow. It cannot tell you function names or signatures without decoding the type, function and name sections;
it cannot validate the module — a file with correct framing can still contain invalid code; and it cannot read the code section’s instructions
without a full decoder. For those questions, use WebAssembly.validate for validity, the reflection API for imports and exports, and a real
toolkit for everything else — wasm-tools or WABT, described in
inspecting modules with wasm-tools.
The value of the small parser is that it answers the structural questions instantly and anywhere.
Expected output
For a valid release module, checkHeader returns { ok: true, kind: "module" } and listSections returns the table above. For a PNG renamed to
.wasm, checkHeader returns { ok: false, reason: "not a WebAssembly module" }.
Gotchas
- Using
>>or<<for large LEB128 values. JavaScript bitwise operators truncate to 32 bits. Use multiplication for shifts beyond 28 bits. - Reading the header as big-endian. WebAssembly is little-endian throughout. Pass
truetoDataViewgetters. - Assuming known section ids only. New proposals add sections (such as
tag, id 13). Treat unknown ids as opaque and keep walking. - Off-by-one after the size. The section content starts after the LEB128 size bytes. Use the reader’s returned position.
- Trusting framing as validity. Correct section framing does not mean the module is valid. Call
WebAssembly.validatebefore relying on it.
Performance note
Walking the sections of a 12 MB module took 0.3 ms in Node, because the parser touches only a few bytes per section. Compiling the same module to answer “does it have debug sections?” took over 400 ms. For structural questions, reading the framing is three orders of magnitude cheaper.
Frequently Asked Questions
Is the section order fixed? Known sections must appear in a fixed order, each at most once; custom sections may appear anywhere. A parser that records order can flag malformed modules.
What is the datacount section?
Section 12, added by the bulk memory proposal, declares the number of data segments so engines can validate memory.init in a single pass.
Can this parser handle components? It recognises them by version and stops. Components have a different section layout that needs its own parser.
Should I use this to check uploads for safety? It checks format, not safety. Treat uploaded modules as untrusted and run them only with restricted imports, as discussed in restricting what a module can import.
Related
- How to decode .wasm files manually — the manual version of this walk.
- Reading the type section — decoding one section’s contents.
- Catching size regressions in CI — using section sizes as budgets.
- Validating binaries with wasm-validate — full validation.
← Back to Wasm Binary Format Deep Dive