Writing a Wasm Binary by Hand in JavaScript
This page answers one task: you want to understand the WebAssembly binary format from the inside — not by reading a specification, but by producing a working
module yourself, one byte at a time, from a few lines of JavaScript. The result is a tiny emitter you can extend, and a much clearer picture of what tools like
compilers and wat2wasm produce.
Prerequisites
- [ ] Node.js or a browser console.
- [ ] Basic familiarity with WAT (to know what you are encoding).
- [ ] Optionally
wasm-toolsorwasm2watto check the output.
The shape of a module
A module is an 8-byte header followed by sections. The header is the magic number \0asm (00 61 73 6d) and the version 1 as a little-endian 32-bit integer
(01 00 00 00). Each section is one byte of section id, the section’s size in bytes as unsigned LEB128, then its contents. Lists inside sections — types,
functions, exports, bodies — are vectors: a LEB128 count followed by the items. Names are vectors of UTF-8 bytes. That is nearly all the structure you need
for a minimal module.
Step 1 — write the helpers
function uleb(n) { // unsigned LEB128
const out = [];
do { let b = n & 0x7f; n >>>= 7; if (n) b |= 0x80; out.push(b); } while (n);
return out;
}
const str = (s) => { const b = [...new TextEncoder().encode(s)]; return [...uleb(b.length), ...b]; };
const vec = (items) => [...uleb(items.length), ...items.flat()];
const section = (id, bytes) => [id, ...uleb(bytes.length), ...bytes];
const I32 = 0x7f;
LEB128 stores seven bits per byte, low bits first, with the high bit set on every byte except the last. Small values (under 128) are a single byte, which is why most counts and indices in small modules are one byte.
Step 2 — encode the sections
const typeSec = section(1, vec([[0x60, ...vec([I32, I32]), ...vec([I32])]])); // func type: (i32 i32) -> (i32)
const funcSec = section(3, vec([[0x00]])); // function 0 uses type 0
const exportSec = section(7, vec([[...str("add"), 0x00, 0x00]])); // "add", kind func (0), index 0
const body = [0x00, 0x20, 0x00, 0x20, 0x01, 0x6a, 0x0b]; // 0 locals; local.get 0; local.get 1; i32.add; end
const codeSec = section(10, vec([[...uleb(body.length), ...body]]));
0x60 introduces a function type; parameter and result lists are vectors of value types. The export entry is a name, a kind byte (0 = function, 1 = table,
2 = memory, 3 = global) and an index. The body starts with its local declarations (zero groups) and ends with 0x0b.
Step 3 — assemble, validate and run
const bytes = new Uint8Array([0x00, 0x61, 0x73, 0x6d, 0x01, 0x00, 0x00, 0x00,
...typeSec, ...funcSec, ...exportSec, ...codeSec]);
console.log(bytes.length, WebAssembly.validate(bytes)); // 41 true
const { instance } = await WebAssembly.instantiate(bytes);
console.log(instance.exports.add(2, 3)); // 5
The complete module is 41 bytes:
00 61 73 6d 01 00 00 00 header
01 07 01 60 02 7f 7f 01 7f type section
03 02 01 00 function section
07 07 01 03 61 64 64 00 00 export section
0a 09 01 07 00 20 00 20 01 6a 0b code section
If validate returns false, compile with new WebAssembly.Module(bytes) instead to get an error message with the failing byte offset.
Step 4 — extend the emitter
Add a memory and a function that stores a value:
const memSec = section(5, vec([[0x00, 0x01]])); // limits flag 0 (min only), min 1 page
// export the memory: name "mem", kind 2 (memory), index 0
// body for (func (param i32 i32) i32.store): 0 locals; local.get 0; local.get 1; i32.store align=2 offset=0; end
const storeBody = [0x00, 0x20, 0x00, 0x20, 0x01, 0x36, 0x02, 0x00, 0x0b];
Section order matters: sections must appear in increasing id order (type 1, import 2, function 3, table 4, memory 5, global 6, export 7, start 8, element 9, data count 12, code 10, data 11 — with data count placed before code). Imports go in section 2 with a module name, field name, kind and type index; imported functions take the lowest function indices, shifting the indices of defined functions.
Step 5 — check with real tools
Write the bytes to a file and disassemble them:
node emit.mjs > /dev/null # writes add.wasm in your script
wasm-tools print add.wasm
# (module (type (;0;) (func (param i32 i32) (result i32))) (func (;0;) (type 0) ... i32.add) (export "add" (func 0)))
Seeing your hand-made bytes round-trip through standard tools confirms the encoding is right.
When generating Wasm at runtime is useful
Hand-written emitters are not just a learning exercise. Generating small modules at runtime powers JIT-style optimisations in JavaScript applications: compiling a user’s formula, a filter expression or a regular expression into Wasm instead of interpreting it, then calling it millions of times. Feature-detection probes are tiny hand-encoded modules. And code generators for domain-specific languages often emit Wasm directly. For anything larger than a few functions, use a library — Binaryen’s JavaScript API or a small assembler — rather than raw bytes, but the principles are the same.
Adding an import and calling it
Imports make the emitter far more useful, because generated code can call back into JavaScript. An import entry is a module name, a field name, a kind byte and,
for functions, a type index. To import env.log of type (i32) -> () and call it from the exported function, add a second type, an import section, and use
call 0 in the body — function index 0 now refers to the import, so the defined function becomes index 1 and the export must point at 1:
const types = section(1, vec([
[0x60, ...vec([I32, I32]), ...vec([I32])], // type 0: (i32 i32) -> i32
[0x60, ...vec([I32]), ...vec([])], // type 1: (i32) -> ()
]));
const imports = section(2, vec([[...str("env"), ...str("log"), 0x00, 0x01]])); // func import of type 1
const funcs = section(3, vec([[0x00]])); // defined func uses type 0
const exports = section(7, vec([[...str("add"), 0x00, 0x01]])); // export function index 1
const body2 = [0x00, 0x20, 0x00, 0x20, 0x01, 0x6a, 0x22, 0x00, 0x10, 0x00, 0x20, 0x00, 0x0b];
The body computes the sum, uses local.tee (0x22) to keep it, calls the import with call 0 (0x10 0x00), and returns the value. Getting the index shift
wrong is the classic mistake, and validation catches it with a type mismatch.
Testing an emitter
An emitter is a compiler, and deserves compiler-style tests: round-trip each generated module through wasm-tools print or wasm2wat and compare with expected
WAT; validate every output with WebAssembly.validate; and run generated functions against a reference implementation (the JavaScript interpreter they replace)
on random inputs. Property-based tests are particularly effective here — generate random expressions, compile each to Wasm and interpret it in JavaScript, and
assert identical results.
Expected output
A 41-byte module built from JavaScript arrays validates, instantiates and returns 5 from add(2, 3); extending it with a memory section and a store function
works once sections are in the right order; and wasm-tools print shows the expected WAT for your bytes.
Gotchas
- Signed versus unsigned LEB128. Counts and sizes are unsigned;
i32.constimmediates are signed. - Wrong section order. Validation fails. Follow the id order (with data count before code).
- Forgetting the body’s local count. Every body starts with local declarations, even zero.
- Missing the final
end. Bodies end with0x0b. - Index shifts from imports. Imported functions come first.
- Untested generated code. Emitters are compilers. Validate and compare outputs against a reference.
Performance note
Generating and compiling a small specialised module at runtime takes well under a millisecond; for a filter expression evaluated over a million rows, compiling it to Wasm ran about 8× faster than interpreting its syntax tree in JavaScript.
Frequently Asked Questions
Is this how compilers emit Wasm? Conceptually yes; they use libraries with the same section and LEB128 logic.
Does CSP allow runtime-generated Wasm?
Compiling requires 'wasm-unsafe-eval' in script-src when a CSP is present.
Can I generate WAT text and assemble it instead? Yes, with an assembler library; emitting binary directly avoids the dependency.
Where are function names stored? Optionally in the custom “name” section, which you can also emit by hand.
Why does my export point at the wrong function after adding an import? Imported functions take the lowest indices, so every defined function’s index shifts by the number of imports.
How should a runtime code generator be tested? Validate every module, round-trip through a disassembler, and compare results with the JavaScript implementation on random inputs.
What does local.tee do? It stores the top of the stack into a local and leaves the value on the stack.
How do I find the byte that makes validation fail?
Compile with new WebAssembly.Module(bytes); the error message includes the offset of the failing byte.
Related
- Understanding LEB128 encoding — the integer format.
- Parsing a Wasm module header in JavaScript — the reverse direction.
- Reading the code section — function bodies.
- Feature-detecting Wasm at startup — hand-encoded probes.
← Back to Wasm Binary Format Deep Dive