Reading the Code Section
This page answers one task: you want to understand exactly how function bodies are encoded in a .wasm file — to write a tool, debug a toolchain, understand
size, or simply see what lies beneath the WAT text. You want to walk the code section byte by byte and know what each byte means.
Prerequisites
- [ ] A small module to inspect (write it in WAT and assemble with
wat2wasmorwasm-tools parse). - [ ] A hex viewer (
xxd,hexdump) andwasm-objdump -dorwasm-tools dumpfor comparison. - [ ] Familiarity with LEB128 integer encoding.
Where function bodies live
A module splits each function across two sections. The function section (id 3) lists, for each function defined in the module, the index of its type in the type section — signatures only. The code section (id 10) holds the bodies, in the same order. Imported functions appear in neither; they occupy the first function indices, so the first body in the code section belongs to function index number of imported functions.
The code section starts with the section id byte 0x0a, the section size as unsigned LEB128, and the number of bodies as unsigned LEB128. Each body is then:
its size in bytes (LEB128), its local declarations, its instructions, and a final 0x0b (end).
Step 1 — assemble a small example
(module
(memory 1)
(func $sum (param $p i32) (param $n i32) (result i32)
(local $acc i32)
(block $done
(loop $next
(br_if $done (i32.eqz (local.get $n)))
(local.set $acc (i32.add (local.get $acc) (i32.load (local.get $p))))
(local.set $p (i32.add (local.get $p) (i32.const 4)))
(local.set $n (i32.sub (local.get $n) (i32.const 1)))
(br $next)))
(local.get $acc)))
wasm-tools parse sum.wat -o sum.wasm
xxd sum.wasm
wasm-tools dump sum.wasm | sed -n '/code section/,$p'
wasm-tools dump prints each byte range with its meaning, which is ideal for checking your manual reading.
Step 2 — decode the header and body size
Find the byte 0x0a that starts the code section (sections appear in a fixed order, so it follows the function, memory, global and export sections when
present). Then read:
0a section id: code
2d section size: 45 bytes (LEB128)
01 number of bodies: 1
2b body size: 43 bytes
Body sizes let a parser skip a function without decoding it — engines use this to compile functions in parallel or lazily.
Step 3 — decode local declarations
Locals beyond the parameters are declared as run-length groups: a count of groups, then for each group a count and a type.
01 1 group
01 7f 1 local of type i32 (0x7f)
Parameters are not declared here; they come from the function’s type. Local indices continue after the parameters: $p is 0, $n is 1, $acc is 2. Grouping
keeps encodings compact for functions with many locals of the same type (compilers often emit 0a 7f — ten i32 locals — in one group).
Step 4 — decode instructions
Each instruction is an opcode byte (some use a prefix byte plus a LEB128 sub-opcode) followed by immediates. The example’s loop begins:
02 40 block, block type 0x40 (empty: no params, no results)
03 40 loop, empty block type
20 01 local.get 1 ($n)
45 i32.eqz
0d 01 br_if 1 (depth 1: the outer block)
20 02 local.get 2 ($acc)
20 00 local.get 0 ($p)
28 02 00 i32.load align=2^2 offset=0
6a i32.add
21 02 local.set 2
Memory instructions carry two immediates: the alignment hint as a power of two (02 means 4-byte alignment) and the offset (LEB128). Branch depths count
enclosing blocks outward from the instruction. Constants such as i32.const 4 encode as 41 04 — opcode plus signed LEB128.
Step 5 — check against tools
Compare your decoding with wasm-objdump -d sum.wasm, which prints offsets and decoded instructions. Mismatches usually come from LEB128 mistakes (forgetting
that values above 127 take several bytes, or that i32.const uses signed LEB128) or from miscounting branch depths.
Prefixed opcodes
Newer instructions live behind prefix bytes so the single-byte opcode space is not exhausted: 0xFC for saturating conversions, bulk memory and table
operations (0xFC 0x0A is memory.copy), 0xFD for SIMD (0xFD plus a LEB128 sub-opcode), 0xFE for atomics. Seeing these prefixes in a hex dump tells you
which feature families a function uses — useful when checking a build against a target engine’s support.
What the code section tells you about size
Since each body records its size, you can compute size per function without decoding instructions — exactly what size-analysis tools do. Large bodies point to inlining or generated code; many locals declared in many small groups hint at compiler output that could be grouped better (Binaryen’s optimiser does this).
Validating while decoding
Decoding bytes is only half of what an engine does with the code section; it also validates every instruction against types. Validation tracks a stack of
operand types and a stack of control frames: local.get 1 pushes i32 (the type of local 1), i32.eqz pops an i32 and pushes an i32, br_if 1 pops an
i32 condition and checks that the stack matches the target block’s expected results. Writing a small validator for the instructions in your example — a few
dozen lines — is the best way to understand why Wasm is safe to compile quickly: every instruction’s effect on types is known in advance, so a single forward
pass proves the function is well-typed without running it. It also explains error messages from engines and tools, such as “type mismatch in i32.add” or
“expected 1 values on stack, got 0”, which point to exactly the instruction where the type stack diverged.
Reading real compiler output
Hand-written examples are tidy; compiler output is not. Expect long runs of local.get/local.set pairs, deeply nested blocks for structured control flow
derived from arbitrary control-flow graphs, br_table for switch statements (a vector of target depths plus a default), and many prefixed instructions for
bulk memory and SIMD. Use wasm-tools dump on a real module with a function you know from the source, and match a handful of instructions to source lines; the
patterns become recognisable quickly, and the byte-level view explains size differences between builds that WAT alone can hide.
A tiny decoder as an exercise
Writing a decoder for a handful of opcodes — local.get, local.set, i32.const, i32.add, block, loop, br_if, end — takes an afternoon and
cements the format better than any reading. Test it against the example above and against wasm-objdump -d output.
Expected output
You can locate the code section in a hex dump, read the body count and each body’s size, decode local groups into local indices, decode common instructions
with their immediates — including memory alignment and offset, block types and branch depths — recognise prefixed opcodes, and confirm your reading with
wasm-tools dump or wasm-objdump -d.
Gotchas
- Forgetting imported functions. They shift function indices. Body i is function imports + i.
- Using unsigned LEB128 for constants.
i32.constimmediates are signed. - Counting branch depth inward. Depth counts outward from the branch.
- Ignoring alignment immediates. Every memory instruction has two immediates.
- Assuming one-byte opcodes. Prefixed instructions use a sub-opcode.
- Skipping validation in custom tools. Malformed bodies crash naive decoders. Bound-check every read.
Performance note
Body sizes allow engines to split compilation across threads and to compile lazily — a 10 MB code section can be validated and handed out to compile tasks without decoding every body first, which is part of why large modules start quickly.
Frequently Asked Questions
Why are locals declared separately from parameters? Parameters come from the function type; locals are extra variables, encoded compactly as groups.
Can bodies appear in a different order? No — they must match the function section’s order.
What is the 0x40 block type? The empty block type: no parameters, no results.
Where are function names? In the custom “name” section, not in the code section.
How does an engine know the body is well-typed without running it? It validates in one forward pass, tracking operand and control stacks; every instruction’s type effect is fixed.
How is a switch statement encoded?
As br_table, a vector of branch depths plus a default depth, selected by an index on the operand stack.
What does “type mismatch in i32.add” mean? The validator found operands of the wrong types on the stack at that instruction.
Should I learn from hand-written or compiler output? Both — start with a tiny WAT example, then match compiler output for a function you know to its source.
Related
- Understanding LEB128 encoding — integer encoding.
- Decoding Wasm opcodes for debugging — opcode tables.
- Reading the type section — signatures.
- Reading the data section — the next section.
← Back to Wasm Binary Format Deep Dive