Reading the Data Section
This page answers one task: you want to understand what the data section of a .wasm file contains — how segments are encoded, where they land in memory, and
what your compiler put there — to analyse size, debug wrong static data, or write tools that read or modify modules.
Prerequisites
- [ ] A module with static data (string literals, constant tables) — compiler output or a WAT example.
- [ ]
wasm-objdump -x -j Data,wasm-tools dumpand a hex viewer. - [ ] Familiarity with LEB128 and constant expressions.
What data segments are
Linear memory starts zeroed. Data segments are byte strings that initialise parts of it. The data section (id 11) lists segments; each is either active —
copied into a memory at a given offset during instantiation — or passive — held by the module until code copies it with memory.init (from the bulk-memory
proposal) and then discarded with data.drop. Compilers place string literals, constant arrays, initialised global variables and vtables in data segments.
Zero-initialised data (BSS) needs no segment, because memory starts zeroed.
A separate data count section (id 12), placed before the code section, declares the number of data segments when code uses memory.init or data.drop, so
validators can check segment indices in function bodies before reaching the data section, which comes after the code.
Step 1 — list segments with tools
wasm-objdump -x -j Data app.wasm
# Data[3]:
# - segment[0] memory=0 size=1520 - init i32=1048576
# - segment[1] memory=0 size=24 - init i32=1050096
# - segment[2] passive size=4096
wasm-objdump -s -j Data app.wasm | head -20 # hex + ASCII contents
Offsets show where active segments land; contents in the ASCII column reveal strings — panic messages, format strings, file paths, error text — that often account for surprising size.
Step 2 — decode an active segment by hand
From a small WAT example:
(module
(memory 1)
(data (i32.const 16) "hello"))
The data section encodes as:
0b section id: data
0b section size: 11 bytes
01 1 segment
00 flags: active, memory 0
41 10 0b offset expression: i32.const 16, end
05 5 bytes follow
68 65 6c 6c 6f "hello"
The offset is a constant expression terminated by 0b (end) — usually i32.const, or global.get of an imported global in position-independent modules,
where the loader supplies __memory_base.
Step 3 — understand passive segments
(data $table "\01\02\03\04") ;; passive: no offset
(func $load_table (param $dst i32)
(memory.init $table (local.get $dst) (i32.const 0) (i32.const 4))
(data.drop $table))
Passive segments are not copied at instantiation. Toolchains use them for threads (only the first thread should initialise shared memory, so initialisation is a
function guarded by a flag), for position-independent code, and to defer copying large data until needed. After data.drop, the segment’s bytes can be freed
by the engine.
Step 4 — relate segments to source
Map segment contents back to source: string literals appear verbatim; numeric tables appear as little-endian bytes (a static TABLE: [u32; 4] = [1, 2, 3, 4]
appears as 01 00 00 00 02 00 00 00 …); vtables and function-pointer tables contain table indices. The linker’s map file (wasm-ld --Map) gives symbol names and
addresses, so you can find which symbol occupies a given address range in a segment.
Step 5 — reduce data size
Large data sections slow downloads and instantiation (active segments are copied every time). Common reductions: strip panic messages and file paths
(panic = "abort" and avoiding formatting), remove unused Unicode or locale tables via dependency features, move large assets out to separately fetched files,
and convert rarely used large data to passive segments loaded on demand. wasm-opt merges adjacent segments and removes zero bytes at segment edges, which helps
a little.
Multiple memories
With the multi-memory proposal, flag 2 encodes an explicit memory index, so segments can target a memory other than the first. Toolchains rarely emit this today, but hand-written or specialised modules may.
Data segments in threaded and dynamically linked builds
In threaded builds, every thread is a separate instance importing the same shared memory. If data segments were active, every new thread’s instantiation would
copy them again — overwriting static variables other threads were already using. Toolchains therefore emit passive segments for shared-memory builds and a
small initialisation function that the first instance runs, guarded by an atomic flag, to memory.init each segment once; later instances skip it. When
reading a threaded module’s data section, all segments appearing as passive is expected, not a mistake. Dynamically linked side modules use passive segments or
active segments with global.get $__memory_base offsets for a similar reason: their data address is only known at load time. Seeing global.get in an offset
expression is the signature of position-independent code.
Reading segments programmatically
For tooling — size reports, checks that no debug strings ship, extracting embedded assets — parse the data section in a few lines with a library rather than by
hand. In JavaScript, WebAssembly.Module.customSections only reads custom sections, so use a parser such as the one in wasm-tools’ JavaScript bindings,
@webassemblyjs, or a minimal hand-written reader that walks section headers and decodes segments as shown above. In Rust, the wasmparser crate exposes data
segments with their kinds and byte ranges. A CI check that scans release modules’ data for patterns like source paths (/home/, src/) or development
messages catches leaked build details before release.
Data and compression
Data segments often compress better than code — text and tables are repetitive — so their share of the compressed download is smaller than their share of raw bytes. Measure compressed contribution when deciding what to move out; a large but highly compressible table may matter less for download than for instantiation, where every raw byte is copied.
Checking what ships
A short release check pays for itself: dump the data section of the production module, search for strings that should not be there — absolute build paths, internal hostnames, debug messages, test fixtures — and fail the build if any appear. Static data is easy to inspect, and leaked strings are both a size cost and, occasionally, an information disclosure.
A quick experiment
Assemble the hello example above, change the offset to (i32.const 65536) with a one-page memory, and try to instantiate it: the out-of-bounds segment makes
instantiation fail, showing the bounds check that protects memory initialisation.
Expected output
You can list a module’s data segments and their offsets, decode an active segment’s flags, offset expression and bytes by hand, explain passive segments and the data count section, find strings and tables in segment contents and relate them to source symbols, and pick strategies to shrink data that is copied at instantiation.
Gotchas
- Expecting BSS in the data section. Zeroed data needs no segment.
- Forgetting the end byte in offset expressions. Every constant expression ends with
0b. - Overlooking passive segments. They are not copied at instantiation. Look for
memory.init. - Missing data count section. Required when code uses
memory.init/data.drop. - Large active segments. Copied every instantiation. Move or defer big data.
- Leaked build paths and debug strings. They ship in data segments. Scan release modules in CI.
Performance note
For a module whose data section held a 2.1 MB Unicode table, enabling a dependency feature that dropped the table reduced instantiation from 9 ms to 2 ms and the compressed download by 380 KB.
Frequently Asked Questions
Can data segments overlap? Active segments can overlap; later ones overwrite earlier ones during instantiation.
What happens if a segment is out of bounds? Instantiation fails with an error (in current semantics) before any segment is written partially.
Where do initialised globals live? Language-level initialised globals live in data segments; Wasm globals are separate.
Can I add data to a module after building?
Tools like wasm-tools and Binaryen can, but adjusting addresses safely requires knowing the layout.
Why are all segments passive in my threaded build? So only the first thread initialises shared memory; later instances must not copy segments again.
What does global.get in an offset expression mean? Position-independent code: the data’s address is supplied at load time through an imported base global.
How can I read data segments from a script?
Use a parser such as Rust’s wasmparser or wasm-tools bindings; customSections only covers custom sections.
Do data segments compress well? Usually better than code, so measure their compressed share before moving them out for download reasons.
Does a segment at the end of memory fail? Only if any byte falls outside the memory’s initial size; the whole instantiation then fails.
Related
- Reading the code section — function bodies.
- Using bulk memory operations — memory.init and data.drop.
- Reducing instantiation cost of large modules — segment costs.
- Reading wasm-ld map files — symbols to addresses.
← Back to Wasm Binary Format Deep Dive