Reading the Data Section

This page answers one task: you want to understand what the data section of a .wasm file contains — how segments are encoded, where they land in memory, and what your compiler put there — to analyse size, debug wrong static data, or write tools that read or modify modules.

Prerequisites

  • [ ] A module with static data (string literals, constant tables) — compiler output or a WAT example.
  • [ ] wasm-objdump -x -j Data, wasm-tools dump and a hex viewer.
  • [ ] Familiarity with LEB128 and constant expressions.

What data segments are

Linear memory starts zeroed. Data segments are byte strings that initialise parts of it. The data section (id 11) lists segments; each is either active — copied into a memory at a given offset during instantiation — or passive — held by the module until code copies it with memory.init (from the bulk-memory proposal) and then discarded with data.drop. Compilers place string literals, constant arrays, initialised global variables and vtables in data segments. Zero-initialised data (BSS) needs no segment, because memory starts zeroed.

A separate data count section (id 12), placed before the code section, declares the number of data segments when code uses memory.init or data.drop, so validators can check segment indices in function bodies before reaching the data section, which comes after the code.

Data segment kinds Active segments in memory zero use flag 0 followed by an offset expression and bytes. Passive segments use flag 1 followed directly by bytes and are copied later with memory.init. Active segments for a specific memory index use flag 2 with an explicit memory index, needed with multiple memories. flag kind encoded as used for 0 active (memory 0) offset expr, bytes static data at fixed addresses 1 passive bytes only lazy init, threads, PIC 2 active (explicit memory) memidx, offset expr, bytes multiple memories

Step 1 — list segments with tools

wasm-objdump -x -j Data app.wasm
# Data[3]:
#  - segment[0] memory=0 size=1520 - init i32=1048576
#  - segment[1] memory=0 size=24 - init i32=1050096
#  - segment[2] passive size=4096
wasm-objdump -s -j Data app.wasm | head -20     # hex + ASCII contents

Offsets show where active segments land; contents in the ASCII column reveal strings — panic messages, format strings, file paths, error text — that often account for surprising size.

Step 2 — decode an active segment by hand

From a small WAT example:

(module
  (memory 1)
  (data (i32.const 16) "hello"))

The data section encodes as:

0b          section id: data
0b          section size: 11 bytes
01          1 segment
00          flags: active, memory 0
41 10 0b    offset expression: i32.const 16, end
05          5 bytes follow
68 65 6c 6c 6f   "hello"

The offset is a constant expression terminated by 0b (end) — usually i32.const, or global.get of an imported global in position-independent modules, where the loader supplies __memory_base.

Decoding one active data segment Read the flags value to learn whether the segment is active or passive and which memory it targets. For active segments, decode the offset constant expression up to its end byte. Read the byte count as LEB128 and then that many bytes, which instantiation copies into memory at the offset. flags 0 active, 1 passive, 2 memidx offset expr i32.const 16, end byte count LEB128 bytes "hello" instantiation copies memory[16..21]

Step 3 — understand passive segments

(data $table "\01\02\03\04")                       ;; passive: no offset
(func $load_table (param $dst i32)
  (memory.init $table (local.get $dst) (i32.const 0) (i32.const 4))
  (data.drop $table))

Passive segments are not copied at instantiation. Toolchains use them for threads (only the first thread should initialise shared memory, so initialisation is a function guarded by a flag), for position-independent code, and to defer copying large data until needed. After data.drop, the segment’s bytes can be freed by the engine.

Step 4 — relate segments to source

Map segment contents back to source: string literals appear verbatim; numeric tables appear as little-endian bytes (a static TABLE: [u32; 4] = [1, 2, 3, 4] appears as 01 00 00 00 02 00 00 00 …); vtables and function-pointer tables contain table indices. The linker’s map file (wasm-ld --Map) gives symbol names and addresses, so you can find which symbol occupies a given address range in a segment.

Step 5 — reduce data size

Large data sections slow downloads and instantiation (active segments are copied every time). Common reductions: strip panic messages and file paths (panic = "abort" and avoiding formatting), remove unused Unicode or locale tables via dependency features, move large assets out to separately fetched files, and convert rarely used large data to passive segments loaded on demand. wasm-opt merges adjacent segments and removes zero bytes at segment edges, which helps a little.

Multiple memories

With the multi-memory proposal, flag 2 encodes an explicit memory index, so segments can target a memory other than the first. Toolchains rarely emit this today, but hand-written or specialised modules may.

Data segments in threaded and dynamically linked builds

In threaded builds, every thread is a separate instance importing the same shared memory. If data segments were active, every new thread’s instantiation would copy them again — overwriting static variables other threads were already using. Toolchains therefore emit passive segments for shared-memory builds and a small initialisation function that the first instance runs, guarded by an atomic flag, to memory.init each segment once; later instances skip it. When reading a threaded module’s data section, all segments appearing as passive is expected, not a mistake. Dynamically linked side modules use passive segments or active segments with global.get $__memory_base offsets for a similar reason: their data address is only known at load time. Seeing global.get in an offset expression is the signature of position-independent code.

Reading segments programmatically

For tooling — size reports, checks that no debug strings ship, extracting embedded assets — parse the data section in a few lines with a library rather than by hand. In JavaScript, WebAssembly.Module.customSections only reads custom sections, so use a parser such as the one in wasm-tools’ JavaScript bindings, @webassemblyjs, or a minimal hand-written reader that walks section headers and decodes segments as shown above. In Rust, the wasmparser crate exposes data segments with their kinds and byte ranges. A CI check that scans release modules’ data for patterns like source paths (/home/, src/) or development messages catches leaked build details before release.

Data and compression

Data segments often compress better than code — text and tables are repetitive — so their share of the compressed download is smaller than their share of raw bytes. Measure compressed contribution when deciding what to move out; a large but highly compressible table may matter less for download than for instantiation, where every raw byte is copied.

Checking what ships

A short release check pays for itself: dump the data section of the production module, search for strings that should not be there — absolute build paths, internal hostnames, debug messages, test fixtures — and fail the build if any appear. Static data is easy to inspect, and leaked strings are both a size cost and, occasionally, an information disclosure.

A quick experiment

Assemble the hello example above, change the offset to (i32.const 65536) with a one-page memory, and try to instantiate it: the out-of-bounds segment makes instantiation fail, showing the bounds check that protects memory initialisation.

Expected output

You can list a module’s data segments and their offsets, decode an active segment’s flags, offset expression and bytes by hand, explain passive segments and the data count section, find strings and tables in segment contents and relate them to source symbols, and pick strategies to shrink data that is copied at instantiation.

Gotchas

  • Expecting BSS in the data section. Zeroed data needs no segment.
  • Forgetting the end byte in offset expressions. Every constant expression ends with 0b.
  • Overlooking passive segments. They are not copied at instantiation. Look for memory.init.
  • Missing data count section. Required when code uses memory.init/data.drop.
  • Large active segments. Copied every instantiation. Move or defer big data.
  • Leaked build paths and debug strings. They ship in data segments. Scan release modules in CI.

Performance note

For a module whose data section held a 2.1 MB Unicode table, enabling a dependency feature that dropped the table reduced instantiation from 9 ms to 2 ms and the compressed download by 380 KB.

Data section size before and after trimming Kilobytes in the data section of a module before trimming, after stripping panic messages and paths, and after also dropping an unused Unicode table through a dependency feature. KB of data segments initial 2,450 KB no panic strings, paths 2,280 KB + Unicode table dropped 160 KB

Frequently Asked Questions

Can data segments overlap? Active segments can overlap; later ones overwrite earlier ones during instantiation.

What happens if a segment is out of bounds? Instantiation fails with an error (in current semantics) before any segment is written partially.

Where do initialised globals live? Language-level initialised globals live in data segments; Wasm globals are separate.

Can I add data to a module after building? Tools like wasm-tools and Binaryen can, but adjusting addresses safely requires knowing the layout.

Why are all segments passive in my threaded build? So only the first thread initialises shared memory; later instances must not copy segments again.

What does global.get in an offset expression mean? Position-independent code: the data’s address is supplied at load time through an imported base global.

How can I read data segments from a script? Use a parser such as Rust’s wasmparser or wasm-tools bindings; customSections only covers custom sections.

Do data segments compress well? Usually better than code, so measure their compressed share before moving them out for download reasons.

Does a segment at the end of memory fail? Only if any byte falls outside the memory’s initial size; the whole instantiation then fails.

← Back to Wasm Binary Format Deep Dive