Catching Memory Bugs with Emscripten Sanitizers

This page answers one task: a C or C++ module produces wrong results, crashes intermittently, or behaves differently in Wasm than natively, and you want the toolchain to tell you which line wrote where it should not.

Prerequisites

  • [ ] emsdk 3.1.x; sanitizer runtimes ship with it.
  • [ ] A reproducible input that triggers the bad behaviour, even occasionally.
  • [ ] A debug build configuration you can switch to without touching the release flags.

Why Wasm hides memory bugs

The browser sandbox protects the host from your module, not your module from itself. Inside linear memory there are no guard pages, no read-only segments and no unmapped address zero: a write through a null pointer lands on byte zero, a buffer overflow lands on whatever the allocator put next, and a read past the end of an array returns whatever bytes are there. Native builds would often crash at the first mistake. Wasm builds keep running with corrupted state, and the symptom appears far from the cause.

The sanitizer options restore the early crash. Each adds checks to the compiled code, costs speed and size, and belongs in a dedicated diagnostic build.

Emscripten's memory-checking options, cheapest to most thorough Five layers of checking from ASSERTIONS, which validates runtime API use, through STACK_OVERFLOW_CHECK and SAFE_HEAP, to UndefinedBehaviorSanitizer and AddressSanitizer, which instrument every access and catch overflows and use-after-free precisely. -sASSERTIONS=2 runtime API misuse, bad exports, unaligned HEAP access from JS — almost free -sSTACK_OVERFLOW_CHECK=2 stack cookie plus a check on every stack-pointer move -sSAFE_HEAP=1 null and out-of-range loads and stores, alignment faults — 2-3x slower -fsanitize=undefined integer overflow, bad shifts, misaligned pointers, invalid enum values -fsanitize=address overflows, use-after-free, double free with exact stack traces — 2-4x memory

Step 1 — turn on the cheap checks everywhere in debug

ASSERTIONS and STACK_OVERFLOW_CHECK are cheap enough to leave on in every debug build and in CI.

emcc -O1 -g src/*.c \
  -sASSERTIONS=2 \
  -sSTACK_OVERFLOW_CHECK=2 \
  -o build/debug/app.js

ASSERTIONS=2 catches mistakes at the boundary — calling a function that was not exported, reading HEAPU8 after memory growth without refreshing the view, passing a non-integer where a pointer was expected. STACK_OVERFLOW_CHECK=2 places a cookie at the end of the shadow stack and checks the stack pointer on every function entry, which is how you find the recursion that silently overwrote your static data. The shadow stack itself is explained in understanding the shadow stack in linear memory.

Step 2 — use SAFE_HEAP for null and wild pointers

emcc -O1 -g src/*.c -sSAFE_HEAP=1 -sASSERTIONS=2 -o build/safeheap/app.js

SAFE_HEAP rewrites every load and store into a call that checks the address first. A store to address 0, to an address past the current memory size, or a misaligned access that the source claimed was aligned aborts with a stack trace:

Aborted(segmentation fault storing 4 bytes at address 0)
    at abort (app.js:648:11)
    at SAFE_HEAP_STORE_i32_4_4 (app.wasm:0x4c21)
    at parse_header (src/parse.c:88:17)
    at main (src/main.c:31:9)

It does not detect overflows that stay inside valid memory — a write one past the end of a heap block lands in a valid address. That is AddressSanitizer’s job.

Step 3 — use AddressSanitizer for overflows and use-after-free

emcc -O1 -g -fsanitize=address src/*.c \
  -sALLOW_MEMORY_GROWTH=1 \
  -sINITIAL_MEMORY=256MB \
  -o build/asan/app.js

ASan puts poisoned redzones around every allocation, keeps a quarantine of freed blocks, and checks a shadow map on every access. Memory use rises sharply — budget at least double the normal heap and allow growth.

// src/bug.c — an off-by-one that native code would probably survive
char *dup_name(const char *src) {
  size_t n = strlen(src);
  char *dst = malloc(n);          // forgot the terminator
  memcpy(dst, src, n + 1);        // writes one byte past the block
  return dst;
}
==42==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x0201a3f5
WRITE of size 6 at 0x0201a3f5 thread T0
    #0 0x1c38f in __asan_memcpy+0x1c38f (app.wasm+0x1c38f)
    #1 0x4a1e in dup_name src/bug.c:4:3
    #2 0x50b2 in main src/main.c:12:16
0x0201a3f5 is located 0 bytes after 5-byte region [0x0201a3f0,0x0201a3f5)
allocated by thread T0 here:
    #0 0x1b8f0 in malloc+0x1b8f0 (app.wasm+0x1b8f0)
    #1 0x4a02 in dup_name src/bug.c:3:15

The report names the bad write, the allocation, and the distance between them. That precision is the reason to bother with ASan’s cost.

How AddressSanitizer catches the off-by-one A five-byte heap block is surrounded by poisoned redzones. The memcpy writes six bytes; the sixth lands in the right redzone, which the shadow map marks as poisoned, so the access is reported immediately. heap block from malloc(5) with ASan redzones left redzone 5 bytes owned right redzone (6th byte lands here) next allocation 0x0201a3e0 0x0201a3f0 0x0201a3f5 0x0201a410

Step 4 — add UBSan for the arithmetic bugs

emcc -O1 -g -fsanitize=undefined src/*.c -o build/ubsan/app.js

UndefinedBehaviorSanitizer catches signed overflow, shifts past the type width, misaligned pointer casts and similar bugs that compile to something in Wasm but rarely the something you meant. It is cheaper than ASan and composes with it: -fsanitize=address,undefined in one build is common.

Step 5 — run the diagnostic builds where they can fail loudly

Sanitized builds work in the browser, but the fastest loop is usually Node, where the report goes straight to the terminal and the process exit code fails the CI job:

emcc -O1 -g -fsanitize=address,undefined src/*.c tests/run_all.c \
  -sENVIRONMENT=node -sEXIT_RUNTIME=1 -o build/asan/tests.js
node build/asan/tests.js

EXIT_RUNTIME makes the program run its exit handlers, which is when LeakSanitizer reports leaks; see detecting leaks in Emscripten with LeakSanitizer for that half of the story.

Choosing the build for the symptom

Not every bug needs the heaviest tool. The symptom usually points at the right configuration, and starting with the cheapest one that can catch the bug keeps the loop fast.

Picking a diagnostic build from the symptom A decision tree from the observed symptom to the cheapest Emscripten checking mode that catches it — stack checks for corruption after deep recursion, SAFE_HEAP for crashes near address zero, AddressSanitizer for corruption that moves with allocation patterns, UBSan for wrong arithmetic results. What does the bug look like? breaks after deep recursion STACK_OVERFLOW_CHECK=2 cheap; leave it on in debug values near address 0 SAFE_HEAP=1 null and wild pointers moves when allocs change -fsanitize=address overflow, use-after-free arithmetic is off -fsanitize=undefined overflow, shifts, casts

A bug whose symptom moves when you add a printf or change an unrelated allocation is the classic sign of a heap overflow or use-after-free: the corruption lands on whatever happens to be adjacent, and adjacency changes with every edit. Go straight to ASan for those. A crash that always happens in the same place with a small address in the message is almost always a null or freed pointer, and SAFE_HEAP will find it faster.

Keep one more cheap trick in reserve for builds you cannot rebuild: -sASSERTIONS=2 alone catches a surprising number of boundary bugs, especially stale HEAP* views in JavaScript glue after memory growth — which look exactly like memory corruption from the C side but are entirely a JavaScript mistake.

Expected output

A clean run prints the test results and nothing else — sanitizers are silent unless they find something. A build where ASan is genuinely active can be confirmed by looking for its runtime in the binary:

wasm-objdump -x build/asan/tests.wasm | grep -c '__asan_'

A count in the hundreds means the instrumentation is present; zero means the flag did not reach the link.

Gotchas

  • Aborted(Cannot enlarge memory arrays) under ASan. The shadow map and redzones need room. Raise INITIAL_MEMORY and allow growth.
  • SAFE_HEAP reports alignment faults in correct code. Code that casts a char * into an int * and dereferences it is relying on unaligned access being legal, which it is in Wasm. Either fix the cast or silence alignment checks; do not ignore the report blindly.
  • Mixing sanitized and unsanitized objects. A static library built without -fsanitize=address will not report its own bugs and may produce false positives at the boundary. Rebuild dependencies, including ports, in the diagnostic configuration.
  • Shipping a sanitized build. Easy to do by accident in a shared CI cache. Keep the output directories separate and check for __asan_ symbols in the release artifact.

Wiring the builds into CI

Diagnostic builds pay off most when they run on every change rather than when someone remembers them. A matrix job that builds the test runner three ways keeps the cost contained and the coverage broad:

jobs:
  sanitize:
    strategy:
      matrix:
        mode: [asan, ubsan, safeheap]
    steps:
      - uses: actions/checkout@v4
      - uses: mymindstorm/setup-emsdk@v14
        with: { version: 3.1.61 }
      - run: make test-${{ matrix.mode }}
SAN_asan     = -fsanitize=address -sINITIAL_MEMORY=256MB -sALLOW_MEMORY_GROWTH=1
SAN_ubsan    = -fsanitize=undefined
SAN_safeheap = -sSAFE_HEAP=1

test-%:
	emcc -O1 -g $(SAN_$*) -sASSERTIONS=2 -sENVIRONMENT=node -sEXIT_RUNTIME=1 \
	     src/*.c tests/run_all.c -o build/$*/tests.js
	node build/$*/tests.js

Each mode fails the job with a non-zero exit when it finds something, and the report lands in the job log with file and line numbers because of -g. Run the three in parallel; even ASan on a modest test suite finishes in a minute or two.

Performance note

On an image-decoding benchmark, the release build ran in 48 ms, SAFE_HEAP in 131 ms, UBSan in 70 ms and ASan in 162 ms with peak memory up from 38 MB to 109 MB. Those numbers make ASan impractical for anything interactive and perfectly reasonable for a test suite — which is where it belongs. The two-line regression test that reproduces the bug afterwards runs in the fast suite forever.

Frequently Asked Questions

Do sanitizers work with -O2 or -O3? They do, but optimizations can merge or remove the accesses being checked and make reports harder to read. -O1 -g is the usual compromise between speed and clear stack traces.

Can I use ASan with pthreads? Yes; build with -pthread alongside the sanitizer flags and expect a higher memory cost per thread.

Is there an equivalent for Rust? Rust’s safe code does not need it, and unsafe code is best checked with Miri or sanitizers on the native target. The differential testing approach catches Wasm-only differences.

← Back to C/C++ to Wasm with Emscripten