Preventing Memory Corruption Exploits Inside Wasm

This page answers one task: you ship a WebAssembly module compiled from C or C++ that parses untrusted input — images, documents, network messages — and you know the sandbox protects the host, but you also want to stop attackers from corrupting the module’s own memory to change its behaviour. You want the hardening options available for Wasm builds and a testing strategy that finds the bugs before attackers do.

Prerequisites

  • [ ] A C/C++ codebase built with Emscripten, wasi-sdk or clang targeting wasm32.
  • [ ] Untrusted input reaching the module (files, messages, user content).
  • [ ] A test and fuzzing setup you can extend.

Why linear memory is softer than native memory

Native processes enjoy several exploit mitigations: guard pages around stacks and heaps, non-executable data, address-space layout randomisation (ASLR) so attackers cannot predict addresses, and hardened allocators. Linear memory has fewer: it is one flat, readable and writable array, address 0 is valid memory, there are no guard pages inside it, and layouts are deterministic — the same module places the same data at the same addresses on every run. A buffer overflow in a heap object reliably overwrites the neighbouring object; a stack buffer overflow on the shadow stack overwrites neighbouring locals.

WebAssembly’s control-flow integrity prevents these bugs from hijacking execution into arbitrary code, but data-only attacks remain: flipping a “user is admin” flag, changing a length so later code reads further, swapping a function-table index for another of the same type, or making the module produce attacker-chosen output. For modules that make security decisions or output data others trust, that matters.

Native mitigations and their linear-memory equivalents Native code has guard pages, ASLR, non-executable memory and hardened allocators. Linear memory has no guard pages inside it and no ASLR, while code is never in linear memory so non-executable data is inherent. Equivalents come from build flags such as stack-first layout and stack protectors, hardened allocators, and testing with sanitizers. native mitigation in linear memory what to do instead guard pages none inside memory stack-first layout + stack protector ASLR deterministic layout assume attackers know addresses non-executable data inherent (code not in memory) — hardened allocator plain malloc by default checked allocator in risky modules

Step 1 — put the stack first and add canaries

Link with the stack before static data (-Wl,--stack-first, the default for some toolchains) so a shadow-stack overflow runs off the bottom of memory and traps, instead of silently overwriting globals. Compile with -fstack-protector-strong, which places canaries next to stack buffers and aborts if a function returns with a damaged canary; Emscripten’s -sSTACK_OVERFLOW_CHECK adds stack-limit checks.

emcc src/*.c -O2 -fstack-protector-strong -sSTACK_OVERFLOW_CHECK=1 -Wl,--stack-first -o parser.js

Step 2 — prefer bounds-checked code paths

Replace unchecked APIs with checked ones: memcpy with explicit size checks, strncpy/snprintf with correct sizes, _FORTIFY_SOURCE where the libc supports it, and C++ containers with .at() or hardened modes (libc++'s hardening modes add bounds checks to operator[] and iterators). In parsers, validate every length and offset read from input before use, with checked arithmetic so offset + length cannot overflow 32-bit size_t.

Step 3 — test with AddressSanitizer and fuzzing

Sanitizers find memory bugs reliably, and Emscripten supports AddressSanitizer for Wasm builds:

emcc src/*.c tests/fuzz_harness.c -g -fsanitize=address -sALLOW_MEMORY_GROWTH -o asan_tests.js
node asan_tests.js corpus/*

ASan instruments every load and store, reporting overflows and use-after-free with stack traces. Combine it with fuzzing: build the parser natively with libFuzzer and ASan (fuzzing is far faster natively), run it continuously, and re-test crashes in the Wasm build. Most memory bugs are platform-independent; the native fuzzer finds them, and the Wasm ASan build confirms the shipped configuration.

Finding memory bugs before attackers do The parser is fuzzed natively with libFuzzer and AddressSanitizer for speed. Crashing inputs are minimised and replayed against the Wasm build compiled with AddressSanitizer. Fixes are verified in both builds, and the inputs become regression tests that run in CI against the release Wasm build. native fuzzing libFuzzer + ASan crash found minimised input replay in Wasm ASan build confirm in target fix + verify both builds regression corpus CI on release build

Step 4 — use a hardened or checked allocator for risky modules

Heap overflows that corrupt allocator metadata enable stronger attacks. For modules that parse hostile input, consider an allocator that validates metadata and fills freed memory, or a debug allocator in testing that detects double frees and use-after-free. Isolating parsing of each input in a fresh instance — instantiation is cheap — means corruption from one input cannot influence the next.

Step 5 — shrink what a corrupted module can do

Assume a bug will be exploited and limit the damage: give the module only the imports it needs (a parser needs no network access), validate its outputs on the host side, run it in a worker with a time limit, and treat its results as untrusted data. A corrupted parser that can only return a structure the host validates is far less dangerous than one that can call fetch or write to storage.

Memory-safe languages

The most effective mitigation is not writing memory-unsafe code. Rust (without unsafe in parsing code) eliminates buffer overflows and use-after-free by construction; rewriting the input-handling layer in Rust, even while keeping the rest in C++, removes the most exposed attack surface. Where a rewrite is not feasible, the hardening and testing steps above reduce risk substantially.

Address zero and null pointers

In native processes, dereferencing a null pointer crashes immediately because the first page of the address space is unmapped. In linear memory, address 0 is ordinary memory: reading through a null pointer returns whatever is stored there, and writing through it silently corrupts it. Toolchains leave the first kilobyte or so unused by default (--global-base defaults to 1024 in wasm-ld), so small null-pointer offsets usually land in unused space, but nothing traps, and a null pointer plus a large offset reaches real data. Treat null checks as mandatory in C code compiled to Wasm rather than relying on crashes, enable UBSan’s null checks in test builds, and consider writing a recognisable pattern into the reserved low memory at startup and checking it periodically in debug builds — a changed pattern means some code wrote through a null or near-null pointer.

Integer overflow on 32-bit targets

Wasm32 has 32-bit pointers and size_t, while much C and C++ code is written and tested on 64-bit platforms. Size calculations that cannot overflow natively can overflow in Wasm: width * height * 4 for a large image, count * sizeof(struct) for a large array, or offset + length from a file header. An overflowed size allocates a small buffer that later code overruns — a classic exploitable pattern. Use checked arithmetic helpers (__builtin_mul_overflow in C, checked operations in Rust) for every size computed from input, and reject inputs whose sizes exceed sensible limits before allocating. Fuzzing with inputs that declare huge dimensions finds these quickly, especially in the Wasm build where they actually overflow.

Keeping hardening in place

Hardening flags disappear quietly during build refactors. Assert them in CI: check the link command or the module for stack-first layout (the stack pointer’s initial value is below the data segments), confirm stack protector symbols are present, and run the regression corpus against the release build on every change.

Expected output

The parser module is built with --stack-first and stack protectors; parsing code uses checked lengths and libc++ hardening; a nightly native fuzzing job with ASan feeds crashing inputs into a Wasm ASan replay; each untrusted file is parsed in a fresh instance in a worker with no network imports; and the host validates every structure the parser returns.

Gotchas

  • Assuming the sandbox makes C safe. Data inside the module can still be corrupted. Harden and test.
  • Default layout with stack overflows. Globals get overwritten silently. Use --stack-first.
  • Fuzzing only the Wasm build. Slow. Fuzz natively, confirm in Wasm.
  • Unchecked 32-bit size arithmetic. Overflows wrap. Use checked arithmetic.
  • Powerful imports in parsers. Exploits gain reach. Minimise imports.
  • Relying on null-pointer crashes. Address 0 is valid memory in Wasm. Check pointers explicitly.

Performance note

Stack protectors and libc++ hardening added about 3% to parsing time in release builds; the ASan build ran about 2.5× slower and is used only in testing.

Parsing time by build configuration Relative parsing time for the same corpus with a plain release build, a hardened release build with stack protectors and libc++ hardening, and an AddressSanitizer test build. relative run time plain release 1 × hardened release 1.0 × ASan test build 2.5 ×

Frequently Asked Questions

Does Wasm have ASLR? No — linear memory layouts are deterministic; assume attackers know addresses.

Can a buffer overflow escape the sandbox? Not through valid Wasm semantics; it stays within linear memory. Engine bugs are a separate concern.

Is MemorySanitizer or UBSan available? UBSan works with Emscripten; sanitizer support varies by toolchain, so check current documentation.

Should every module be parsed in a fresh instance? For hostile inputs, it is a cheap, effective isolation step.

Does reading a null pointer trap in Wasm? No — address 0 is valid linear memory; null dereferences read or corrupt data silently, so check pointers explicitly.

Why do size overflows matter more in Wasm? size_t is 32 bits, so size calculations that are safe on 64-bit platforms can wrap and allocate buffers that later code overruns.

How do I keep hardening flags from disappearing? Assert them in CI by checking the module’s layout and symbols, and run the regression corpus against every release build.

Can one corrupted input affect the next? Not if each untrusted input is parsed in a fresh instance; otherwise corrupted state persists in the same memory.

Is fuzzing the native build enough? It finds most bugs fastest, but replay crashes and the corpus against the Wasm build, where 32-bit sizes behave differently.

← Back to Browser Sandbox & Security Boundaries