Understanding Wasm Control-Flow Integrity

This page answers one question: security discussions often say WebAssembly has “built-in control-flow integrity” — that a compromised module cannot hijack execution the way native exploits do. You want to understand exactly what that means, which mechanisms provide it, and where its guarantees stop.

Prerequisites

  • [ ] Basic familiarity with WebAssembly modules, functions and linear memory.
  • [ ] A rough idea of how native exploits work (overwriting return addresses, function pointers).
  • [ ] Optionally wasm-tools or wasm2wat to inspect compiled code.

How native control-flow hijacking works

In native code, control flow is data in memory. Return addresses live on the same stack as local arrays; function pointers and vtable pointers live in heap objects; jump targets are arbitrary machine addresses. An attacker who can write out of bounds — a buffer overflow — can overwrite a return address so the function “returns” into attacker-chosen code, or overwrite a function pointer so an indirect call goes somewhere else. Techniques like return-oriented programming chain small fragments of existing code to do arbitrary work. Native control-flow integrity (CFI) schemes add checks to limit where indirect jumps may go, but they are optional, partial and add overhead.

WebAssembly’s design removes most of these opportunities at the instruction-set level, so every module gets them without opting in.

Native control flow versus Wasm control flow In native code return addresses and function pointers live in writable memory and jump targets are arbitrary addresses, so memory corruption can redirect execution. In Wasm the call stack lives outside linear memory, branches target validated structured blocks, and indirect calls go through tables with type checks, so corruption cannot redirect execution to arbitrary code. native code return addresses in writable stack jumps to any address function pointers are raw addresses hijackable WebAssembly call stack outside linear memory branches to validated blocks only indirect calls via typed tables CFI by design

The four mechanisms

Structured control flow. Wasm has no goto to arbitrary addresses. Branches (br, br_if, br_table) target enclosing block, loop and if constructs by nesting depth, and the validator checks every branch target and the types on the operand stack when the module is compiled. There are no instruction addresses in the code that a branch could be redirected to.

A protected call stack. Return addresses and the operand stack are managed by the engine, outside linear memory. A buffer overflow in linear memory — even one that overwrites everything the module can address — cannot touch a return address. (Compilers keep a shadow stack in linear memory for address-taken locals, but it holds data, not return addresses.)

Typed function tables. Indirect calls (call_indirect) do not take an address; they take an index into a table of function references, and the engine checks that the target’s type matches the expected signature before calling. Corrupting a function index can only select another function in the table with the same signature, or trap.

Validation before execution. The whole module is validated when compiled: every instruction’s operand types, every branch, every memory access’s alignment hints and every table reference. Invalid modules never run.

What call_indirect checks

(type $binop (func (param i32 i32) (result i32)))
(table 4 funcref)
(elem (i32.const 0) $add $sub $mul $div)
(func $apply (param $op i32) (param $a i32) (param $b i32) (result i32)
  (call_indirect (type $binop) (local.get $a) (local.get $b) (local.get $op)))

At run time, call_indirect checks that $op is within the table’s bounds (else trap “table index out of bounds”), that the slot is not null (else trap), and that the function’s type equals $binop (else trap “indirect call signature mismatch”). Only then does it call. An attacker who controls $op can choose among $add, $sub, $mul and $div — nothing else.

The checks inside call_indirect An indirect call takes a table index from the operand stack. The engine checks the index is within the table bounds, that the slot is not null, and that the stored function's type matches the expected signature. Only if all checks pass does the call happen; any failure traps instead of jumping elsewhere. index from operand stack maybe attacker-influenced bounds check else trap null check else trap signature check else trap call the function same type only

What remains possible

Control-flow integrity protects how code runs, not what data it computes with. Inside a module compiled from C or C++, memory-safety bugs still corrupt linear memory: overflowing a buffer can overwrite neighbouring variables, heap metadata, or function indices stored in memory. An attacker can therefore:

  • Change data that drives program logic (flags, lengths, permissions) — data-only attacks.
  • Redirect an indirect call to another function with the same signature in the table — coarse-grained CFI still allows that, and large C++ programs have many functions sharing common signatures.
  • Make the module call its imports with attacker-chosen arguments, if the module’s own logic does so.

The module cannot escape its sandbox — it cannot read the host’s memory or call functions it was not given — but within its own memory and imports, a corrupted module can misbehave in all the ways its code allows. Defence inside the module therefore still matters; see preventing memory corruption exploits inside Wasm.

How this shapes threat models

For the host — the browser or a Wasm runtime — the guarantee is strong: a module, however buggy or malicious, executes only validated code, cannot jump into the engine’s machine code, and touches only its own memory and the imports it received. For the module’s own data, the guarantee is weaker than a memory-safe language provides. That is why the most important security decision for an embedder is which imports to give a module, and the most important one for a module author is writing memory-safe code (Rust, or hardened C).

Engine bugs

These guarantees assume the engine is correct. JIT compilers are complex, and engine bugs have occasionally allowed escapes from the Wasm sandbox. Browsers add defence in depth — guard regions around linear memory, process isolation, site isolation — so that a single bug does not compromise the system. Keeping browsers and runtimes updated is part of the security model.

Narrowing indirect-call targets further

Because call_indirect only checks signatures, a large C++ program whose table holds hundreds of functions of type (i32) -> i32 gives a corrupted index many possible targets. Toolchains and authors can narrow that set. Fewer functions in the table means fewer targets: linkers only place functions in the table when their address is taken, so avoiding unnecessary function pointers shrinks it. Distinct types help: wrapping callbacks in types with unique signatures, or using the typed function references proposal (call_ref with precise types), makes mismatched targets trap. Multiple tables let a program keep unrelated function-pointer families separate. And in Rust, trait objects and closures compile to indirect calls too, but safe Rust’s memory safety means their indices cannot be corrupted in the first place — the strongest mitigation is still avoiding the memory bug that would corrupt an index.

Comparison with other sandboxes

Process sandboxes (separate OS processes with restricted system calls) isolate code with hardware and kernel mechanisms and can run anything, including JIT compilers, at the cost of heavier isolation boundaries. Language-level sandboxes (a JavaScript engine running untrusted scripts) rely on the language’s memory safety. WebAssembly sits between: arbitrary compiled code — including unsafe C — runs at near-native speed in-process, with isolation enforced by validation, bounds-checked memory and the control-flow rules described here. Many systems combine them: a Wasm sandbox for fine-grained, cheap isolation inside a process, and process isolation around the whole runtime as a second layer.

Expected output

You can explain why a buffer overflow in a Wasm module cannot overwrite a return address, read a WAT call_indirect and list the three checks it performs, and describe the data-only and same-signature attacks that remain possible inside a module compiled from unsafe code.

Gotchas

  • Treating CFI as memory safety. Linear memory can still be corrupted. Write safe code.
  • Assuming indirect calls are fully protected. Same-signature targets remain reachable.
  • Giving modules powerful imports. A corrupted module can call them. Minimise imports.
  • Ignoring the shadow stack. Address-taken locals in linear memory can be overwritten.
  • Assuming engines are bug-free. Keep browsers and runtimes patched.
  • Huge tables of same-signature functions. Many reachable targets. Take fewer addresses and use distinct types.

Performance note

The checks behind these guarantees are cheap: structured control flow and validation cost nothing at run time, and a call_indirect adds a bounds and signature check — about a nanosecond — over a direct call in current engines.

Cost of a call by kind Approximate nanoseconds per call for a direct call, a call_indirect through a table with its bounds and signature checks, and an imported JavaScript function call in a current browser engine. ns per call (approximate) direct call 1 ns call_indirect (checked) 2 ns call to JS import 10 ns

Frequently Asked Questions

Can Wasm code generate new code at runtime? Not inside the module; it can only ask the host (via imports) to compile and instantiate another module.

Is Wasm CFI equivalent to native CFI schemes? It is coarse-grained by signature, similar to some native schemes, but always on and enforced by the engine.

Does Rust need these protections? Safe Rust avoids memory corruption; CFI still protects against bugs in unsafe code and dependencies.

Can a module read the engine’s memory? Not through valid code; only engine bugs could allow it.

How can the set of reachable indirect-call targets be reduced? Take fewer function addresses, use distinct signatures or typed function references, and separate function families into different tables.

Is Wasm a replacement for process sandboxing? It complements it: Wasm gives cheap in-process isolation, and a process sandbox around the runtime adds a second layer.

Where are return addresses kept? In the engine’s own stack, outside linear memory, where module code cannot read or write them.

← Back to Browser Sandbox & Security Boundaries