Understanding the Threads Proposal

This page answers one question: “WebAssembly threads” are mentioned everywhere, yet WebAssembly has no instruction to start a thread. You want to understand what the threads proposal actually adds to the core specification, what it leaves to the host, and how those pieces become the pthreads, rayon pools and worker threads that applications use.

Prerequisites

  • [ ] Familiarity with linear memory and the basics of multithreading (data races, locks).
  • [ ] Optionally, WAT reading skills for the instruction examples.
  • [ ] A browser or runtime with threads support for experiments.

What the proposal adds

The threads proposal (now part of the standard and supported in all major engines) adds three things to WebAssembly. Shared memories: a memory declared shared can be imported by several instances at once, which may run on different threads; non-shared memories cannot. Atomic instructions: loads, stores and read-modify-write operations (add, sub, and, or, xor, xchg, cmpxchg) on 8-, 16-, 32- and 64-bit values in memory, plus a fence. Wait and notify: memory.atomic.wait32/wait64 suspend the calling thread until another thread calls memory.atomic.notify on the same address (or a timeout expires), the building block for efficient locks and condition variables.

It also defines a memory model: how reads and writes by different threads to shared memory may be observed. Atomic operations are sequentially consistent; plain (non-atomic) accesses to shared memory have weaker guarantees, and racing plain accesses produce unspecified but well-defined values (no undefined behaviour at the engine level, unlike C).

What the threads proposal specifies and what the host provides The core specification adds shared memories, atomic instructions, wait and notify, and a memory model. It does not create threads. The host provides threads by running several instances on workers that import the same shared memory. Toolchains build pthreads and thread pools on top of both. toolchain libraries pthreads, rayon, std::thread host: create threads Web Workers, worker_threads, runtime threads wait / notify memory.atomic.wait32/64, notify atomic instructions load/store/rmw/cmpxchg, fence shared memory + memory model memory (shared)

Shared memory declarations

A shared memory must declare a maximum size (so engines can reserve address space up front) and is created in JavaScript with shared: true:

(module
  (import "env" "memory" (memory 17 16384 shared))
  (func (export "inc") (param $addr i32) (result i32)
    (i32.atomic.rmw.add (local.get $addr) (i32.const 1))))
const memory = new WebAssembly.Memory({ initial: 17, maximum: 16384, shared: true });
// memory.buffer is a SharedArrayBuffer; pass `memory` to instances in several workers

Atomic accesses must be naturally aligned (a 32-bit atomic at an address divisible by 4); misaligned atomics trap.

Atomic instructions

Atomic operations make updates indivisible and ordered. i32.atomic.rmw.add adds to a value in memory and returns the old value, without another thread interleaving; i32.atomic.rmw.cmpxchg replaces a value only if it equals an expected one — the basis of lock-free algorithms and of locks themselves. Atomic loads and stores provide ordering: once a thread observes an atomic store, it also observes the writes that preceded it in the storing thread.

Wait and notify

;; wait while the value at $addr equals $expected, up to $timeout_ns (-1 = forever)
(memory.atomic.wait32 (local.get $addr) (local.get $expected) (i64.const -1))   ;; returns 0 ok, 1 not-equal, 2 timed-out
(memory.atomic.notify (local.get $addr) (i32.const 1))                           ;; wake up to 1 waiter, returns count woken

Waiting is a blocking operation; hosts may forbid it on threads that must not block. In browsers, the main thread cannot wait (it traps in Wasm, and Atomics.wait throws in JavaScript); only workers can. That restriction comes from the host, not the core proposal.

What the core spec defines versus what hosts decide The core specification defines shared memories, atomic instructions, wait and notify semantics and the memory model. Hosts decide how threads are created, whether a given thread may block, and in browsers whether shared memory is available at all, which requires cross-origin isolation. core specification shared memory type atomics + fence wait/notify + memory model portable semantics host / embedder creating threads (workers) which threads may block availability (COOP/COEP) platform policy

Threads come from the host

The proposal does not include thread.spawn. A “thread” is another instance of the module (or a related module) running on another host thread and importing the same shared memory. In browsers, the host thread is a Web Worker; in Node, a worker_threads worker; in Wasmtime, an OS thread created by the embedder (the wasi-threads proposal defines a thread-spawn import that WASI runtimes implement). Toolchains hide this: Emscripten’s pthreads create workers and instances; Rust’s wasm-bindgen-rayon builds a pool of workers; each thread needs its own stack and thread-local storage inside the shared memory, which the toolchain allocates.

Browser availability

Because shared memory enables high-resolution timers (and therefore Spectre-style attacks), browsers make SharedArrayBuffer — and shared Wasm memory — available only to cross-origin isolated pages (served with COOP and COEP headers). That deployment requirement, not engine support, is usually the obstacle to using threads on the web.

Shared-everything threads is a later proposal exploring shared tables, globals and GC objects across threads and a native way to spawn threads, addressing limits of the current design (each thread is a separate instance; only memory is shared). Check its status before relying on it.

Memory growth with shared memories

Growing a shared memory is more involved than growing a private one. All threads see the same memory, so growth must be atomic with respect to them: when one thread executes memory.grow, the new pages become visible to every thread, and the memory’s size changes for all. That is why shared memories must declare a maximum — engines reserve address space for the maximum up front, so growth never moves the memory, and pointers held by other threads stay valid. In JavaScript, the SharedArrayBuffer behind memory.buffer is growable rather than detached: existing typed-array views remain valid, but they keep their original length, so code that needs the new size must create new views (or use length-tracking views where supported). Choose maxima deliberately: too small and allocation fails under load; very large maxima reserve virtual address space, which can fail on 32-bit hosts or constrained devices.

Thread-local state

Because each thread is a separate instance, every instance has its own globals — including the stack pointer global that compilers use for the shadow stack. That is how thread-local storage works in practice: toolchains allocate a stack and a TLS block in shared memory for each new thread and initialise that instance’s globals to point at them. Static data, by contrast, lives in shared memory and is shared by all threads, which is what makes static mutable data in C or Rust a shared resource that needs synchronisation. Data segments are initialised only once — by the first instance — so later instances must not re-run initialisation that would overwrite memory other threads are using; toolchains arrange this with passive segments and a flag checked at startup.

Interaction with other features

Bulk-memory instructions (memory.copy, memory.fill) work on shared memories but are not atomic as a whole. Exceptions and traps in one thread do not automatically stop others; a trap in a worker leaves the rest running, possibly holding locks the failed thread was supposed to release.

Expected output

You can explain that the threads proposal provides shared memories, atomics, wait/notify and a memory model; that threads themselves are created by the host (workers or runtime threads) running instances that share one memory; why atomics must be aligned; why the main browser thread cannot wait; and why cross-origin isolation is required on the web.

Gotchas

  • Looking for a spawn instruction. There is none. Hosts create threads.
  • Misaligned atomics. They trap. Align atomic data.
  • Waiting on the browser main thread. It traps. Wait only in workers.
  • Plain accesses for shared data. Races give unspecified values. Use atomics or locks.
  • Forgetting isolation headers. Shared memory is unavailable. Configure COOP/COEP.
  • Re-running data initialisation in every thread. It overwrites shared state. Initialise once.

Performance note

An uncontended atomic increment costs a few nanoseconds; a contended one, bouncing a cache line between cores, can cost an order of magnitude more — the reason per-thread counters merged at the end beat a single shared counter.

Cost of an atomic increment Approximate nanoseconds per atomic add on one thread with no contention and with four threads incrementing the same counter. ns per increment (approximate) uncontended 5 ns 4 threads, same counter 60 ns

Frequently Asked Questions

Are 64-bit atomics supported? Yes — i64.atomic.* instructions operate on aligned 64-bit values.

Can non-shared memories use atomics? Yes, atomics validate on non-shared memories too; they just do not involve other threads.

Is the memory model the same as C++'s? Atomics behave like sequentially consistent C++ atomics; compilers map C++ orderings onto them.

Do WASI runtimes support threads? Several do, via wasi-threads and shared memories; check your runtime.

Why must shared memories declare a maximum? Engines reserve address space for the maximum so growth never moves memory, keeping other threads’ pointers valid.

Do typed-array views detach when a shared memory grows? No — the SharedArrayBuffer grows; existing views stay valid but keep their old length, so create new views for new pages.

How does thread-local storage work? Each thread is its own instance with its own globals, which toolchains point at a per-thread stack and TLS block in shared memory.

Does a trap in one thread stop the others? No — other threads keep running, possibly waiting on locks the failed thread held; restart the pool on traps.

Are memory.copy and memory.fill atomic on shared memory? No — they are not atomic as a whole; synchronise around them like any other shared writes.

← Back to Post-MVP Wasm Proposals in Practice