Understanding the Threads Proposal
This page answers one question: “WebAssembly threads” are mentioned everywhere, yet WebAssembly has no instruction to start a thread. You want to understand what the threads proposal actually adds to the core specification, what it leaves to the host, and how those pieces become the pthreads, rayon pools and worker threads that applications use.
Prerequisites
- [ ] Familiarity with linear memory and the basics of multithreading (data races, locks).
- [ ] Optionally, WAT reading skills for the instruction examples.
- [ ] A browser or runtime with threads support for experiments.
What the proposal adds
The threads proposal (now part of the standard and supported in all major engines) adds three things to WebAssembly. Shared memories: a memory declared
shared can be imported by several instances at once, which may run on different threads; non-shared memories cannot. Atomic instructions: loads, stores and
read-modify-write operations (add, sub, and, or, xor, xchg, cmpxchg) on 8-, 16-, 32- and 64-bit values in memory, plus a fence. Wait and notify:
memory.atomic.wait32/wait64 suspend the calling thread until another thread calls memory.atomic.notify on the same address (or a timeout expires), the
building block for efficient locks and condition variables.
It also defines a memory model: how reads and writes by different threads to shared memory may be observed. Atomic operations are sequentially consistent; plain (non-atomic) accesses to shared memory have weaker guarantees, and racing plain accesses produce unspecified but well-defined values (no undefined behaviour at the engine level, unlike C).
Shared memory declarations
A shared memory must declare a maximum size (so engines can reserve address space up front) and is created in JavaScript with shared: true:
(module
(import "env" "memory" (memory 17 16384 shared))
(func (export "inc") (param $addr i32) (result i32)
(i32.atomic.rmw.add (local.get $addr) (i32.const 1))))
const memory = new WebAssembly.Memory({ initial: 17, maximum: 16384, shared: true });
// memory.buffer is a SharedArrayBuffer; pass `memory` to instances in several workers
Atomic accesses must be naturally aligned (a 32-bit atomic at an address divisible by 4); misaligned atomics trap.
Atomic instructions
Atomic operations make updates indivisible and ordered. i32.atomic.rmw.add adds to a value in memory and returns the old value, without another thread
interleaving; i32.atomic.rmw.cmpxchg replaces a value only if it equals an expected one — the basis of lock-free algorithms and of locks themselves. Atomic
loads and stores provide ordering: once a thread observes an atomic store, it also observes the writes that preceded it in the storing thread.
Wait and notify
;; wait while the value at $addr equals $expected, up to $timeout_ns (-1 = forever)
(memory.atomic.wait32 (local.get $addr) (local.get $expected) (i64.const -1)) ;; returns 0 ok, 1 not-equal, 2 timed-out
(memory.atomic.notify (local.get $addr) (i32.const 1)) ;; wake up to 1 waiter, returns count woken
Waiting is a blocking operation; hosts may forbid it on threads that must not block. In browsers, the main thread cannot wait (it traps in Wasm, and
Atomics.wait throws in JavaScript); only workers can. That restriction comes from the host, not the core proposal.
Threads come from the host
The proposal does not include thread.spawn. A “thread” is another instance of the module (or a related module) running on another host thread and importing
the same shared memory. In browsers, the host thread is a Web Worker; in Node, a worker_threads worker; in Wasmtime, an OS thread created by the embedder
(the wasi-threads proposal defines a thread-spawn import that WASI runtimes implement). Toolchains hide this: Emscripten’s pthreads create workers and
instances; Rust’s wasm-bindgen-rayon builds a pool of workers; each thread needs its own stack and thread-local storage inside the shared memory, which the
toolchain allocates.
Browser availability
Because shared memory enables high-resolution timers (and therefore Spectre-style attacks), browsers make SharedArrayBuffer — and shared Wasm memory —
available only to cross-origin isolated pages (served with COOP and COEP headers). That deployment requirement, not engine support, is usually the obstacle to
using threads on the web.
Related proposals
Shared-everything threads is a later proposal exploring shared tables, globals and GC objects across threads and a native way to spawn threads, addressing limits of the current design (each thread is a separate instance; only memory is shared). Check its status before relying on it.
Memory growth with shared memories
Growing a shared memory is more involved than growing a private one. All threads see the same memory, so growth must be atomic with respect to them: when one
thread executes memory.grow, the new pages become visible to every thread, and the memory’s size changes for all. That is why shared memories must declare a
maximum — engines reserve address space for the maximum up front, so growth never moves the memory, and pointers held by other threads stay valid. In
JavaScript, the SharedArrayBuffer behind memory.buffer is growable rather than detached: existing typed-array views remain valid, but they keep their
original length, so code that needs the new size must create new views (or use length-tracking views where supported). Choose maxima deliberately: too small
and allocation fails under load; very large maxima reserve virtual address space, which can fail on 32-bit hosts or constrained devices.
Thread-local state
Because each thread is a separate instance, every instance has its own globals — including the stack pointer global that compilers use for the shadow stack.
That is how thread-local storage works in practice: toolchains allocate a stack and a TLS block in shared memory for each new thread and initialise that
instance’s globals to point at them. Static data, by contrast, lives in shared memory and is shared by all threads, which is what makes static mutable data in
C or Rust a shared resource that needs synchronisation. Data segments are initialised only once — by the first instance — so later instances must not
re-run initialisation that would overwrite memory other threads are using; toolchains arrange this with passive segments and a flag checked at startup.
Interaction with other features
Bulk-memory instructions (memory.copy, memory.fill) work on shared memories but are not atomic as a whole. Exceptions and traps in one thread do not
automatically stop others; a trap in a worker leaves the rest running, possibly holding locks the failed thread was supposed to release.
Expected output
You can explain that the threads proposal provides shared memories, atomics, wait/notify and a memory model; that threads themselves are created by the host (workers or runtime threads) running instances that share one memory; why atomics must be aligned; why the main browser thread cannot wait; and why cross-origin isolation is required on the web.
Gotchas
- Looking for a spawn instruction. There is none. Hosts create threads.
- Misaligned atomics. They trap. Align atomic data.
- Waiting on the browser main thread. It traps. Wait only in workers.
- Plain accesses for shared data. Races give unspecified values. Use atomics or locks.
- Forgetting isolation headers. Shared memory is unavailable. Configure COOP/COEP.
- Re-running data initialisation in every thread. It overwrites shared state. Initialise once.
Performance note
An uncontended atomic increment costs a few nanoseconds; a contended one, bouncing a cache line between cores, can cost an order of magnitude more — the reason per-thread counters merged at the end beat a single shared counter.
Frequently Asked Questions
Are 64-bit atomics supported?
Yes — i64.atomic.* instructions operate on aligned 64-bit values.
Can non-shared memories use atomics? Yes, atomics validate on non-shared memories too; they just do not involve other threads.
Is the memory model the same as C++'s? Atomics behave like sequentially consistent C++ atomics; compilers map C++ orderings onto them.
Do WASI runtimes support threads?
Several do, via wasi-threads and shared memories; check your runtime.
Why must shared memories declare a maximum? Engines reserve address space for the maximum so growth never moves memory, keeping other threads’ pointers valid.
Do typed-array views detach when a shared memory grows? No — the SharedArrayBuffer grows; existing views stay valid but keep their old length, so create new views for new pages.
How does thread-local storage work? Each thread is its own instance with its own globals, which toolchains point at a per-thread stack and TLS block in shared memory.
Does a trap in one thread stop the others? No — other threads keep running, possibly waiting on locks the failed thread held; restart the pool on traps.
Are memory.copy and memory.fill atomic on shared memory? No — they are not atomic as a whole; synchronise around them like any other shared writes.
Related
- Using Atomics for Wasm thread synchronization — atomics in practice.
- Implementing a mutex with Atomics — wait/notify locks.
- Spectre and cross-origin isolation for Wasm — why isolation.
- Building a Wasm thread pool — host-side threads.
← Back to Post-MVP Wasm Proposals in Practice