Porting pthreads Code with Emscripten

This page answers one task: take C or C++ code that already uses POSIX threads and get it running, multi-threaded, in a browser through Emscripten — without deadlocking the page on its first join.

Prerequisites

  • [ ] emsdk 3.1.50 or newer, activated in the shell (emcc -v prints the version).
  • [ ] C or C++ source that uses pthread_create, std::thread or OpenMP-style worker threads.
  • [ ] A dev server that can send Cross-Origin-Opener-Policy and Cross-Origin-Embedder-Policy headers.
  • [ ] Chrome, Firefox or Safari 16.4+ for testing; all three support shared WebAssembly memory once isolated.

How Emscripten maps a pthread onto the web

A native pthread is an operating-system thread sharing one address space. The browser has no such primitive. What it has is the Web Worker — a separate JavaScript event loop with its own global scope — and SharedArrayBuffer, a block of memory more than one worker can see. Emscripten builds threads out of those two pieces: the module’s linear memory becomes a shared memory, every pthread becomes a worker that instantiates the same module against it, and the synchronization primitives in libc are implemented with the memory.atomic.wait32 and memory.atomic.notify instructions.

The consequence is that a “thread” costs a worker: a few megabytes of JavaScript heap, a script fetch and a module instantiation. Creating one is not the microsecond operation C programmers expect, which is why Emscripten keeps a pool of pre-started workers.

A pthread in the browser is a worker sharing one memory The main thread creates a pthread through libc, Emscripten hands the start routine to a pooled Web Worker, the worker instantiates the same module against the shared memory, and synchronization runs through atomic wait and notify instructions. pthread_create() main thread calls libc worker from pool pre-started or spawned same module instance imports the shared memory start routine runs in parallel with main atomics wait/notify mutexes, condvars, join Every thread shares one SharedArrayBuffer — which is why the page must be cross-origin isolated.

The shared memory is also why threads need cross-origin isolation. Browsers only expose SharedArrayBuffer to pages served with the right headers, as explained in configuring COOP/COEP headers; without them the module fails at startup with a clear but easily misread error.

The flag must appear on every compile and on the link. A translation unit compiled without it uses non-atomic loads and stores; linking it into a threaded build produces a module that validates but races.

emcc -O2 -pthread -c worker_pool.c -o worker_pool.o
emcc -O2 -pthread -c main.c        -o main.o
emcc -O2 -pthread worker_pool.o main.o \
  -sPTHREAD_POOL_SIZE=4 \
  -sINITIAL_MEMORY=64MB \
  -sEXPORTED_RUNTIME_METHODS=ccall \
  -o app.js

The output is app.js (the loader), app.wasm (the module) and, in older emsdk releases, a separate app.worker.js. Newer releases fold the worker bootstrap into app.js itself, which simplifies hosting: there is one script for both the main thread and the workers.

INITIAL_MEMORY matters more than usual. A shared memory must declare a maximum, and growing it is slower than growing an unshared one because every worker has to observe the new size. Sizing the initial memory for the expected peak avoids growth on the hot path; -sALLOW_MEMORY_GROWTH still works with threads but carries a measurable cost on every access from JavaScript, since views must be re-checked.

Step 2 — size the worker pool

PTHREAD_POOL_SIZE pre-creates workers while the module loads, so pthread_create can hand a start routine to an idle worker immediately. Without a pool, creation has to spawn a worker and wait for it to load — and that wait needs the main thread’s event loop to turn, which is where ports deadlock.

// main.c — creates four workers and joins them
#include <pthread.h>
#include <stdio.h>

static void *work(void *arg) {
  int id = (int)(intptr_t)arg;
  long sum = 0;
  for (long i = 0; i < 50000000; i++) sum += i % (id + 3);
  printf("thread %d done: %ld\n", id, sum);
  return NULL;
}

int main(void) {
  pthread_t t[4];
  for (int i = 0; i < 4; i++) pthread_create(&t[i], NULL, work, (void *)(intptr_t)i);
  for (int i = 0; i < 4; i++) pthread_join(t[i], NULL);
  return 0;
}

Match the pool to the number of threads the program creates early. For code that sizes itself from the hardware, -sPTHREAD_POOL_SIZE=navigator.hardwareConcurrency is accepted as a JavaScript expression, though it is worth capping: a sixteen-core desktop will happily start sixteen workers it never uses.

Startup cost grows with the worker pool Time from navigation to the module being ready, for pool sizes of zero, two, four, eight and sixteen workers, measured on a mid-range laptop with a 1.2 MB module. Each pre-started worker adds an instantiation. time to module ready (ms, lower is better) pool size 0 118 ms pool size 2 164 ms pool size 4 205 ms pool size 8 291 ms pool size 16 468 ms Each pooled worker instantiates the same module, so a large pool trades startup time for instant thread creation later.

Step 3 — never block the main thread on a thread that has not started

The single most common failure in a pthreads port is a pthread_join or a mutex wait on the browser’s main thread. The main thread is not allowed to use Atomics.wait; Emscripten works around this with a busy-wait, which burns CPU but functions — unless the thread being waited for needs the main thread to do something first, such as spawning its worker. Then nothing progresses and the tab hangs.

There are two robust fixes. The first is to keep main() off the main thread entirely:

emcc -O2 -pthread main.o worker_pool.o \
  -sPROXY_TO_PTHREAD \
  -sPTHREAD_POOL_SIZE=4 \
  -o app.js

PROXY_TO_PTHREAD runs main() in a worker, leaving the browser thread free to service the runtime. Blocking calls inside main() become legal because they happen in a worker. The cost is that anything touching the DOM must be proxied back, which Emscripten does for its own library functions and you do with emscripten_async_run_in_main_runtime_thread for your own.

The second fix, when main() must stay put, is to restructure so the main thread never waits: start the threads, return to the event loop, and have the last thread post a completion message.

Why a join on the main thread can deadlock The main thread calls pthread_create with an empty pool, then immediately joins. The new worker cannot be spawned because spawning needs the main event loop, which is stuck in the busy-wait, so the join never returns. main thread Emscripten runtime new worker pthread_create (pool empty) queue a worker spawn for the next event-loop turn pthread_join busy-waits; event loop never turns spawn never happens completion never sent Either pre-start the worker with PTHREAD_POOL_SIZE or move main() to a worker with PROXY_TO_PTHREAD.

Step 4 — serve it with isolation headers

A threaded build will not start on a page that is not cross-origin isolated. For local testing, any of the servers in local development server configurations works once the two headers are set:

npx http-server . -p 8080 \
  -H "Cross-Origin-Opener-Policy: same-origin" \
  -H "Cross-Origin-Embedder-Policy: require-corp"

Then confirm in the console before chasing any other bug:

console.log("isolated:", self.crossOriginIsolated);   // must be true
console.log("SAB:", typeof SharedArrayBuffer);          // must be "function"

Every subresource on an isolated page must opt in to being embedded, which in practice means same-origin assets or third-party assets served with Cross-Origin-Resource-Policy. Fonts, analytics scripts and images from other origins are the usual casualties.

Step 5 — send DOM work back to the main thread

Workers have no document. Code that ran on a desktop thread and called into a UI toolkit now needs an explicit hop back to the browser thread for anything visible. Emscripten provides a family of proxying calls for this; the synchronous variant blocks the calling worker until the main thread has run the function, which is safe precisely because the caller is not the main thread.

#include <emscripten/threading.h>

static void set_status(int pct) {
  EM_ASM({ document.getElementById('status').textContent = $0 + '%'; }, pct);
}

// called from a worker thread
void report_progress(int pct) {
  emscripten_sync_run_in_main_runtime_thread(EM_FUNC_SIG_VI, set_status, pct);
}

Keep these hops coarse. Each one is a message through the main thread’s event loop, so updating a progress bar once per percent is fine while updating it once per pixel will cost more than the work.

Expected output

With the four-thread example, the console shows interleaved completions and the loader reports no warnings:

thread 1 done: 99999999
thread 0 done: 74999997
thread 3 done: 124999990
thread 2 done: 99999997

The order varies between runs — the threads genuinely run in parallel. In the Sources panel of Chrome DevTools, the threads list shows the main thread plus one entry per pooled worker, each with the same app.wasm loaded. A single-threaded result where the lines always arrive in index order is a hint that the build silently fell back, usually because -pthread was missing from the link.

Gotchas

  • SharedArrayBuffer is not defined or Atomics is not defined. The page is not cross-origin isolated. Check self.crossOriginIsolated first; it is the cause far more often than the build.
  • The tab freezes at the first join. A blocking call on the main thread is waiting for a worker that needs the main thread to start. Use PROXY_TO_PTHREAD or raise PTHREAD_POOL_SIZE to cover every early pthread_create.
  • Tried to spawn a new thread, but the thread pool is exhausted. The program created more threads than the pool holds and the extra spawn is waiting on the main loop. Enlarge the pool or create threads lazily from a worker.
  • Random corruption that disappears in debug builds. One object file was compiled without -pthread, so its accesses are not atomic. Rebuild every dependency, including ports and static libraries, with the flag.
Choosing how main() and blocking calls coexist A decision tree. If main() blocks on joins or locks, either run it in a worker with PROXY_TO_PTHREAD or restructure it to return to the event loop. If it never blocks, a sized thread pool on the main thread is enough. Does main() block on pthread_join, mutexes or condition variables? yes, and it touches the DOM rarely -sPROXY_TO_PTHREAD main() runs in a worker; blocking is legal there yes, and it drives the UI Restructure into callbacks start threads, return, finish on a message no PTHREAD_POOL_SIZE only pool covers every early pthread_create

Performance note

The example kernel took 1.92 s single-threaded and 0.53 s with four threads on an eight-core laptop — a 3.6× speedup, close to linear because the threads share nothing. Workloads that contend on one mutex scale far worse in the browser than natively, because the main thread’s busy-wait and the cost of notify across workers add latency to every handoff. If profiling shows time in emscripten_futex_wait, the fix is less sharing, not more threads. Compare with the benchmark harness before and after so the speedup is measured rather than assumed.

Frequently Asked Questions

Can I use std::thread instead of raw pthreads? Yes. libc++ implements std::thread, std::mutex and std::condition_variable on top of Emscripten’s pthreads, so the same flags apply. std::async with the default launch policy also works but tends to create threads on demand, which interacts badly with an undersized pool.

Do threaded builds work in browsers that cannot be isolated? No. Ship a second, single-threaded build and choose between them at load time, as described in handling browsers without SharedArrayBuffer.

How is this different from Rust threads with wasm-bindgen? The underlying mechanism — workers sharing one memory — is the same. Rust’s ecosystem exposes it through Rayon-based thread pools rather than a libc pthreads layer, and requires nightly flags to rebuild the standard library with atomics.

← Back to C/C++ to Wasm with Emscripten