Porting pthreads Code with Emscripten
This page answers one task: take C or C++ code that already uses POSIX threads and get it running,
multi-threaded, in a browser through Emscripten — without deadlocking the page on its first join.
Prerequisites
- [ ] emsdk 3.1.50 or newer, activated in the shell (
emcc -vprints the version). - [ ] C or C++ source that uses
pthread_create,std::threador OpenMP-style worker threads. - [ ] A dev server that can send
Cross-Origin-Opener-PolicyandCross-Origin-Embedder-Policyheaders. - [ ] Chrome, Firefox or Safari 16.4+ for testing; all three support shared WebAssembly memory once isolated.
How Emscripten maps a pthread onto the web
A native pthread is an operating-system thread sharing one address space. The browser has no such
primitive. What it has is the Web Worker — a separate JavaScript event loop with its own global scope —
and SharedArrayBuffer, a block of memory more than one worker can see. Emscripten builds threads out
of those two pieces: the module’s linear memory
becomes a shared memory, every pthread becomes a worker that instantiates the same module against it,
and the synchronization primitives in libc are implemented with the memory.atomic.wait32 and
memory.atomic.notify instructions.
The consequence is that a “thread” costs a worker: a few megabytes of JavaScript heap, a script fetch and a module instantiation. Creating one is not the microsecond operation C programmers expect, which is why Emscripten keeps a pool of pre-started workers.
The shared memory is also why threads need cross-origin isolation. Browsers only expose
SharedArrayBuffer to pages served with the right headers, as explained in
configuring COOP/COEP headers;
without them the module fails at startup with a clear but easily misread error.
Step 1 — compile and link with -pthread
The flag must appear on every compile and on the link. A translation unit compiled without it uses non-atomic loads and stores; linking it into a threaded build produces a module that validates but races.
emcc -O2 -pthread -c worker_pool.c -o worker_pool.o
emcc -O2 -pthread -c main.c -o main.o
emcc -O2 -pthread worker_pool.o main.o \
-sPTHREAD_POOL_SIZE=4 \
-sINITIAL_MEMORY=64MB \
-sEXPORTED_RUNTIME_METHODS=ccall \
-o app.js
The output is app.js (the loader), app.wasm (the module) and, in older emsdk releases, a separate
app.worker.js. Newer releases fold the worker bootstrap into app.js itself, which simplifies hosting:
there is one script for both the main thread and the workers.
INITIAL_MEMORY matters more than usual. A shared memory must declare a maximum, and growing it is
slower than growing an unshared one because every worker has to observe the new size. Sizing the initial
memory for the expected peak avoids growth on the hot path; -sALLOW_MEMORY_GROWTH still works with
threads but carries a measurable cost on every access from JavaScript, since views must be re-checked.
Step 2 — size the worker pool
PTHREAD_POOL_SIZE pre-creates workers while the module loads, so pthread_create can hand a start
routine to an idle worker immediately. Without a pool, creation has to spawn a worker and wait for it to
load — and that wait needs the main thread’s event loop to turn, which is where ports deadlock.
// main.c — creates four workers and joins them
#include <pthread.h>
#include <stdio.h>
static void *work(void *arg) {
int id = (int)(intptr_t)arg;
long sum = 0;
for (long i = 0; i < 50000000; i++) sum += i % (id + 3);
printf("thread %d done: %ld\n", id, sum);
return NULL;
}
int main(void) {
pthread_t t[4];
for (int i = 0; i < 4; i++) pthread_create(&t[i], NULL, work, (void *)(intptr_t)i);
for (int i = 0; i < 4; i++) pthread_join(t[i], NULL);
return 0;
}
Match the pool to the number of threads the program creates early. For code that sizes itself from the
hardware, -sPTHREAD_POOL_SIZE=navigator.hardwareConcurrency is accepted as a JavaScript expression,
though it is worth capping: a sixteen-core desktop will happily start sixteen workers it never uses.
Step 3 — never block the main thread on a thread that has not started
The single most common failure in a pthreads port is a pthread_join or a mutex wait on the browser’s
main thread. The main thread is not allowed to use Atomics.wait; Emscripten works around this with a
busy-wait, which burns CPU but functions — unless the thread being waited for needs the main thread to
do something first, such as spawning its worker. Then nothing progresses and the tab hangs.
There are two robust fixes. The first is to keep main() off the main thread entirely:
emcc -O2 -pthread main.o worker_pool.o \
-sPROXY_TO_PTHREAD \
-sPTHREAD_POOL_SIZE=4 \
-o app.js
PROXY_TO_PTHREAD runs main() in a worker, leaving the browser thread free to service the runtime.
Blocking calls inside main() become legal because they happen in a worker. The cost is that anything
touching the DOM must be proxied back, which Emscripten does for its own library functions and you do
with emscripten_async_run_in_main_runtime_thread for your own.
The second fix, when main() must stay put, is to restructure so the main thread never waits: start the
threads, return to the event loop, and have the last thread post a completion message.
Step 4 — serve it with isolation headers
A threaded build will not start on a page that is not cross-origin isolated. For local testing, any of the servers in local development server configurations works once the two headers are set:
npx http-server . -p 8080 \
-H "Cross-Origin-Opener-Policy: same-origin" \
-H "Cross-Origin-Embedder-Policy: require-corp"
Then confirm in the console before chasing any other bug:
console.log("isolated:", self.crossOriginIsolated); // must be true
console.log("SAB:", typeof SharedArrayBuffer); // must be "function"
Every subresource on an isolated page must opt in to being embedded, which in practice means
same-origin assets or third-party assets served with Cross-Origin-Resource-Policy. Fonts, analytics
scripts and images from other origins are the usual casualties.
Step 5 — send DOM work back to the main thread
Workers have no document. Code that ran on a desktop thread and called into a UI toolkit now needs an
explicit hop back to the browser thread for anything visible. Emscripten provides a family of proxying
calls for this; the synchronous variant blocks the calling worker until the main thread has run the
function, which is safe precisely because the caller is not the main thread.
#include <emscripten/threading.h>
static void set_status(int pct) {
EM_ASM({ document.getElementById('status').textContent = $0 + '%'; }, pct);
}
// called from a worker thread
void report_progress(int pct) {
emscripten_sync_run_in_main_runtime_thread(EM_FUNC_SIG_VI, set_status, pct);
}
Keep these hops coarse. Each one is a message through the main thread’s event loop, so updating a progress bar once per percent is fine while updating it once per pixel will cost more than the work.
Expected output
With the four-thread example, the console shows interleaved completions and the loader reports no warnings:
thread 1 done: 99999999
thread 0 done: 74999997
thread 3 done: 124999990
thread 2 done: 99999997
The order varies between runs — the threads genuinely run in parallel. In the Sources panel of Chrome
DevTools, the threads list shows the main thread plus one entry per pooled worker, each with the same
app.wasm loaded. A single-threaded result where the lines always arrive in index order is a hint that
the build silently fell back, usually because -pthread was missing from the link.
Gotchas
SharedArrayBuffer is not definedorAtomics is not defined. The page is not cross-origin isolated. Checkself.crossOriginIsolatedfirst; it is the cause far more often than the build.- The tab freezes at the first join. A blocking call on the main thread is waiting for a worker that
needs the main thread to start. Use
PROXY_TO_PTHREADor raisePTHREAD_POOL_SIZEto cover every earlypthread_create. Tried to spawn a new thread, but the thread pool is exhausted.The program created more threads than the pool holds and the extra spawn is waiting on the main loop. Enlarge the pool or create threads lazily from a worker.- Random corruption that disappears in debug builds. One object file was compiled without
-pthread, so its accesses are not atomic. Rebuild every dependency, including ports and static libraries, with the flag.
Performance note
The example kernel took 1.92 s single-threaded and 0.53 s with four threads on an eight-core laptop —
a 3.6× speedup, close to linear because the threads share nothing. Workloads that contend on one mutex
scale far worse in the browser than natively, because the main thread’s busy-wait and the cost of
notify across workers add latency to every handoff. If profiling shows time in
emscripten_futex_wait, the fix is less sharing, not more threads. Compare with the
benchmark harness
before and after so the speedup is measured rather than assumed.
Frequently Asked Questions
Can I use std::thread instead of raw pthreads?
Yes. libc++ implements std::thread, std::mutex and std::condition_variable on top of Emscripten’s
pthreads, so the same flags apply. std::async with the default launch policy also works but tends to
create threads on demand, which interacts badly with an undersized pool.
Do threaded builds work in browsers that cannot be isolated? No. Ship a second, single-threaded build and choose between them at load time, as described in handling browsers without SharedArrayBuffer.
How is this different from Rust threads with wasm-bindgen? The underlying mechanism — workers sharing one memory — is the same. Rust’s ecosystem exposes it through Rayon-based thread pools rather than a libc pthreads layer, and requires nightly flags to rebuild the standard library with atomics.
Related
- Sharing memory between Wasm and Web Workers — the shared-memory model underneath every pthread.
- Using Atomics for Wasm thread synchronization — what mutexes and condition variables compile to.
- Building Emscripten projects with CMake — passing
-pthreadthrough a CMake toolchain consistently. - Catching memory bugs with Emscripten sanitizers — finding the races a port exposes.
← Back to C/C++ to Wasm with Emscripten