Debugging Deadlocks in Threaded Wasm

This page answers one task: a WebAssembly application using threads — Emscripten pthreads, wasm-bindgen-rayon, or a hand-built worker pool — sometimes stops making progress. No error appears; work never finishes. You want to find out which thread is waiting for what, and why.

Prerequisites

  • [ ] A threaded build (shared memory, atomics) running in a cross-origin isolated page.
  • [ ] Chrome DevTools or Firefox DevTools, which can inspect workers.
  • [ ] The ability to rebuild with extra instrumentation.

Is it really a deadlock?

“It hangs” has several causes that look alike from outside. A deadlock is threads waiting on each other in a cycle — thread A holds lock 1 and waits for lock 2, thread B holds lock 2 and waits for lock 1 — so all of them sleep and CPU usage drops to zero. A livelock or busy loop keeps threads running without progress, and CPU usage stays high. Main-thread starvation happens when worker threads wait for something the main thread must do — proxied calls, worker creation in Emscripten — while the main thread is itself blocked or busy. And a lost wake-up — a notify sent before the waiter started waiting, on a different address, or with the wrong expected value — leaves one thread asleep forever with no cycle at all.

The first diagnostic is CPU usage in the browser’s task manager or the OS: near zero suggests deadlock or lost wake-up; high suggests a busy loop. The second is what each thread is doing, which DevTools can show.

Hang symptoms and likely causes Zero CPU with all workers waiting suggests a lock-ordering deadlock or a lost wake-up. High CPU with no progress suggests a busy loop or livelock. Workers waiting on a proxied call while the main thread is blocked suggests main-thread starvation. Hangs only at startup with Emscripten suggest the pthread pool was not pre-created. symptom likely cause first check CPU ~0, all workers waiting lock-order deadlock / lost wake-up paused stacks of each worker CPU high, no progress busy loop / livelock profile the busy thread workers wait on main thread main-thread starvation is the main thread blocked? hang at thread creation (Emscripten) pool not pre-created PTHREAD_POOL_SIZE

Step 1 — pause every thread and read the stacks

In Chrome DevTools, the Sources panel’s Threads pane lists the main thread and each worker. Pause the page while it hangs, then select each thread and read its call stack. A thread blocked in Atomics.wait (or in Wasm frames ending at a wait instruction inside futex_wait, pthread_mutex_lock, std::sync::Mutex::lock) is waiting on a lock or condition; the frames above it show which function tried to acquire it. With DWARF debug info and the C/C++ DevTools extension, those frames map to source lines. Write down, for each thread: what it waits for, and what it last acquired. A cycle in that list is your deadlock.

Step 2 — record lock ownership in shared memory

Stacks show who is waiting but not who holds a lock. Add ownership tracking to debug builds: when a thread acquires a lock, write its thread ID into a slot next to the lock word; clear it on release. A debug export can then dump every lock’s state:

#[cfg(feature = "lock-debug")]
pub struct DebugMutex<T> { inner: std::sync::Mutex<T>, owner: std::sync::atomic::AtomicU32, name: &'static str }

#[cfg(feature = "lock-debug")]
impl<T> DebugMutex<T> {
    pub fn lock(&self) -> DebugGuard<'_, T> {
        let g = self.inner.lock().unwrap();
        self.owner.store(current_thread_id(), std::sync::atomic::Ordering::SeqCst);
        register_held(self.name);
        DebugGuard { guard: g, owner: &self.owner, name: self.name }
    }
}

When the app hangs, call the dump from the console (exports are callable while workers are blocked, as long as the main thread is not): “lock queue held by thread 3, waited on by threads 1 and 4; lock cache held by thread 1, waited on by thread 3” shows the cycle directly.

Step 3 — add timeouts that report stuck waits

Infinite waits hide problems. In debug builds, wait with a timeout and report when it expires — the lock’s name, the waiting thread, the current owner, and the waiter’s stack — then continue waiting or abort:

function lockDebug(i32, i, name) {
  for (;;) {
    if (Atomics.compareExchange(i32, i, 0, 1) === 0) return;
    if (Atomics.wait(i32, i, 1, 2000) === "timed-out") {
      console.error(`waited 2 s for lock ${name}; owner thread ${Atomics.load(i32, i + 1)}`, new Error().stack);
    }
  }
}

In Rust, parking_lot has an optional deadlock detector (deadlock_detection feature) that finds cycles among its locks; it uses a background thread, so in Wasm it needs to be driven manually by calling its check function periodically from a worker.

A lock-ordering deadlock Worker 1 acquires lock A then requests lock B. Meanwhile worker 2 acquires lock B then requests lock A. Each waits for the other to release, forming a cycle. Neither can proceed, CPU drops to zero, and the job never finishes. worker 1 holds A wants B worker 2 holds B wants A both Atomics.wait sleeping cycle: A→B→A no one releases job never completes CPU ~0

Step 4 — fix the cause

The fixes depend on the cause:

  • Lock-ordering cycles. Give locks a global order and always acquire in that order. If a function needs locks A and B, it takes A first everywhere. Or merge the two locks if they always go together.
  • Holding a lock while waiting for something else. Release locks before waiting on a channel, a condition, a message or the main thread.
  • Lost wake-ups. Wait on the same address the notifier notifies, with the correct expected value, and re-check the condition in a loop after waking. Use condition variables from the standard library rather than hand-rolled wait/notify pairs.
  • Main-thread dependencies. In Emscripten, pre-create workers with -sPTHREAD_POOL_SIZE=N so starting a thread does not need the main thread to yield, and consider -sPROXY_TO_PTHREAD so main itself runs in a worker.
  • Nested parallelism. A rayon task that blocks waiting for another rayon task in a fully occupied pool can stall; avoid blocking calls inside pool tasks and let rayon’s work stealing handle nesting.

Step 5 — reproduce with stress and fewer threads

Deadlocks depend on timing. Make them more frequent: run the scenario in a loop, add random short delays inside critical sections in debug builds, and vary the worker count — including two, which often makes a cycle easier to hit. Once reproducible, confirm the fix by running the stress loop long enough that the old failure rate would have shown several hangs.

Emscripten-specific causes

Emscripten’s pthreads add a few patterns of their own. Creating a thread requires a worker, and if the pool is empty the runtime must create one, which needs the main thread to return to the event loop; a main thread that blocks in pthread_join immediately after pthread_create deadlocks. Some operations are proxied to the main thread (certain file-system and DOM-related calls); a worker waiting on such a call while the main thread waits on that worker deadlocks. -sPTHREAD_POOL_SIZE, -sPROXY_TO_PTHREAD and avoiding blocking calls on the main thread remove most of these. Assertions builds (-sASSERTIONS=2) print warnings when the main thread blocks.

Hangs that are not locks

If no thread is waiting on a lock and CPU is high, profile instead: the Performance panel records worker threads too. Common causes are spin-waits without backoff, a work-stealing loop that never finds work but keeps polling, or a lost termination signal in a hand-built pool where workers check a flag that is never set because it was written non-atomically. If CPU is zero and no locks are involved, check for a promise or message that never arrives — a worker that crashed (check the console for its error) never posts its result.

Preventing deadlocks by design

Debugging is the expensive route; designs that cannot deadlock are cheaper. Keep the number of locks small and document their order in one place. Prefer message passing or lock-free queues for communication between workers, so threads never wait on each other while holding state. Keep critical sections free of calls into code you do not control — callbacks, JavaScript imports, logging that may itself lock. And give every blocking wait a timeout in production too, reporting through the error pipeline, so a hang becomes a diagnosable event with names and stacks instead of a frozen tab.

Expected output

Pausing during a hang shows worker 2 in Mutex::lock called from Cache::insert and worker 4 in Mutex::lock from Queue::push; the lock dump shows each holding the lock the other wants; introducing a global order (queue before cache) removes the cycle; a 30-minute stress loop that previously hung every few minutes completes; and debug builds report any wait longer than two seconds with names and stacks.

Gotchas

  • Assuming every hang is a deadlock. Check CPU first.
  • Blocking the main thread in threaded Emscripten builds. Workers that need it stall.
  • Holding locks across waits. Release before waiting on anything else.
  • Hand-rolled wait/notify. Lost wake-ups. Use standard condition variables.
  • Testing only at one thread count. Timing hides cycles. Vary it.

Performance note

The debug lock wrapper added about 15 ns per acquisition — too much for release builds in hot loops, negligible for diagnosis. Release builds keep only the timeout-based stuck-lock report, at no cost on the fast path.

Cost per lock acquisition with debug instrumentation Nanoseconds per uncontended lock and unlock with a plain mutex and with the debug wrapper that records owners and held locks. ns per lock + unlock plain mutex 10 ns debug wrapper 25 ns

Frequently Asked Questions

Can DevTools show which thread holds a lock? No — only who is waiting. Record ownership yourself in debug builds.

Does Firefox show worker stacks? Yes — its debugger lists workers as separate threads you can pause.

Are deadlocks possible without locks? Yes, with any circular wait — channels, joins or messages.

Why does it only hang on some machines? Thread timing differs with core counts and load; stress testing exposes it.

Should production builds keep wait timeouts? Yes — a timed-out wait that reports and recovers is far easier to diagnose than a silent hang.

← Back to SharedArrayBuffer, Atomics & Threading