Building Rust Wasm with Threads Using Rayon
This page answers one task: Rust code that uses Rayon’s par_iter should run in parallel in the browser, using several cores through Web Workers
that share one WebAssembly memory — with a fallback for browsers or pages where threads are unavailable.
Prerequisites
- [ ] A Rust crate using Rayon (
rayon = "1") for data-parallel work. - [ ] A nightly Rust toolchain (needed to rebuild the standard library with atomics).
- [ ] A page served with cross-origin isolation headers (COOP
same-origin, COEPrequire-corp). - [ ] A bundler that handles module workers (Vite, webpack 5, Parcel 2) or a no-bundler setup with
--target web.
How threads work in browser WebAssembly
WebAssembly threads are not created by the module. They are Web Workers, each running its own instance of the same module, all instantiated with one
shared WebAssembly.Memory. Because they share memory, Rust’s threads in different workers see the same heap, and atomics and locks work as they do
natively. What the browser does not provide is std::thread::spawn — there is no way for Wasm to create a worker by itself. So something in JavaScript
must start the workers, instantiate the module in each, and hand them to the Rust thread pool.
wasm-bindgen-rayon does exactly that for Rayon. It exports an initThreadPool(n) function that starts n workers, instantiates the module in each
with the shared memory, and registers them as Rayon’s global pool. After it resolves, par_iter and rayon::join run in parallel just as they would
natively.
Step 1 — add the dependency and export the pool initialiser
[dependencies]
wasm-bindgen = "0.2"
rayon = "1.10"
wasm-bindgen-rayon = "1.2"
pub use wasm_bindgen_rayon::init_thread_pool; // exported to JS as initThreadPool
use rayon::prelude::*;
use wasm_bindgen::prelude::*;
#[wasm_bindgen]
pub fn sum_of_squares(data: &[f64]) -> f64 {
data.par_iter().map(|x| x * x).sum()
}
The re-export is all the Rust code needs. Everything else is in the build.
Step 2 — build with atomics and a rebuilt standard library
The precompiled standard library for wasm32-unknown-unknown is built without atomics, so it cannot be used in a threaded module. Rebuild it with the
right target features on nightly. Pin the toolchain and put the flags in configuration so every build uses them:
# rust-toolchain.toml
[toolchain]
channel = "nightly-2025-06-01"
components = ["rust-src"]
targets = ["wasm32-unknown-unknown"]
# .cargo/config.toml
[target.wasm32-unknown-unknown]
rustflags = ["-C", "target-feature=+atomics,+bulk-memory", "-C", "link-arg=--max-memory=1073741824"]
[unstable]
build-std = ["panic_abort", "std"]
wasm-pack build --target web --release
+atomics enables shared memory and atomic instructions; +bulk-memory is needed for the passive data segments a shared memory uses during
initialisation. A shared memory must declare a maximum size, which --max-memory sets. Pinning the nightly date avoids surprises when a newer nightly
changes something — the reasons are covered in
pinning Wasm toolchain versions.
Step 3 — initialise the pool in JavaScript
import init, { initThreadPool, sum_of_squares } from "./pkg/compute.js";
await init();
await initThreadPool(navigator.hardwareConcurrency); // starts the workers
const data = Float64Array.from({ length: 10_000_000 }, (_, i) => i / 1e6);
console.log(sum_of_squares(data));
initThreadPool must be awaited before any parallel function runs, and should be called once. Choose the pool size deliberately: hardwareConcurrency
counts logical cores, including the one the main thread uses; on phones, fewer threads than cores often give better results because of thermal limits
and big.LITTLE cores. Run the heavy calls from a worker rather than the main thread, so that the main thread is never blocked waiting for the pool — a
blocked main thread also cannot service the workers’ messages during startup.
Step 4 — serve with cross-origin isolation
Shared memory requires a cross-origin isolated page. Without the headers, SharedArrayBuffer is unavailable and the module’s shared memory cannot be
created — instantiation fails with an error about shared memory. Serve:
Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp
Every subresource the page loads must then be same-origin or send Cross-Origin-Resource-Policy or CORS headers. The dev-server configuration is in
configuring the Vite dev server for Wasm,
and runtime detection in
detecting cross-origin isolation at runtime.
Step 5 — ship a single-threaded fallback
Pages that cannot be isolated, and browsers or embedded webviews without shared memory, need a version that works without threads. Build the crate
twice: once with the threaded configuration above, once with stable Rust and no atomics, where Rayon runs everything on the current thread. A Cargo
feature can switch wasm-bindgen-rayon on only for the threaded build. Then choose at runtime:
const threaded = crossOriginIsolated && typeof SharedArrayBuffer !== "undefined";
const mod = threaded ? await import("./pkg-threads/compute.js") : await import("./pkg/compute.js");
await mod.default();
if (threaded) await mod.initThreadPool(Math.min(navigator.hardwareConcurrency, 8));
Both builds expose the same functions, so the rest of the application does not care which one loaded.
Choosing work that parallelises well
Threads in the browser have the same economics as natively, plus a few extra costs. Starting the pool takes tens of milliseconds, because each worker
must instantiate the module; do it once, early, and keep the pool. Each parallel call pays a small scheduling overhead, so parallelise coarse work — image
filters over megapixels, simulations over thousands of bodies, compression of large buffers — not loops of a few hundred cheap iterations, where the
overhead exceeds the gain. Memory bandwidth is shared, so memory-bound loops scale worse than compute-bound ones; a simple sum over a large array may
improve by 2× on eight cores while a per-pixel convolution improves by 6×. Rayon’s with_min_len on parallel iterators sets a minimum chunk size,
which is an easy way to stop it from splitting work too finely. Measure scaling with one, two, four and eight threads; the curve shows quickly where
adding threads stops paying.
Debugging a threaded build
Threaded builds fail in ways single-threaded ones do not, and most failures happen at startup. If initThreadPool never resolves, a worker failed to
load — open the browser’s worker list in DevTools and look for a script error, usually a path the bundler did not rewrite or a missing isolation header
on the worker script itself, which must also be served with COEP-compatible headers. If instantiation throws about shared memory or “memory import must
be shared”, the module was built without +atomics or the glue was generated from a different build than the .wasm file. If the page works but runs no
faster, check that the parallel function actually runs off the main thread, and that the work is large enough to split. A panic in any worker aborts
that worker’s instance; with console_error_panic_hook installed in each thread the message appears in the worker’s console, and the panicking call
never returns, so add timeouts around parallel calls during development. DevTools’ Performance panel shows each worker as its own track, which makes it
easy to confirm that all threads are busy — and to spot one thread doing most of the work because the data was split unevenly.
Expected output
On an 8-core laptop, sum_of_squares over 10 million values takes about 3 ms with the pool versus 17 ms single-threaded; the Performance panel shows
eight worker threads active during the call; and on a non-isolated page the fallback build loads and produces the same result single-threaded.
Gotchas
SharedArrayBuffer is not defined. The page is not cross-origin isolated. Check the headers.- Build errors about atomics in
std. The standard library was not rebuilt. Use nightly withbuild-stdandrust-src. - Blocking the main thread. Waiting on the pool from the main thread can deadlock during startup. Call parallel code from a worker.
- No maximum memory. Shared memory requires one. Set
--max-memory. - Pool too large on phones. More threads than effective cores slow things down. Measure and cap.
Performance note
Applying a 5×5 convolution to a 24-megapixel image took 410 ms single-threaded and 68 ms with 8 Rayon threads in Chrome on an 8-core laptop (6× faster). Summing 10 million floats scaled less well — 17 ms to 3 ms — because it is limited by memory bandwidth. Pool start-up cost 45 ms once.
Frequently Asked Questions
Can I use std::thread::spawn directly?
Not without help — spawning requires JavaScript to create a worker. Rayon with wasm-bindgen-rayon, or crates such as wasm_thread, provide that.
Does this work in Node? wasm-bindgen-rayon targets browsers. In Node, use worker_threads with a shared memory, or run native code instead.
Will this need nightly forever?
Rebuilding std with atomics currently needs nightly. Track the Rust project’s progress on a threads-enabled Wasm target.
Why is the threaded build larger?
Atomics, thread-local storage setup and the rebuilt std add size, typically 10–30 KB.
Can the pool size change later?
No — Rayon’s global pool is initialised once. Choose the size at startup, or build a custom ThreadPool for separate workloads.
Related
- Building a Wasm thread pool — the mechanism without Rayon.
- Sharing memory between Wasm and Web Workers — the shared-memory basics.
- Porting pthreads code with Emscripten — the C/C++ equivalent.
- Avoiding data races in shared Wasm memory — correctness under threads.
← Back to SharedArrayBuffer, Atomics & Threading