Worker Threads

A worker thread is a second JavaScript thread inside your process, with its own V8 86,723 isolate, heap and event loop. A worker cannot see your variables, so ordinary objects have no data races. node:worker_threads is stable since Node 12.11.0 and needs no flag.

One file plays both roles: as the entry point it builds a pool, loaded as a worker (isMainThread is false) it answers messages. Counting primes by trial division is deliberately dumb work; the shape of the curve is the point.

A worker pool measured against one threadJavaScript
import { Worker, isMainThread, parentPort } from 'node:worker_threads';
import { availableParallelism } from 'node:os';
function count({ from, to }) {
  let n = 0;
  outer: for (let i = Math.max(from, 2); i < to; i++) {
    for (let d = 2; d * d <= i; d++) if (i % d === 0) continue outer;
    n++;
  }
  return n;
}
const tasks = Array.from({ length: 48 }, (_, i) => ({ from: i * 250e3, to: (i + 1) * 250e3 }));
if (!isMainThread) parentPort.on('message', (t) => parentPort.postMessage(count(t)));
else {
  const pool = (size) => new Promise((resolve) => {
    const queue = tasks.slice(); let left = tasks.length;
    const ws = Array.from({ length: size }, () => new Worker(new URL(import.meta.url)));
    const feed = (w) => queue.length && w.postMessage(queue.shift());
    for (const w of ws) {
      w.on('message', () => --left ? feed(w) : (ws.forEach((x) => x.terminate()), resolve()));
      feed(w);
    }
  });
  const s = (ms) => (ms / 1000).toFixed(2); const t0 = performance.now();
  const primes = tasks.reduce((n, t) => n + count(t), 0), base = performance.now() - t0;
  console.log(`${availableParallelism()} cores; ${primes} primes; 1 thread ${s(base)} s`);
  for (const size of [4, 8, 16, 32]) {
    const t = performance.now(); await pool(size);
    const ms = performance.now() - t;
    console.log(`${size} workers: ${s(ms)} s, speedup ${(base / ms).toFixed(1)}x`);
  }
}
Output
36 cores; 788060 primes; 1 thread 6.96 s
4 workers: 2.31 s, speedup 3.0x
8 workers: 1.36 s, speedup 5.1x
16 workers: 0.87 s, speedup 8.0x
32 workers: 0.64 s, speedup 10.8x

That run is an 18-core, 36-thread Intel Core i9-7980XE under Windows 11. Thirty-two workers on 36 logical cores returned 10.8x, not 32x: half those "cores" are hyperthreads sharing one execution unit, the clock drops as more cores wake, and each new Worker builds a fresh V8 isolate — 34 ms apiece here, and 36 MB of RSS became 49 MB for eight of them.

Build the pool once, never a worker per request, and send tasks worth several milliseconds: postMessage copies its payload through structured clone, and a 1 ms task drowns in that overhead.