Clustering Across CPU Cores

node:cluster forks copies of your own program and lets all of them serve one port. Each is a full process with its own heap, so a crash or a leak takes down one worker, not the server, and the primary replaces it. For an HTTP server it beats worker threads: no shared state, no shared-state bugs.

A CPU-bound server, one process or manyJavaScript
import cluster from 'node:cluster'; import http from 'node:http';
import { availableParallelism } from 'node:os'; import { pbkdf2Sync } from 'node:crypto';
const WORKERS = Number(process.env.WORKERS || 0);
if (WORKERS > 1 && cluster.isPrimary) {
  console.log(`primary ${process.pid}: forking ${WORKERS} of ${availableParallelism()} cores`);
  for (let i = 0; i < WORKERS; i++) cluster.fork();
  cluster.on('exit', (w, code, sig) =>
    (console.log(`worker ${w.process.pid} died (${sig || code})`), cluster.fork()));
} else {
  http.createServer((req, res) => {
    const hash = pbkdf2Sync('password', 'salt', 6000, 32, 'sha256');
    res.end(`${process.pid} ${hash.toString('hex').slice(0, 8)}`);
  }).listen(3000, () => console.log('serving on pid', process.pid));
}
Output
primary 46044: forking 4 of 36 cores
serving on pid 55104
serving on pid 63856
serving on pid 43740
serving on pid 35308

Every request costs one synchronous pbkdf2 at 6,000 iterations, so the handler is pure CPU. A driver opened 24 keep-alive connections, looped requests for 10 seconds and recorded which pid answered, against the same 18-core, 36-thread Core i9-7980XE.

Measured throughput of the same handler at different cluster sizes
Configuration Requests in 10 s Throughput Workers that got traffic
single process 3,486 347 req/s 1 of 1
4 workers, round robin 10,677 1,063 req/s 4 of 4
8 workers, round robin 19,027 1,899 req/s 8 of 8
16 workers, round robin 31,990 3,196 req/s 16 of 16
8 workers, OS choice 10,996 1,095 req/s 3 of 8

Four workers gave 3.1x, sixteen gave 9.2x. The last row is the trap: under the default policy on Windows the primary hands the listening socket to every worker and lets the operating system pick the winner of each accept, and five of the eight never saw a connection. Node's round-robin scheduler — the default everywhere except Windows — balanced them perfectly, worth 73% more throughput. Set cluster.schedulingPolicy = cluster.SCHED_RR before the first fork(), or export NODE_CLUSTER_SCHED_POLICY=rr, and verify the balance rather than assume it. Roll restarts one worker at a time — worker.disconnect(), wait for 'exit', cluster.fork() — so the port is never unserved.