One Node process runs JavaScript on one core, so a CPU-bound Express 24,430 app on an 18-core machine wastes 17 of them. node:cluster forks workers that share a listening socket; Clustering Across CPU Cores covers the API, the primary/worker split and graceful shutdown. What belongs here is the effect on an Express app and the constraint it imposes.
if (cluster.isPrimary) {
for (let i = 0; i < Number(process.env.WORKERS); i += 1) cluster.fork();
cluster.on('exit', () => cluster.fork()); // replace a crashed worker
} else {
const app = express();
app.get('/top', (req, res) => res.json({ pid: process.pid, data: aggregate() }));
app.listen(4309, '127.0.0.1');
}$ node clusterbench.mjs # autocannon -c 50 -d 5 against /top, i9-7980XE 1 worker(s): 1282 req/s 38.43 ms avg p99 65 ms 4 worker(s): 4248 req/s 11.29 ms avg p99 53 ms 8 worker(s): 8031 req/s 5.78 ms avg p99 22 ms
Throughput scaled 3.3x at four workers and 6.3x at eight, and the 99th percentile fell from 65 ms to 22 ms — unusually clean, because the route is pure computation with nothing shared behind it. A route that queries MongoDB 1,815 stops scaling once the database is the bottleneck, so measure instead of assuming. These runs are from Windows, where cluster.schedulingPolicy defaults to SCHED_NONE and the OS distributes accepted connections; elsewhere the default is SCHED_RR, the primary handing them out round-robin.
The constraint is that the app must hold no request state in process memory. In-memory sessions, a rate-limit counter in a Map and the per-worker cache of Server-Side Caching with Redis all break when a client's next request lands on another worker; move them to Redis 2,763 (Session Stores for Production, Rate Limiting) and Socket.IO 24,482 to its Redis adapter (Scaling Real Time).
In production something must restart workers, cap their memory and start them at boot. pm2 29,762 (https://github.com/Unitech/pm2 43,298 ) 7.0.4 does all three and forks for you, so the app file stays a plain single-process server: {"script": "./src/server.js", "instances": "max", "exec_mode": "cluster", "max_memory_restart": "400M", "kill_timeout": 5000} in ecosystem.config.json. Then pm2 reload <name> restarts workers one at a time, so a deploy drops no connections — provided the app closes its server on SIGTERM within kill_timeout. In a container you skip pm2: the orchestrator is the process manager, one process per container (Docker).