Load Testing an HTTP Server

A micro-benchmark measures a function; a load test measures a server, which is a different animal. Under concurrency you find out whether the event loop stalls and whether a fix that worked in isolation survives fifty simultaneous clients.

autocannon 8,524 is the standard Node tool for this: an HTTP/1.1 load generator by Matteo Collina, MIT licensed, at github.com/mcollina/autocannon (https://github.com/mcollina/autocannon 8,524 ), version 8.0.0. Run it with npx autocannon -c 50 -d 10 http://localhost:3100/search?category=alpha. The flags that matter are -c (concurrent connections), -d (seconds), -p (pipelined requests per connection) and -w (spread the load over worker threads so the generator is not the bottleneck). The programmatic API takes the same options.

Starting a server, driving it, and shutting it down by pidJavaScript
const child = spawn(process.execPath, ['server.mjs'], { env: { ...process.env, MODE: mode } });
await delay(900);                                       // let it bind the port
try {
  const r = await autocannon({
    url: 'http://127.0.0.1:3100/', connections: 50, duration: 10,
    requests: cats.map((c) => ({ path: `/search?category=${c}` })),   // five categories
  });
  console.log(mode, r.requests.average, r.latency.p50, r.latency.p99, r.non2xx);
} finally { child.kill('SIGKILL'); }

The server under test is the endpoint from Measure Before You Optimize behind a MODE switch: naive reparses catalog.json on every request, warm parses it once at startup. Both ran for ten seconds with fifty keep-alive connections, generator and server sharing the 18-core i9-7980XE under Node 25.8.0.

Effect of moving one JSON.parse out of the request path
Mode Throughput p50 latency p99 latency Requests in 10 s
naive: read and parse per request 244 req/s 204 ms 245 ms 2,437
warm: parse once at startup 8,507 req/s 5 ms 13 ms 85,066

Thirty-five times the throughput and a p99 down from a quarter of a second to 13 milliseconds, from deleting one line. The collapse is so complete because readFileSync and JSON.parse both block the event loop: while one request parses for 3.8 ms the other forty-nine wait. That is also why the naive p50 and p99 nearly coincide — every request queues behind the same bottleneck, so the distribution has no tail, just a wall.

Read the results in this order. Errors and non-2xx first — a fast server returning 500s is not fast. Then the percentiles, never the mean: p97_5 and p99 are what users feel. Then throughput, which means something only alongside the latency it was achieved at. Finally confirm the generator was not the limit: if -w 4 raises the number, you were measuring autocannon.