A micro-benchmark measures a function; a load test measures a server, which is a different animal. Under concurrency you find out whether the event loop stalls and whether a fix that worked in isolation survives fifty simultaneous clients.
autocannon 8,524 is the standard Node tool for this: an HTTP/1.1 load generator by Matteo Collina, MIT licensed, at github.com/mcollina/autocannon (https://github.com/mcollina/autocannon 8,524 ), version 8.0.0. Run it with npx autocannon -c 50 -d 10 http://localhost:3100/search?category=alpha. The flags that matter are -c (concurrent connections), -d (seconds), -p (pipelined requests per connection) and -w (spread the load over worker threads so the generator is not the bottleneck). The programmatic API takes the same options.
const child = spawn(process.execPath, ['server.mjs'], { env: { ...process.env, MODE: mode } });
await delay(900); // let it bind the port
try {
const r = await autocannon({
url: 'http://127.0.0.1:3100/', connections: 50, duration: 10,
requests: cats.map((c) => ({ path: `/search?category=${c}` })), // five categories
});
console.log(mode, r.requests.average, r.latency.p50, r.latency.p99, r.non2xx);
} finally { child.kill('SIGKILL'); }The server under test is the endpoint from Measure Before You Optimize behind a MODE switch: naive reparses catalog.json on every request, warm parses it once at startup. Both ran for ten seconds with fifty keep-alive connections, generator and server sharing the 18-core i9-7980XE under Node 25.8.0.
| Mode | Throughput | p50 latency | p99 latency | Requests in 10 s |
|---|---|---|---|---|
| naive: read and parse per request | 244 req/s | 204 ms | 245 ms | 2,437 |
| warm: parse once at startup | 8,507 req/s | 5 ms | 13 ms | 85,066 |
Thirty-five times the throughput and a p99 down from a quarter of a second to 13 milliseconds, from deleting one line. The collapse is so complete because readFileSync and JSON.parse both block the event loop: while one request parses for 3.8 ms the other forty-nine wait. That is also why the naive p50 and p99 nearly coincide — every request queues behind the same bottleneck, so the distribution has no tail, just a wall.
Read the results in this order. Errors and non-2xx first — a fast server returning 500s is not fast. Then the percentiles, never the mean: p97_5 and p99 are what users feel. Then throughput, which means something only alongside the latency it was achieved at. Finally confirm the generator was not the limit: if -w 4 raises the number, you were measuring autocannon.