Two numbers describe a server under load, and they trade against each other. Throughput is requests served per second; latency is how long one client waits. Pushed past its capacity a server keeps throughput roughly flat while latency climbs, because the extra requests queue in the accept backlog. An average latency is therefore nearly useless alone: report the median and the 99th percentile, the one request in a hundred users complain about.
autocannon 8,524 (https://github.com/mcollina/autocannon 8,524 ) is the load generator to reach for from a Node project: one npm 2,036 i -D autocannon, no runtime to install, and it reports exactly those percentiles. -c is the concurrent connection count, -d the duration in seconds; fifty connections for five seconds is a good first probe, long enough for the JIT to warm up and short enough to repeat after every change.
const books = Array.from({ length: 20 }, (_, i) =>
({ id: i + 1, title: `Book ${i + 1}`, author: 'Anon', year: 1990 + i }));
const app = express();
app.get('/books', (req, res) => res.json({ data: books }));
app.listen(4302, '127.0.0.1');$ npx autocannon -c 50 -d 5 http://127.0.0.1:4302/books ┌─────────┬──────┬──────┬───────┬───────┬─────────┬─────────┬───────┐ │ Stat │ 2.5% │ 50% │ 97.5% │ 99% │ Avg │ Stdev │ Max │ ├─────────┼──────┼──────┼───────┼───────┼─────────┼─────────┼───────┤ │ Latency │ 2 ms │ 4 ms │ 11 ms │ 13 ms │ 4.57 ms │ 2.28 ms │ 40 ms │ └─────────┴──────┴──────┴───────┴───────┴─────────┴─────────┴───────┘ │ Req/Sec │ 8,535 │ 8,535 │ 10,055 │ 10,735 │ 9,858.4 │ 860.31 │ ... 49k requests in 5.02s, 66.5 MB read
9,858 requests per second, a 4 ms median and a 13 ms 99th percentile. Is that ceiling Express 24,430 or Node? Replace the app with a node:http server answering the same pre-serialized body and the same test reports 21,083 req/s with a 1 ms median. Express costs about half the throughput of a hand-written handler on a route this trivial — the router, the req/res prototype augmentation and (ETags and Conditional Requests) the default ETag hash each charge per request. On a route that also queries MongoDB 1,815 the framework disappears into the noise, which is the whole argument for using it.
Two other open-source generators are worth knowing: k6 112,597 (Go, JavaScript scenarios) for staged ramps and pass/fail thresholds in CI, and oha or wrk 40,414 (Rust and C) when the generator must not become the limit.
Loopback has no bandwidth limit, so use it only to compare versions of your own code, never to predict what a remote client sees. Give the generator its own cores, or autocannon at 50 connections starves the server it is testing. And never compare a development run with a production one: view caching and error formatting differ.