Three endpoints answer the platform, not the user. Liveness says the process can still respond; readiness says this instance should receive traffic; metrics expose counters a scraper reads. Keep them apart — a load balancer that restarts a pod because its database is slow turns a dependency blip into an outage.
import express from 'express';
import client from 'prom-client';
const app = express(), startedAt = Date.now();
let ready = false; // flipped when warm-up finishes
client.collectDefaultMetrics(); // event-loop lag, heap, handles
const httpDuration = new client.Histogram({
name: 'http_request_duration_seconds', help: 'Request duration in seconds',
labelNames: ['method', 'route', 'status'], buckets: [0.005, 0.025, 0.1, 0.5, 1, 5] });
app.use((req, res, next) => {
const end = httpDuration.startTimer();
res.on('finish', () => end({ method: req.method, status: res.statusCode,
route: req.route?.path ?? req.path })); // the pattern, never the raw URL
next(); });
app.get('/healthz', (req, res) => res.json({ status: 'ok',
uptime: (Date.now() - startedAt) / 1000 }));
app.get('/readyz', (req, res) => res.status(ready ? 200 : 503)
.json({ status: ready ? 'ready' : 'not_ready', checks: [{ name: 'store', ok: ready }] }));
app.get('/metrics', async (req, res) => res.type(client.register.contentType)
.send(await client.register.metrics()));
app.listen(3115, () => setTimeout(() => { ready = true; }, 1500)); // 1.5 s warm-upProbing during and after the warm-up shows the two endpoints disagreeing on purpose, and /metrics answering in the Prometheus 16,091 text format:
$ curl -s -w " [HTTP %{http_code}]\n" localhost:3115/readyz # during warm-up
{"status":"not_ready","checks":[{"name":"store","ok":false}]} [HTTP 503]
$ curl -s -w " [HTTP %{http_code}]\n" localhost:3115/healthz
{"status":"ok","uptime":2.987} [HTTP 200]
$ curl -s -w " [HTTP %{http_code}]\n" localhost:3115/readyz # after warm-up
{"status":"ready","checks":[{"name":"store","ok":true}]} [HTTP 200]
$ curl -s localhost:3115/metrics | grep healthz
http_request_duration_seconds_sum{method="GET",status="200",route="/healthz"} 0.0008168Labeling with req.originalUrl instead of the route pattern creates one time series per book id, and a few thousand ids exhaust a Prometheus server's memory; the failure has a name, cardinality explosion. collectDefaultMetrics adds event-loop lag and heap size, usually where you find that a synchronous JSON parse, not the database, is the bottleneck. Metrics tell you that p99 latency doubled; tracing tells you where: OpenTelemetry 36,171 's Node SDK (opentelemetry-js (https://github.com/open-telemetry/opentelemetry-js 3,479 )) patches http, Express 24,430 and common drivers at load time, so a request becomes a span tree with no handler changes.