The stateless API suits an HPA on CPU: two replicas as the floor, six as the ceiling, and a 60-second scale-down window so the demonstration fits on a page:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: api, labels: { app: booknest, tier: api } }
spec:
scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: api }
minReplicas: 2
maxReplicas: 6
metrics:
- type: Resource
resource: { name: cpu, target: { type: Utilization, averageUtilization: 70 } }
behavior:
scaleDown: { stabilizationWindowSeconds: 60 }The Deployment manifest must stop setting replicas, or each kubectl 5,150 apply resets it to 3. Deleting the line alone makes the next apply remove the field, dropping the Deployment to one replica for a moment; updating the last-applied annotation first avoids that:
sed -i '/^ replicas: 3$/d' k8s/api-deployment.yaml
kubectl apply set-last-applied -f k8s/api-deployment.yaml >/dev/null
kubectl apply -f k8s/api-hpa.yamlhorizontalpodautoscaler.autoscaling/api created
Now one Pod fires 100 small requests a second at the API Service for 100 seconds, while a loop samples the HPA:
kubectl run load --image=localhost:33500/booknest-api:1.4 --restart=Never --command -- \
node -e "setInterval(() => fetch('http://api:3000/api/books').catch(() => 0), 10)" >/dev/null
J='{.status.currentMetrics[0].resource.current.averageUtilization}% '
J+='{.status.currentReplicas} -> {.status.desiredReplicas}{"\n"}'
for i in $(seq 11); do
kubectl get hpa api -o jsonpath="$J" | sed "s/^/$(date +%T) /"
[ "$i" = 5 ] && kubectl delete pod load --wait=false >/dev/null
sleep 20
done
git add k8s/api-hpa.yaml k8s/api-deployment.yaml
git commit -qm "Autoscale the API between 2 and 6 replicas on CPU"01:20:24 2% 3 -> 3 01:20:44 2% 3 -> 3 01:21:04 32% 3 -> 3 01:21:25 74% 3 -> 3 01:21:45 86% 3 -> 4 01:22:05 81% 4 -> 4 01:22:26 56% 4 -> 4 01:22:46 50% 4 -> 4 01:23:06 2% 4 -> 4 01:23:26 2% 4 -> 3 01:23:46 2% 3 -> 2
Utilization climbed as metrics caught up, the HPA added a fourth replica about 80 seconds into the load, and the average fell back below the target. Once the load stopped, it stepped down to the floor of two. In production keep the five-minute default or longer for scale-down and scale up fast: slow scale-up is felt by users, slow scale-down on the bill.