Autoscaling the API

Autoscaling BookNest's API Deployment

The stateless API suits an HPA on CPU: two replicas as the floor, six as the ceiling, and a 60-second scale-down window so the demonstration fits on a page:

k8s/api-hpa.yaml: two to six API replicas at 70% of the CPU requestYAML
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: api, labels: { app: booknest, tier: api } }
spec:
  scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: api }
  minReplicas: 2
  maxReplicas: 6
  metrics:
  - type: Resource
    resource: { name: cpu, target: { type: Utilization, averageUtilization: 70 } }
  behavior:
    scaleDown: { stabilizationWindowSeconds: 60 }

The Deployment manifest must stop setting replicas, or each kubectl 5,150 apply resets it to 3. Deleting the line alone makes the next apply remove the field, dropping the Deployment to one replica for a moment; updating the last-applied annotation first avoids that:

Handing the replica count to the HPAYAML
sed -i '/^  replicas: 3$/d' k8s/api-deployment.yaml
kubectl apply set-last-applied -f k8s/api-deployment.yaml >/dev/null
kubectl apply -f k8s/api-hpa.yaml
Output
horizontalpodautoscaler.autoscaling/api created

Now one Pod fires 100 small requests a second at the API Service for 100 seconds, while a loop samples the HPA:

A hundred seconds of load, then quietShell
kubectl run load --image=localhost:33500/booknest-api:1.4 --restart=Never --command -- \
  node -e "setInterval(() => fetch('http://api:3000/api/books').catch(() => 0), 10)" >/dev/null
J='{.status.currentMetrics[0].resource.current.averageUtilization}% '
J+='{.status.currentReplicas} -> {.status.desiredReplicas}{"\n"}'
for i in $(seq 11); do
  kubectl get hpa api -o jsonpath="$J" | sed "s/^/$(date +%T) /"
  [ "$i" = 5 ] && kubectl delete pod load --wait=false >/dev/null
  sleep 20
done
git add k8s/api-hpa.yaml k8s/api-deployment.yaml
git commit -qm "Autoscale the API between 2 and 6 replicas on CPU"
Output
01:20:24 2% 3 -> 3
01:20:44 2% 3 -> 3
01:21:04 32% 3 -> 3
01:21:25 74% 3 -> 3
01:21:45 86% 3 -> 4
01:22:05 81% 4 -> 4
01:22:26 56% 4 -> 4
01:22:46 50% 4 -> 4
01:23:06 2% 4 -> 4
01:23:26 2% 4 -> 3
01:23:46 2% 3 -> 2

Utilization climbed as metrics caught up, the HPA added a fourth replica about 80 seconds into the load, and the average fell back below the target. Once the load stopped, it stepped down to the floor of two. In production keep the five-minute default or longer for scale-down and scale up fast: slow scale-up is felt by users, slow scale-down on the bill.