Scaling StatefulSets

Scaling and Upgrading a StatefulSet Safely

Scaling a StatefulSet adds Pods, not replication. A partition limits a rolling update to Pods whose index is at least the partition number, so a change can be tried on the highest replica first:

A second replica, a partitioned update, and a careful scale-downShell
kubectl scale sts postgres --replicas=2 >/dev/null && kubectl rollout status sts/postgres | tail -1
kubectl exec postgres-0 -- getent hosts postgres-1.postgres
sql() { kubectl exec -i postgres-$1 -- psql -U booknest -d booknest -tAc "$2"; }
sql 1 'SELECT count(*) FROM books' 2>&1 | head -1
kubectl patch sts postgres -p '{"spec": {"updateStrategy": {"rollingUpdate": {"partition": 1}}}}'
kubectl set env sts/postgres TZ=UTC >/dev/null && kubectl rollout status sts/postgres | tail -1
kubectl get pods -l tier=db -o custom-columns=NAME:.metadata.name,\
REVISION:.metadata.labels.controller-revision-hash
kubectl scale sts postgres --replicas=1 >/dev/null && kubectl set env sts/postgres TZ- >/dev/null
kubectl patch sts postgres -p '{"spec": {"updateStrategy": {"rollingUpdate": {"partition": 0}}}}'
PV=$(kubectl get pvc data-postgres-1 -o jsonpath='{.spec.volumeName}')
kubectl delete pvc data-postgres-1 && kubectl delete pv $PV
Output
partitioned roll out complete: 2 new pods have been updated...
10.244.1.17     postgres-1.postgres.booknest.svc.cluster.local
ERROR:  relation "books" does not exist
statefulset.apps/postgres patched
partitioned roll out complete: 1 new pods have been updated...
NAME         REVISION
postgres-0   postgres-5b79c96694
postgres-1   postgres-67f8796596
statefulset.apps/postgres patched
persistentvolumeclaim "data-postgres-1" deleted from booknest namespace
persistentvolume "pvc-cf3908f6-12c6-4373-a3d9-397e4db638a1" deleted

postgres-1 came up after postgres-0 and answers by its own DNS name, but it is a second, empty PostgreSQL 1,289 . With the partition at 1, the template change (a TZ variable standing in for a new image tag) reached only postgres-1. Scaling down left its claim and, with the Retain class, its volume for manual deletion.

A minor upgrade (18.x) is an image change applied Pod by Pod, highest index first; a major one (18 to 19) needs pg_upgrade or a dump and restore. OnDelete leaves restarts to you, maxUnavailable (beta since 1.35) speeds up large sets, and a PodDisruptionBudget limits what a node drain may take. Replication and failover come from the database or an operator (PostgreSQL Operator), never from replicas alone.