Node Groups and Spot

Node Groups, Spot Instances and Cost Controls

The control-plane fee is the small line; nodes are the big one. Nodes come in groups of one shape: an EKS 24 managed node group (an EC2 24 Auto Scaling group), a GKE 1 node pool or an AKS 6 node pool (a virtual machine scale set), each sized by the Cluster Autoscaler 8,982 or Karpenter 676,493 (Cluster Autoscaler).

Spot capacity is the provider's spare machines at a discount, taken back when needed: AWS 24 advertises up to 90% off On-Demand prices and gives a two-minute interruption notice. Each provider marks its Spot nodes:

Spot node groups on the three managed services
Provider Turning Spot on Mark on the node
EKS eksctl 24 --spot or CapacityType: SPOT Label eks.amazonaws.com/capacityType=SPOT
GKE --spot on a cluster or node pool Label cloud.google.com/gke-spot=true
AKS --priority Spot on an extra node pool Label and NoSchedule taint scalesetpriority=spot

Only AKS taints Spot nodes by default, so there a Pod must tolerate the taint to land on one; on EKS and GKE, Pods land on Spot nodes unless you keep them off. BookNest's API and front end qualify for Spot: stateless, at least two replicas, quick to stop. PostgreSQL 1,289 does not, so pin its StatefulSet to On-Demand nodes with a nodeSelector of eks.amazonaws.com/capacityType: ON_DEMAND (on EKS).

Give a Spot group several instance types of the same size, so one type's shortage does not empty it, and keep terminationGracePeriodSeconds short: AWS warns that Pods may not get the full two minutes. Beyond Spot, right-size requests (Probes and Resources), let the HPA scale in (Autoscaling Pods and Nodes), and delete test clusters: a forgotten one bills 730 hours a month.