CPU is a proxy; an API may scale better on requests per second, a worker on queue depth. Besides Resource, an autoscaling/v2 HPA accepts ContainerResource, Pods (a per-Pod custom metric), Object (a value from another object) and External (a value from outside the cluster). The last three need an adapter serving custom.metrics.k8s.io or external.metrics.k8s.io, and this cluster has none:
kubectl get apiservices | grep -e NAME -e metricsNAME SERVICE AVAILABLE AGE v1beta1.metrics.k8s.io kube-system/metrics-server True 33s
Prometheus Adapter 2,094 (github.com/kubernetes-sigs/prometheus-adapter (https://github.com/kubernetes-sigs/prometheus-adapter 2,094 ), v0.12.0) turns PromQL queries into custom metrics. KEDA 434,961 (github.com/kedacore/keda (https://github.com/kedacore/keda 10,549 ), a CNCF graduated project, v2.21.0) brings 70+ scalers (Kafka 129 , RabbitMQ 28,807 , SQS, Prometheus 16,091 , cron) and can scale a Deployment to zero, which a plain HPA cannot. With several metrics, the HPA takes the largest replica count.