A request is what the scheduler reserves (Scheduler); a limit is a ceiling the kernel enforces through cgroup v2. At the ceiling, CPU is compressible: the container is throttled to its quota per 100 ms. Memory is not: the kernel's OOM killer ends the container with exit code 137:
kubectl apply -f - <<'EOF' >/dev/null
apiVersion: v1
kind: Pod
metadata: { name: hog }
spec:
restartPolicy: Never
containers:
- name: hog
image: localhost:33500/booknest-api:1.4
command: [node, -e, 'const a = []; setInterval(() => { a.push(Buffer.alloc(8 << 20, 1));
console.log(a.length * 8, "MiB") }, 1000)']
resources: { requests: { cpu: 100m, memory: 64Mi }, limits: { cpu: 500m, memory: 128Mi } }
EOF
kubectl wait --for=condition=Ready pod/hog --timeout=30s >/dev/null
kubectl exec hog -- sh -c 'cd /sys/fs/cgroup && cat cpu.max cpu.weight memory.max'
sleep 20; kubectl logs hog | tail -1
kubectl get pod hog -o jsonpath='{.status.containerStatuses[0].state.terminated}' \
| jq -c '{reason, exitCode}'Output
50000 100000
17
134217728
112 MiB
{"reason":"OOMKilled","exitCode":137}cpu.max allows 50 ms per 100 ms (the 500m limit), the 100m request became a relative cpu.weight of 17 (a full core maps to 100), and memory.max is 128 MiB. Set memory limits above the measured peak so leaks die early; for CPU, many teams set a request and no limit, as BookNest does, so idle cores are not wasted.