Full lesson
Explore the full explanation, examples, and visuals at your own pace.
Traffic rises; two Pods serve it
More requests reach the same two Pods, but traffic volume alone doesn’t tell an HPA how busy they are. What signal can it use to decide whether two replicas are still enough?
The signal is average CPU utilization
As traffic rises, the HPA averages CPU measurements from the two Pods. Utilization is relative to each Pod’s CPU request, not the total CPU capacity of its node.
HPA compares the metric with its target
As traffic rises across the two Pods, their measured average CPU utilization is reported to the HPA. It compares that reading, 100 percent of the CPU requests, with its configured 50 percent target.
Two Pods become four desired Pods
Measured CPU utilization is twice the 50 percent target, so two times 100 divided by 50 gives four desired Pods. HPA adjusts the desired Pod count, then checks the load again.
Chart values
| Label | Percent of CPU request |
|---|---|
| Measured | 100 |
| Target | 50 |
Desired replicas change before capacity does
The HPA sets the Deployment’s desired count to four. The Deployment passes that target to its ReplicaSet, which requests two additional Pods. Desired state changes first, not available capacity. Kubernetes still needs time for those Pods to become ready before they can serve traffic.
Four Pods become ready and traffic stays steady. What should HPA check next?
Let's think this through. Four Pods become ready and traffic stays steady. What should HPA check next? A: A new measurement of average CPU utilization. B: Whether the original two Pods were deleted. C: The total CPU capacity of every node. Choose an answer, or just think it through. I'll explain in a moment.
- A new measurement of average CPU utilization
- Whether the original two Pods were deleted
- The total CPU capacity of every node
Four Pods become ready and traffic stays steady. What should HPA check next?
The answer is A: A new measurement of average CPU utilization. HPA uses later metric readings to decide whether four desired Pods are still appropriate. Adding Pods does not finish the decision; it changes the load each Pod may experience, which the next measurement can reveal.
- A new measurement of average CPU utilization
- Whether the original two Pods were deleted
- The total CPU capacity of every node
New measurements close the loop
With traffic spread across four ready Pods, average CPU utilization may fall. HPA uses that later CPU measurement to reassess the desired replica count, then updates the Deployment again. That’s the feedback loop: each scaling change shapes the next measurement.
Why scaling may lag
Scaling isn’t instantaneous: metrics arrive over time, and new Pods must become ready before serving traffic. Minimum and maximum set bounds, while downscale stabilization prevents removing Pods after a brief dip.
- Metrics and Pod readiness take time
- Minimum and maximum bound desired replicas
- Downscale stabilization slows removal
Measure, adjust, measure again
HPA repeatedly compares a configured metric with its target and adjusts desired replicas within configured limits. Rising traffic changes CPU utilization; HPA updates desired replicas, new Pods become ready, and the next measurement guides another adjustment.




