Full lesson
Explore the full explanation, examples, and visuals at your own pace.
Traffic rises; the same two Pods do more work
More requests don’t directly add Pods. They make the same two Pods work harder, raising CPU load; HPA uses that measurable change to decide when the application needs more capacity.
HPA checks a metric against a target
HPA doesn’t react to each request directly. Pods provide CPU data through the Metrics API, and HPA periodically asks that API for a reading. It compares measured utilization with its configured target. Traffic matters only because it can change how busy the Pods are.
Observed CPU is above the target
The chart shows average CPU utilization at an illustrative 80 percent, above the 50 percent target. These are CPU utilization percentages, not request counts; the gap gives the HPA a signal to recommend more capacity.
Chart values
| Label | Percent |
|---|---|
| Target | 50 |
| Observed | 80 |
Two Pods lead to a four-Pod recommendation
With two Pods, HPA applies observed CPU divided by the target: two times 80 over 50 is 3.2, which rounds up to a recommendation of four Pods. Configured limits and scaling rules can affect that recommendation.
With two Pods at 80 percent average CPU and a 50 percent target, how many Pods does HPA recommend?
Let's think this through. With two Pods at 80 percent average CPU and a 50 percent target, how many Pods does HPA recommend? A: Two Pods, because traffic is not measured directly. B: Three Pods, rounding the calculation down. C: Four Pods, rounding the calculation up. Choose an answer, or just think it through. I'll explain in a moment.
- Two Pods, because traffic is not measured directly
- Three Pods, rounding the calculation down
- Four Pods, rounding the calculation up
With two Pods at 80 percent average CPU and a 50 percent target, how many Pods does HPA recommend?
The answer is C: Four Pods, rounding the calculation up. Two times 80 divided by 50 is 3.2, which rounds up to four replicas. This is a recommendation, not a guarantee that two new Pods are already ready.
- Two Pods, because traffic is not measured directly
- Three Pods, rounding the calculation down
- Four Pods, rounding the calculation up
A recommendation becomes new, ready Pods
The HPA changes the Deployment’s desired count to four; it doesn’t create instant capacity. The Deployment must create two additional Pods, and each needs to be scheduled, start successfully, and become ready before it can serve requests. That delay separates the recommendation from usable capacity.
The next measurement closes the loop
HPA steers replica count through repeated measurements and corrections. After the Deployment adjusts the Pods, HPA measures CPU again; new ready Pods may share incoming requests and lower average CPU.
When traffic falls
When traffic eases, lower measured CPU can make HPA recommend fewer Pods. Scale-down stabilization can delay removal, while missing metrics or configured bounds can change the decision.
- Lower CPU can recommend fewer Pods
- Scale-down stabilization avoids quick reversals
- Limits and available metrics constrain decisions
Measure, adjust, measure again
HPA repeatedly compares an observed metric with a target and adjusts a workload's desired replica count. Traffic changes CPU, HPA changes desired replicas, ready Pods change capacity, and that capacity shapes the next measurement.




