Full lesson
Explore the full explanation, examples, and visuals at your own pace.
100 ms alone. Seconds under load.
A request that takes 100 milliseconds at low traffic can stretch to seconds under load. The Client still reaches the API, and the API still queries the DB, but the extra time can come from waiting for API capacity, a DB connection, or the DB itself.
Requests begin waiting
A request can be quick once the API starts processing it, yet still spend time waiting in the Queue. If requests arrive faster than the API can handle them, each new arrival waits its turn before the query reaches the DB.
Connections are limited too
After the Queue admits a request to the API, it may still wait before reaching the DB. The API must acquire a connection from the Pool, and only a limited number of queries can use those connections at once. When they're busy, the request waits.
More traffic, less CPU time per request
At low traffic, API work gets CPU time promptly. As traffic rises, runnable requests compete for the same CPU, so they may wait or progress more slowly. That’s different from waiting for a DB connection in the pool.
The API CPU has spare capacity, but requests wait before querying the DB. Where should you look first?
Let's think this through. The API CPU has spare capacity, but requests wait before querying the DB. Where should you look first? A: The DB connection pool. B: The API CPU. C: The client's screen. Choose an answer, or just think it through. I'll explain in a moment.
- The DB connection pool
- The API CPU
- The client's screen
The API CPU has spare capacity, but requests wait before querying the DB. Where should you look first?
The answer is A: The DB connection pool. A request can reach the API and still wait for a free DB connection. Spare CPU does not rule out a connection-pool bottleneck; measuring pool wait time would help confirm it.
- The DB connection pool
- The API CPU
- The client's screen
A slow DB holds connections longer
The API acquires a connection from the pool before querying the DB. If the DB responds slowly, that connection stays occupied longer, so fewer are available for other requests. They wait for a release, even though the API code hasn’t changed.
Response time includes every wait
End-to-end response time is the work plus every wait along the request path: the Client can wait in the Queue, the API can wait for the Pool, and the DB takes time to return a result. Together, those delays accumulate.
Find the bottleneck
Measure queue and pool wait times to see where requests are stalled. Compare CPU use with database latency; adding API instances won’t help if database capacity or its connections are still constrained.
- Measure queue and pool wait times
- Check CPU use and DB latency
- Fix the constrained resource first
Work plus waiting
Response time includes both work and waiting. When incoming demand approaches or exceeds a resource’s capacity, queues grow; a slow downstream service can also tie up scarce upstream connections. The 100-millisecond baseline excludes these added waits.





