Near full utilization, load turns randomness into long waits

Donald Reinertsen's chapter on queues in The Principles of Product Development Flow [1] targets a habit most managers share: idle people look like waste, so teams get loaded until everyone is busy. His chart of queue size against capacity utilization stays low through the middle, then turns nearly vertical near 100%. Why does the last stretch cost so much when capacity still exceeds the work?

Because load only amplifies the waiting; randomness creates it. If work arrived like clockwork and every task took the same time, each item would arrive just as the last one finished, and nothing would wait even at 99% busy. Real work clumps and varies, so bursts leave backlogs that only idle time can clear. At utilization ρ\rho, new work keeps arriving at ρ\rho of the service rate, so a backlog drains at just 1−ρ1-\rho of it.

Kingman's formula for the average wait near saturation makes this a product of the two [2]:

Wq≈ρ1−ρ⋅ca2+cs22⋅τW_q \approx \frac{\rho}{1-\rho}\cdot\frac{c_a^2+c_s^2}{2}\cdot\tau

where τ\tau is the average service time and cac_a and csc_s are the coefficients of variation (standard deviation over mean) of the gaps between arrivals and of the service times. Clockwork makes the middle factor 0; fully random arrivals and service make it 1.

Waiting time against utilizationAverage wait in units of service time, rho over one minus rho times the middle factor, for factors 2, 1, one half and 0. All curves stay low through the middle range and turn nearly vertical approaching 100 percent; the last 15 percent of utilization is shaded. At factor 1 the wait is 4 service times at 80 percent and 19 at 95 percent. At factor 0, clockwork work, the wait is zero at every utilization.last 15%0102030400%20%40%60%80%100%wait ÷ service timeutilization ρ419
  • Middle factor:
  • 2 burstier than random
  • 1 fully random
  • ½ random arrivals, same-size tasks
  • 0 clockwork
Kingman’s approximation of the average wait, in units of service time. The shaded band is the last 15% of utilization; clockwork work lies flat on the axis.

With fully random work, going from 80% to 95% busy raises the wait from 4 to 19 service times: nearly five times as long for 19% more load. I'd have said the wait blows up exponentially near saturation, but the curve is a hyperbola: the wait roughly doubles each time the idle share halves, and it goes to infinity at 100%.

For capacity planning, the headroom a system needs depends on how variable its work is. At 95% busy, making every task the same size while arrivals stay random halves the factor, cutting the wait from 19 to 9.5 service times, the same as dropping to about 90% busy.

References

  1. The Principles of Product Development Flow: Second Generation Lean Product Development
    Reinertsen, D. G., 2009. Celeritas Publishing.

  2. The single server queue in heavy traffic [link]
    Kingman, J. F. C., 1961. Mathematical Proceedings of the Cambridge Philosophical Society, Vol 57(4), pp. 902–904. Cambridge University Press (CUP). DOI: 10.1017/s0305004100036094