The 5 Stages of Load-Management
The load-management clock seems to perpetually stuck at Amateur Hour. Most of the implementation in this area seems to be something along the following lines
- 1. We’ll deal with load when we have load
Which is not all that bad an idea, especially given that the other option tends to usually be “Design something appropriate for Google” - 2. (Now that we have load) Put A Queue On ItAnd spend a huge amount of time here, increasing the number of workers, optimizing their performance, “scaling horizontally”, and whatnot, till your queue-limit is the problem. Somewhere along the way, you also end up at
- 3. Increase the Buffer Size
Which promptly leads to bufferbloat — basically, the latency that comes from having your requests spend waaaaay too much time in the queue, and all the emergent properties that come from this (the same request ends up in the queue multiple times as clients start saturate-bombing, etc.).
The sad thing is that bufferbloat is quite straightforward to deal with, but that requires that you know that it exists in the first place. Instead, you invariably end up with - 4. Discover Tail Drop (aka: “Drop Stuff When The Queue Fills Up”)
Which is not all that bad an idea, except that you are now firmly in emergent properties land — if your clients implement any sort of backoff, they’ll probably all go into back-off, clobbering performance all over the place (and that’s just the start).
By now, you think people would have learned, except this is almost exactly when folks end up with - 5. Reinvent Queue ManagementBy now, we are firmly in the domain of the #CowboyDeveloper, where using Teh Googles is consider heresy of the highest order.
Mind you, it’s good to see that every now and then people actually Act Smart. Uber, for example, implemented load-shedding by not doing Tail Drop, and by not reinventing Queue Management. Instead, they basically plugged in CoDel to deal with overloads. I won’t go into details (read the article for more on this) but the key is that they focus on latencies and not on traffic rates, queue size, and the like. (Note, Netflix does something similar too!).
For an excellent talk on queue management, go check out this video. And remember, always use Teh Googles!

Comments