Hacking the Thundering Herd
You know what the Thundering Herd issue is, right? Basically, you have a whole bunch of different events that are gated on a single resource — a database for example. If all those events fire at the same time, then they all go thundering along to the database at the same time, trying to get in, and causing the database to choke. Ideally, you’d get the events to fire one at a time, but sometimes, well, shit happens.
My first exposure to this was way in the past I was the CTO of a phone company (•). Imagine if you will hundreds of thousands of VoIP phones scattered around the country, each of them connected to our data-center. Each of these phones registered themselves to our service, basically letting our service know where the phone was located on the InterTubes, and which Userthe phone was associated with. Whenever there was a call to/from a phone, we’d spin up a process for the associated User that did a bunch-a housekeeping — think fun stuff like Billing, status updates, UI notifications, and whatnot. And to keep things sane, after the call ended, we’d leave the User process up and running for a while, to avoid the overhead of starting it up again on the next call.
So, imagine steady state, at 2pm on a Wednesday afternoon. The east-coast and west-coast are both online and at work, there are tens of thousands of people on calls, everything is working well, when all of a sudden our main data-center has a power-outage, and thanks to a wiring issue, our cage lost all power (No, not this outage that I’ve written about before. A previous outage. Life was fun in those days!). One panic-stricken drive to the co-lo facility later, our Ops person figures out that a circuit-breaker has tripped in our cage. They un-trip it, and we’re back online.
Or, well, almost back online. Because, here’s basically what happens
- 1. In a well orchestrated sequence, all of our servers come up. First the database, then the application servers, the call workers, the edge-proxies, etc. all smooth as silk, because we know from orchestration.
- 2. The edge-proxies switch into ingress mode once all the health-checks have passed, and the calls start coming in.
- 3. A lot of calls, that is. Remember the people who were on calls, and got so rudely dropped? They’re all trying to get back on those calls — all tens of thousands of them.
- 4. Even more calls in fact. These same tens of thousands of people are alsocalling us on their cell-phones, calls which route to us via — surprise! — our own servers. Ugh.
- 5. Every single one of these calls tries to spin up one of those tiny Userprocesses to do housekeeping. Which promptly overloads and crashes our application servers.
- 6. Which — smooth as silk — bring themselves back up, and open themselves to the world again, repeating the above cycle (with the occasional database crash thrown in for good measure)
- 7. HELLO THUNDERING HERD…
Mind you, these were early days. We didn’t have any of the fun stuff we should have to avoid these scenarios, things like load-shedding, request caching, exponential backoff, active queue management, and the like. Even worse, we didn’t have the time to implement any of this, since, after all the customers were having issues now.
(And in case you’re wondering, thanks to a whole bunch of Fault-Tolerance that we’d baked into the infrastructure, we’d never faced this before. We were quite stupid in those days…)
So, what genius hack do we come up with to deal with this? In real time? Is it something out of Hackers (greatest movie ever!) where, under impossible time constraints, we invent something radically new?
Well, actually no.
We hacked it old school and implemented load-shedding by, literally, yanking out our uplink network connection every few seconds.
Seriously.
Our Ops guy — Stuart, who I’ll forever be grateful to — would yank out the uplink to our upstream ISP ever few seconds, and then plug it back in after a minute or so. Over and over again. Lather, Rinse, Repeat. Each time the uplink was plugged in, a couple of hundred calls would get through, and when yanked out, it would give the servers enough time to deal with spinning up those calls’ User processes.
It was profoundly low-tech, but, far more importantly, it worked
Well, actually no.
We hacked it old school and implemented load-shedding by, literally, yanking out our uplink network connection every few seconds.
Seriously.
Our Ops guy — Stuart, who I’ll forever be grateful to — would yank out the uplink to our upstream ISP ever few seconds, and then plug it back in after a minute or so. Over and over again. Lather, Rinse, Repeat. Each time the uplink was plugged in, a couple of hundred calls would get through, and when yanked out, it would give the servers enough time to deal with spinning up those calls’ User processes.
It was profoundly low-tech, but, far more importantly, it worked
Mind you, the first thing we did afterwards was figure out how to do load-shedding.
(Ok, technically speaking, the first thing we did was deal with a lot of irate customers. And then, load-shedding…)
(Ok, technically speaking, the first thing we did was deal with a lot of irate customers. And then, load-shedding…)
Ah, those were the days…
(•) If anybody, suggests that you get into the phone business, back away without making contact, and at the first available opportunity, RUN.


Comments