Scaling

What stalled Salesforce for ten hours?

One internal login service saturated and a global platform queued behind it. The September outage, read as a briefing on shared dependencies.

By TIS Partners · · 1 min

On 16 September, hundreds of Salesforce instances stalled from roughly 08:30 to 19:20 UTC, across the US, Japan, India and most of Europe, during the company's own Dreamforce week. The interim explanation is one sentence, and it is the whole briefing: requests were stalling while waiting on a response from an internal login service, which was using up available server resources.

That is saturation by proxy. The login service did not crash; it slowed. Every request that queued behind it held a thread, a connection, a slice of capacity on some healthier machine, until the platform's spare resources were spent waiting on one dependency. A slow dependency with no timeout and no shedding is more dangerous than a dead one, because dead things fail fast and slow things recruit their callers into the outage.

Our standing question fits in the margin: which of your internal services sits on the request path of everything, and what happens upstream when it answers in seconds instead of milliseconds? A timeout budget is an architecture decision. Recovery took region by region, with manual restarts. The queue always drains slower than it filled.