‹ BackHN Continuity

Thread

Salesforce Global Outage

280 points · 184 comments · mabil

  1. stmw · · focus · HN ↗
    Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.
    1. gibsonf1 · · focus · HN ↗
      Hmm, could the use of genAI have anything to do with this failure and the inability to quickly fix it?
      1. Jach · · focus · HN ↗
        It's not impossible, but Salesforce has had big outages before LLMs. For a disruption that began at 1am pacific, the response time isn't that bad. 3 hours total to give up on restarts, 4 hours total to validate a quick fix and begin rollout, and the rest of the time since has been waiting for the rollout + addressing subsets of instances that had some issues with restarting+the quick fix. It's nearly 9am pacific now, so Dreamforce is saved~ (It's Dreamforce week this week. Most devs are either focused on that or on soft-vacation / working on lower priority non-feature-work items, it's surprising anything would be updated to production this week that could do this.) The architecture and approval process of everything there has long been setup so that things can't be changed quickly.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.