Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.
What I'm curious about is why it is a single-PaaS; I'd have expected Salesforce to have the customers quite isolated so the chance of bringing down multiple customers at once was much smaller.
The customers are quite isolated, but it doesn't mean that some services or errors do not propagate. In public cloud terms, think back on some AWS or Azure or even Gmail outages - you probably wouldn't even hear about them if it didn't affect millions of users at once, across security and availability boundaries.
I think it's easy to underestimate how many smaller outages there are in any period of time but which do not affect anyone or only a small number of people due to all the mitigations, due diligence that developers and SREs do, and the self-repairing nature of modern systems.
stmw · · focus · HN ↗
trebligdivad · · focus · HN ↗
stmw · · focus · HN ↗
Cthulhu_ · · focus · HN ↗