‹ BackHN Continuity

Thread

Salesforce Global Outage

280 points · 184 comments · mabil

  1. stmw · · focus · HN ↗
    Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.
    1. trebligdivad · · focus · HN ↗
      What I'm curious about is why it is a single-PaaS; I'd have expected Salesforce to have the customers quite isolated so the chance of bringing down multiple customers at once was much smaller.
      1. kakwa_ · · focus · HN ↗
        You still have fleet wide management which can cause issues. Plus there are always a few core services like queues, authz, presentation layer.

        Also, mono-tenant architectures is no golden bullet either. Such architecture (often coming from a formerly on-premise product that was SaaS-ified) can easily become hell to operate as it multiplies the integration points (DB parameters, URLs, allowlists, etc).

        It's also quite wasteful in terms of resource utilization and hosting costs.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.