‹ BackHN Continuity

Thread

Salesforce Global Outage

280 points · 184 comments · mabil

  1. stmw · · focus · HN ↗
    Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.
    1. Anon1096 · · focus · HN ↗
      Hacker News is much easier to read when you realize that 95% of people have never worked on a "high" (maybe we could say >1B requests per day as a starting point) scale distributed service and think it's trivial to run one with more than 2 nines. You see comments all the time here mentioning that their own desktop at home is achieving more than that which belies deep misunderstanding of how systems are measured. Or that unofficial github status page repeatedly posted here that counts all github services together into one number.
      1. torginus · · focus · HN ↗
        Well ackchually.. I get that large scale systems pose their own challenges on their own, but it also matters what's the smallest isolable unit.

        What I mean by this is a CDN consists of nodes that are horizontally replicable and don't really talk to each other, and thus are easy to run even at scale.

        In contrast, something like a bank or social media isn't really reducible - every user needs to be able to interact with every other user in a consistent manner.

        So running a midsize bank's backend which processes 10m transactions per day, might be as if not more complex (all consistent, repeatable, and must never fail), that having a product which is a 10-10k org's IT infra replicated a thousand times.

        And yes, lots of people have worked at banks and other fintech companies of this scale, including me.

        I am not an expert, as I never worked on the 'core' systems but I know folks who did, and everyone told me there's an arcane database monolith that sits at the heart of these, very expensive and exotic big box SW & HW (at least for us unwashed rubes used to EC2 instances)

        1. supriyo-biswas · · focus · HN ↗
          > What I mean by this is a CDN consists of nodes that are horizontally replicable and don't really talk to each other, and thus are easy to run even at scale.

          They do though! They mostly try to avoid it since hitting the network to serve any kind of latency would unacceptably increase latency, but you wildly underestimated the amount of complexity there is to running a CDN.

          1. torginus · · focus · HN ↗
            I probably underestimated the complexity and I didn't mean to knock on CDNs - I just wanted to say that not all distributed systems have equal complexity, and some require essentially almost serializable transactions, while others are fine with small channels of eventual consistency
            1. dastbe · · focus · HN ↗
              the complexity is in different places. For these systems the sheer scale of throughput makes reasoning about them challenging.

              the OP mentioned 1B requests/day, where there are systems handling 1B requests a second.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.