‹ BackHN Continuity

Thread

Salesforce Global Outage

280 points · 184 comments · mabil

  1. stmw · · focus · HN ↗
    Despite all of the snark here, in my experience Salesforce SRE team is quite competent. The engineering challenges of running a large PaaS - not just with own apps, but with millions of customer-written apps running on it - are quite interesting, and sadly things happen. The status page makes sense to actual customers, it's the particular "pods" where a given service runs.
    1. Anon1096 · · focus · HN ↗
      Hacker News is much easier to read when you realize that 95% of people have never worked on a "high" (maybe we could say >1B requests per day as a starting point) scale distributed service and think it's trivial to run one with more than 2 nines. You see comments all the time here mentioning that their own desktop at home is achieving more than that which belies deep misunderstanding of how systems are measured. Or that unofficial github status page repeatedly posted here that counts all github services together into one number.
      1. ocdtrekkie · · focus · HN ↗
        > which belies deep misunderstanding

        I think you are missing the point. When I state my Exchange server is more reliable than Exchange Online, I don't think I'm a better engineer. I recognize Microsoft has harder problems to solve than I do. I think building overengineered, oversized SaaS environments is introducing extreme risk. It's an inherent flaw of the current approach.

        Smaller is, in fact, better, because it's easier to operate reliably.

        1. toomuchtodo · · focus · HN ↗
          Indeed, the scale Anon1096 refers to wrt distributed systems is anti pattern. It is designed to vacuum up revenue and create enterprise value with scale, not to create resiliency for customers (although resiliency might be a byproduct of a well architected and operated distributed system at scale).

          "Simplicity is the ultimate sophistication." -- Da Vinci

          1. a_conservative · · focus · HN ↗
            Hidden in this discussion around self-hosting reliability are other options as well.

            Depending on your time and appetite for tinkering with all of this, it's not hard to imagine a home setup that fails over to a cheap Hetzner or DO VM. A manual failover at the DNS level isn't overly complex, and could be scripted.

            Keeping a database in sync between home and the instance might be simple or more complex depending on needs, but would it really be that hard to have Claude help you setup a replicating Postgres server? If your database (or data files) are 1 gigabyte and don't update that often... maybe just rsync it every night or something

            There's a thread you and others are pulling on here, and we need to pull it. Hosting doesn't have to be the domain of the big vendors anymore.

            1. toomuchtodo · · focus · HN ↗
              That was my intent, pull the thread.
        2. jedberg · · focus · HN ↗
          Is it? When your internet is out for five days because your ISP takes a few days to get to you, do you acknowledge that you're now at 98.5% availability for the year, far worse than any SaaS email service?

          I think people forget that those large environments are there for a reason. To make sure the service stays up in the face of problems outside your own control.

          1. toomuchtodo · · focus · HN ↗
            In my entire adult lifetime (mid 40s), my ISP has never been out for five days. Compare to Github, Microsoft, Salesforce, and AWS outages that are always occurring in some fashion. Reddit is down constantly in various ways and still continues to operate as a business, public no less, so I disagree about the need to chase five nines and broadly speaking, large distributed systems that are potentially unnecessary for the use case and target outcome.

            <a href="https:&#x2F;&#x2F;hn.algolia.com&#x2F;?dateRange=all&amp;page=0&amp;prefix=false&amp;query=%22is%20down%22&amp;sort=byDate&amp;type=story" rel="nofollow">https:&#x2F;&#x2F;hn.algolia.com&#x2F;?dateRange=all&amp;page=0&amp;prefix=false&amp;qu...

            1. jedberg · · focus · HN ↗
              Consider yourself lucky that you’ve never been the victim of a fiber cut. But what about if the power to your house goes out? Or what if your server blows the power supply?

              My entire point is that you have no redundancy in your system and you also aren’t big enough to have any pull with the vendors who can fix these types of outages so you’re basically at the mercy of your providers with no recourse.

              That’s why these systems are built the way they are.

              And generally four nines is considered the gold standard these days. I can tell you for sure that both Netflix and Ebay would lose money anytime they drop below four nines because I have at some point been responsible for both. You’re correct that Reddit has a lot more leeway and outage time before they start losing money but not that much leeway.

              1. toomuchtodo · · focus · HN ↗
                I&#x27;ve contributed to building out data centers, as well as managed colos for others, primarily in downtown Chicago at Level3 and at 350 E Cermak. I am familiar with architecture required for reliability and diversity, from power and fiber in all the way up the stack to the Kubernetes cluster and software defined networking. If you participate in the capital markets, your data traverses systems I&#x27;ve participated in designing and implementing. There is a time and place for complexity (in this context, large&#x2F;global distributed systems), but too often, complexity exists where it need not (imho).

                &quot;What are you optimizing for?&quot; is always an important question, as is &quot;The Five Whys.&quot;

              2. ocdtrekkie · · focus · HN ↗
                I&#x27;ve dealt with a fiber cut, it wasn&#x27;t nearly that bad. Fiber cuts impacting my SaaS providers were worse because there was nothing I could do about it.
              3. Semaphor · · focus · HN ↗
                I mean, if you really need redundancy, isn’t a second instance on a VPS somewhere that you manually switch over to, enough?

                Over multiple ISPs, so far internet outages for more then a few minutes is very rare (though the few minutes would make me not want to host something requiring high availability; and a cut cable is really annoying because there simply is no quick fix), power outages even rares, I experienced 3 in 40 years, and the longest was 6 hours.

                1. chias · · focus · HN ↗
                  Generally, no.

                  Ignoring for now how you are synchronizing the database and filesystem, and how doing so may well result in your duplicate experiencing the same failure as the original, you can maybe recover from a small class of availability issues that could knock you out of an SLA.

                  But that assumes you can get online and can fully orchestrate the transition within less than 53 minutes of it starting. Including the time you took to become aware of it. And including the time to diagnose and decide that a switchover would resolve the problem. Including the time it takes for DNS caches to expire and point to the new host. Including the DNS caches which may ignore your TTL. And including all these things again when you switch back.

                  And assuming, of course, that it doesn&#x27;t happen again for a whole year.

              4. torginus · · focus · HN ↗
                &gt; That’s why these systems are built the way they are.

                Built how? Because I can state with confidence that I have cleaned up a ton of failed upgrades&#x2F;zombie terraform deploys of these serverless kubernetes wonders that followed every best practice under the sun, and these things are not really considered even moderately reliable (as designed by imperfect mortals under real world conditions), meanwhile professionally, people who stand to lose a lot of money should their systems go down generally operate systems whose architectures were designed decades ago, are generally horizontally scaled monoliths, and are extremely conservative in software choice.

                Also downtime often is no biggie, as long as it&#x27;s planned and or don&#x27;t lose (too much) customer critical data.

                Like nodobody cares if your test db cluster goes down for the weekend. We even shut down our db instances to save money.

              5. fc417fc802 · · focus · HN ↗
                None of that is an argument against small being more reliable. Rather it&#x27;s an argument that the smaller you are the more important being distributed becomes when managing mundane day to day failures.

                Even if a solar flare takes out an entire continent or two I think it&#x27;s safe to say that the bittorrent network will still be running in some form. Can you be so certain about any given SaaS product?

            2. temp_praneshp · · focus · HN ↗
              Curious, what AWS outage has affected you for days?

              (I hope you&#x27;ll agree that the middle east outage is a true outlier)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.