‹ BackHN Continuity

Thread

Salesforce Global Outage

280 points · 184 comments · mabil

  1. raffraffraff · · focus · HN ↗
    Have you tried turning it off and then on again?

    > We're no longer pursuing restarts as a path to remediation.

    Oh you have

    1. cube00 · · focus · HN ↗
      Kind of surprised they admit they're going to try restarting and see what happens. I'm sure it happens everywhere but nobody admits it.

      > We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue.

      At least it didn't fix the problem so they can actually start finding the real cause.

      > We're no longer pursuing restarts as a path to remediation.

      Why isn't the AI they sell telling them what's wrong? Why do they need to take shots in the dark to "see if that resolves the issue"?

      1. b112 · · focus · HN ↗
        Restart should be a very last emergency step, as if it works, a restart often might wipe out evidence of why.

        So hopefully it's not done often.

        1. CoffeeOnWrite · · focus · HN ↗
          I wouldn't say so, rather you need to balance recovery time and evidence preservation. A good incident manager will give the service owning team a chance or two to debug, but not let them fall into the trap of needing to understand the problem fully before attempt a clumsy potential fix. And of course will take into account the total business impact of the ongoing disruption and the known and unknown risks of the proposed clumsy fix (it could make things worse).
          1. b112 · · focus · HN ↗
            Yes, that's why it's the last emergency step. We're not disagreeing.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.