‹ BackHN Continuity

Thread

Salesforce Global Outage

280 points · 184 comments · mabil

  1. raffraffraff · · focus · HN ↗
    Have you tried turning it off and then on again?

    > We're no longer pursuing restarts as a path to remediation.

    Oh you have

    1. cube00 · · focus · HN ↗
      Kind of surprised they admit they're going to try restarting and see what happens. I'm sure it happens everywhere but nobody admits it.

      > We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue.

      At least it didn't fix the problem so they can actually start finding the real cause.

      > We're no longer pursuing restarts as a path to remediation.

      Why isn't the AI they sell telling them what's wrong? Why do they need to take shots in the dark to "see if that resolves the issue"?

      1. b112 · · focus · HN ↗
        Restart should be a very last emergency step, as if it works, a restart often might wipe out evidence of why.

        So hopefully it's not done often.

        1. wookmaster · · focus · HN ↗
          The requirement is to get customers out of impact as #1 priority. If there's suspicions around memory/thread states a restart makes a lot of sense. Digging through logs and flight records takes a lot of time, customers are losing business in that time. If you're afraid to restart your service you need to work on your telemetry.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.