Kind of surprised they admit they're going to try restarting and see what happens. I'm sure it happens everywhere but nobody admits it.
> We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue.
At least it didn't fix the problem so they can actually start finding the real cause.
> We're no longer pursuing restarts as a path to remediation.
Why isn't the AI they sell telling them what's wrong? Why do they need to take shots in the dark to "see if that resolves the issue"?
I wouldn't say so, rather you need to balance recovery time and evidence preservation. A good incident manager will give the service owning team a chance or two to debug, but not let them fall into the trap of needing to understand the problem fully before attempt a clumsy potential fix. And of course will take into account the total business impact of the ongoing disruption and the known and unknown risks of the proposed clumsy fix (it could make things worse).
raffraffraff · · focus · HN ↗
> We're no longer pursuing restarts as a path to remediation.
Oh you have
cube00 · · focus · HN ↗
> We've attempted a rolling restart on one of the impacted instances to see if that resolves the issue.
At least it didn't fix the problem so they can actually start finding the real cause.
> We're no longer pursuing restarts as a path to remediation.
Why isn't the AI they sell telling them what's wrong? Why do they need to take shots in the dark to "see if that resolves the issue"?
b112 · · focus · HN ↗
So hopefully it's not done often.
CoffeeOnWrite · · focus · HN ↗
b112 · · focus · HN ↗