‹ BackHN Continuity

Thread

Can we stop with the uptime percentages?

146 points · 115 comments · surprisetalk

  1. proxysna · · focus · HN ↗
    Wider audience started to use status pages because the service unreliability became so much more noticeable than before and not the other way around. I never had to use a status page for bear blog or protonmail because i never had and issue with it or just i never noticed.

    I am now _required_ to consult status page of github, circleci or MS services etc because i need to know why a build is not passing, why i cannot open a repo, why is my work stalling.

    Percentages matter, it is just so much more obvious why they matter when it comes down to important pieces of the internet like github. And i highly doubt the number of 12 hours in the last month. MS has been downplaying the issues they have with GH performance for a while now and i don't think it is time to start to believe them yet. Maintaining these pieces of infrastructure is responsibility and a burden.

    Overall i would be careful with "nonlinear significance of numbers near 100%" we are talking gh being well into the 90's this year and one number that infra people are also often being reminded about is that "1% is 3.5 days".

    Things are tough for gh people and i feel for them but they are not a startup or a underdog of some sort to receive sympathy in that case.

    1. rdmuser · · focus · HN ↗
      Yeah I'm at a point where I have a bookmark folder of status pages mostly for very large orgs because I've been using those pages relatively regularly. This is not something I felt the need for historically.
      1. fmbb · · focus · HN ↗
        Nothing beats downdetector.com anyway. Always quicker!
        1. hinkley · · focus · HN ↗
          Downdetector also doesn't have a motivation to lie.

          I've yet to find a status page that wasn't lying about the actual status.

          Also 97% up is bullshit for the 3% of people who are offline.

          Saucelabs was doubly bad for this because I'm absolutely certain based on traces that they had some sort of demux bug where they would send events from their tunnel to the wrong job. I could see it in the logs that a test timeout was often the cause of an event firing that was looking for something that never happened, because the event immediately preceding it in the script was never fired. Which meant it was either dropped or went somewhere it shouldn't.

          Then it stopped one day and there was nothing in their release notes about it. Lies compounded by further lies.

          That's just the most memorable example I have. Stuff like this happens all the time and with many services it plays out the same. There's a perverse incentive not to be transparent about problems with the service, so the status pages play down the intensity of the situation.

          1. Melatonic · · focus · HN ↗
            Isnt downdetector just people reporting its down though? Its useful for sure but not actually hooking into any officialy API or anything. Great for when the status page also goes down but surely a lag time
            1. fmbb · · focus · HN ↗
              No lag time for popular services. Always faster than e.g. GitHub’s or Chat GPT’s status pages.
              1. Melatonic · · focus · HN ↗
                Do they do alerts ? Been also using UpDog (based on DataDog) which actually has been quite good. They have a single status page for many services who integrate their products
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.