Wider audience started to use status pages because the service unreliability became so much more noticeable than before and not the other way around. I never had to use a status page for bear blog or protonmail because i never had and issue with it or just i never noticed.
I am now _required_ to consult status page of github, circleci or MS services etc because i need to know why a build is not passing, why i cannot open a repo, why is my work stalling.
Percentages matter, it is just so much more obvious why they matter when it comes down to important pieces of the internet like github. And i highly doubt the number of 12 hours in the last month. MS has been downplaying the issues they have with GH performance for a while now and i don't think it is time to start to believe them yet. Maintaining these pieces of infrastructure is responsibility and a burden.
Overall i would be careful with "nonlinear significance of numbers near 100%" we are talking gh being well into the 90's this year and one number that infra people are also often being reminded about is that "1% is 3.5 days".
Things are tough for gh people and i feel for them but they are not a startup or a underdog of some sort to receive sympathy in that case.
Yeah I'm at a point where I have a bookmark folder of status pages mostly for very large orgs because I've been using those pages relatively regularly. This is not something I felt the need for historically.
Downdetector also doesn't have a motivation to lie.
I've yet to find a status page that wasn't lying about the actual status.
Also 97% up is bullshit for the 3% of people who are offline.
Saucelabs was doubly bad for this because I'm absolutely certain based on traces that they had some sort of demux bug where they would send events from their tunnel to the wrong job. I could see it in the logs that a test timeout was often the cause of an event firing that was looking for something that never happened, because the event immediately preceding it in the script was never fired. Which meant it was either dropped or went somewhere it shouldn't.
Then it stopped one day and there was nothing in their release notes about it. Lies compounded by further lies.
That's just the most memorable example I have. Stuff like this happens all the time and with many services it plays out the same. There's a perverse incentive not to be transparent about problems with the service, so the status pages play down the intensity of the situation.
Isnt downdetector just people reporting its down though? Its useful for sure but not actually hooking into any officialy API or anything. Great for when the status page also goes down but surely a lag time
I go to status pages to find out if 1) I’m crazy, 2) if our IT fucked up DNS.
Every service I’ve ever paid for or someone paid for on my behalf has gaslit me about their status page because it’s impolitic and bad for sales to update the page before you know what’s going on, just because some users are reporting issues.
So a third party doesn’t have to deal with VPs kneecapping the engineers’ access to the status page. Or some services can’t update the status page when the site is hard down because they are so obsessed with keeping it up that they have no mitigations when they are down.
I was the one at my biggest gig that had to push to get static 404 and 500 pages uploaded to S3 so we could show something for vanity URLs even if customer ID lookup was down. And then a customer noticed they hadn’t updated since they changed their contact info and I found the job was timing out without an alert or deployment failure for five months. Five. Months. The guy who wrote it had quit, and he didn’t follow my advice on copying a batch job I’d poured way too much effort into. The damned thing was timing out after 50 minutes. I followed my own advice and got it to 4.5 minutes. Almost all of that time delta was waiting for fanout calls, which were pounding the shit out of consumer facing services. 90% of the calls he was making didn’t need to be made.
Do they do alerts ? Been also using UpDog (based on DataDog) which actually has been quite good. They have a single status page for many services who integrate their products
proxysna · · focus · HN ↗
I am now _required_ to consult status page of github, circleci or MS services etc because i need to know why a build is not passing, why i cannot open a repo, why is my work stalling.
Percentages matter, it is just so much more obvious why they matter when it comes down to important pieces of the internet like github. And i highly doubt the number of 12 hours in the last month. MS has been downplaying the issues they have with GH performance for a while now and i don't think it is time to start to believe them yet. Maintaining these pieces of infrastructure is responsibility and a burden.
Overall i would be careful with "nonlinear significance of numbers near 100%" we are talking gh being well into the 90's this year and one number that infra people are also often being reminded about is that "1% is 3.5 days".
Things are tough for gh people and i feel for them but they are not a startup or a underdog of some sort to receive sympathy in that case.
rdmuser · · focus · HN ↗
gplk · · focus · HN ↗
fmbb · · focus · HN ↗
hinkley · · focus · HN ↗
I've yet to find a status page that wasn't lying about the actual status.
Also 97% up is bullshit for the 3% of people who are offline.
Saucelabs was doubly bad for this because I'm absolutely certain based on traces that they had some sort of demux bug where they would send events from their tunnel to the wrong job. I could see it in the logs that a test timeout was often the cause of an event firing that was looking for something that never happened, because the event immediately preceding it in the script was never fired. Which meant it was either dropped or went somewhere it shouldn't.
Then it stopped one day and there was nothing in their release notes about it. Lies compounded by further lies.
That's just the most memorable example I have. Stuff like this happens all the time and with many services it plays out the same. There's a perverse incentive not to be transparent about problems with the service, so the status pages play down the intensity of the situation.
Melatonic · · focus · HN ↗
hinkley · · focus · HN ↗
Every service I’ve ever paid for or someone paid for on my behalf has gaslit me about their status page because it’s impolitic and bad for sales to update the page before you know what’s going on, just because some users are reporting issues.
So a third party doesn’t have to deal with VPs kneecapping the engineers’ access to the status page. Or some services can’t update the status page when the site is hard down because they are so obsessed with keeping it up that they have no mitigations when they are down.
I was the one at my biggest gig that had to push to get static 404 and 500 pages uploaded to S3 so we could show something for vanity URLs even if customer ID lookup was down. And then a customer noticed they hadn’t updated since they changed their contact info and I found the job was timing out without an alert or deployment failure for five months. Five. Months. The guy who wrote it had quit, and he didn’t follow my advice on copying a batch job I’d poured way too much effort into. The damned thing was timing out after 50 minutes. I followed my own advice and got it to 4.5 minutes. Almost all of that time delta was waiting for fanout calls, which were pounding the shit out of consumer facing services. 90% of the calls he was making didn’t need to be made.
fmbb · · focus · HN ↗
Melatonic · · focus · HN ↗