I’ve worked in systems which at worse have had three nines for years, but I’ve also worked in systems where five nines is a failure.
This attitude of modern tech claiming 98% is good just doesn’t work in the old tech acceptance. We had individual components fail all the time. We’re still looking at a 230ms outage to a branch office last week caused by a power failure combined with a badly plumbed power distribution.
Modern software people don’t consider 230ms to be an outage. Glad they don’t work in electricity.
(The failure we had was only on the services we guarentee at 99.1%, our lowest sla. After that there’s 99.95 and 99.999.
(In reality we reach five nines year after year on even the lowest levels, but there are major concerns like “large bomb in data centre” which could cause some of our less critical units to drop way more than 5 minutes a year.
That's not really an old versus new thing. The water company isn't going to consider 230ms an outage either.
With tech, like, it's crazily difficult to make sure every http request succeeds, so you build in retry, and look at that 230ms doesn't cause real disruption now. Not to excuse 98%, that's awful.
I'm curious what kind of power distribution you work in. I've lived in suburban, rural, and urban places over the years and I've never experienced five 9's of uptime on the grid.
hdgvhicv · · focus · HN ↗
This attitude of modern tech claiming 98% is good just doesn’t work in the old tech acceptance. We had individual components fail all the time. We’re still looking at a 230ms outage to a branch office last week caused by a power failure combined with a badly plumbed power distribution.
Modern software people don’t consider 230ms to be an outage. Glad they don’t work in electricity.
(The failure we had was only on the services we guarentee at 99.1%, our lowest sla. After that there’s 99.95 and 99.999.
(In reality we reach five nines year after year on even the lowest levels, but there are major concerns like “large bomb in data centre” which could cause some of our less critical units to drop way more than 5 minutes a year.
Dylan16807 · · focus · HN ↗
With tech, like, it's crazily difficult to make sure every http request succeeds, so you build in retry, and look at that 230ms doesn't cause real disruption now. Not to excuse 98%, that's awful.
cco · · focus · HN ↗
You work in grid/industrial?