Separately from how you present the number, the very concept of "uptime" as a single number is a bit muddy in the context of a distributed system, where different components can be differently available for different users.
Also, 0.1% downtime in the form of a 45-minute outage per month is very different from 0.1% of requests failing in brief bursts. You often see downtime reported as "increased error rates" which is so vague as to be meaningless.
Google's "windowed user-uptime" attempts to deal with this a bit better, by exposing different views of the data instead of trying to condense uptime into a single number: <a href="https://www.usenix.org/system/files/nsdi20-paper-hauer.pdf" rel="nofollow">https://www.usenix.org/system/files/nsdi20-paper-hauer.pdf
Exactly. If my build and test CI takes several hours and it gets interrupted, it really doesn’t matter how long the interruption was. It impacts me all the added time of realizing it stopped, investigating and confirming why it stopped, triggering another run, and continued monitoring.
If anyone was tracking that time they'd realize running your own build servers is cheaper. But capex is kryptonite to MBAs so you get a shitty unreliable cloud service instead.
teraflop · · focus · HN ↗
Also, 0.1% downtime in the form of a 45-minute outage per month is very different from 0.1% of requests failing in brief bursts. You often see downtime reported as "increased error rates" which is so vague as to be meaningless.
Google's "windowed user-uptime" attempts to deal with this a bit better, by exposing different views of the data instead of trying to condense uptime into a single number: <a href="https://www.usenix.org/system/files/nsdi20-paper-hauer.pdf" rel="nofollow">https://www.usenix.org/system/files/nsdi20-paper-hauer.pdf
olsondv · · focus · HN ↗
kevin_thibedeau · · focus · HN ↗