For example we had a 6 9 (99.9999%) requirement from a customer for any given 3-6 month period. If we violated that, we owed them their money back (baring the outage wasn’t caused by us - I.e our cloud provider shit the bed).
That’s something like 7.5 seconds. For a contract over $1.5M. Am I the only one who thinks that’s outrageous expectations?
EDIT: The web app was for generating SBOMs of static assets.
Well, we don't know the application. Six nines is appropriate for some applications. If you're doing PSIP lookups to route 911 calls, that's reasonable. If you're sending out paper mailers, it's not.
But six nines gives you 7.9 seconds a quarter. If you run a multihost system, that translates to ~1s dead host detection and switch and 3-4 switches per quarter. It's acheivable with reliable hardware and reasonable software. Otoh, it's very hard to hit if you need to move traffic to a different location to respond to a no notice location failure (failed automatic transfer switch, all fiber paths severed by construction because the redundant paths were in the same bundle, etc). If you have an out for 'cloud provider failure', that probably covers location failures.
Often times a tight uptime promise like that also comes with maintenance windows. Depending on the application, degraded service or no service may be acceptable within the maintenance window.
stairlane · · focus · HN ↗
For example we had a 6 9 (99.9999%) requirement from a customer for any given 3-6 month period. If we violated that, we owed them their money back (baring the outage wasn’t caused by us - I.e our cloud provider shit the bed).
That’s something like 7.5 seconds. For a contract over $1.5M. Am I the only one who thinks that’s outrageous expectations?
EDIT: The web app was for generating SBOMs of static assets.
toast0 · · focus · HN ↗
But six nines gives you 7.9 seconds a quarter. If you run a multihost system, that translates to ~1s dead host detection and switch and 3-4 switches per quarter. It's acheivable with reliable hardware and reasonable software. Otoh, it's very hard to hit if you need to move traffic to a different location to respond to a no notice location failure (failed automatic transfer switch, all fiber paths severed by construction because the redundant paths were in the same bundle, etc). If you have an out for 'cloud provider failure', that probably covers location failures.
Often times a tight uptime promise like that also comes with maintenance windows. Depending on the application, degraded service or no service may be acceptable within the maintenance window.