‹ BackHN Continuity

Thread

Can we stop with the uptime percentages?

146 points · 115 comments · surprisetalk

  1. stairlane · · focus · HN ↗
    These metrics tend to be bullshit in contracts.

    For example we had a 6 9 (99.9999%) requirement from a customer for any given 3-6 month period. If we violated that, we owed them their money back (baring the outage wasn’t caused by us - I.e our cloud provider shit the bed).

    That’s something like 7.5 seconds. For a contract over $1.5M. Am I the only one who thinks that’s outrageous expectations?

    EDIT: The web app was for generating SBOMs of static assets.

    1. usernametaken29 · · focus · HN ↗
      I worked in realtime trading. No. Not at all outrageous. Quite reasonable actually. If that’s what we agreed and I need you to be reliable I will charge you back for being unreliable. I’m happy to pay top and extra dollar for the SLA but that means it needs to be acted on.
      1. stairlane · · focus · HN ↗
        I think in your case it’s reasonable. Also a $1.5M contract is likely chump change in that realm.

        This particular case was in cybersecurity- specifically static analysis of assets, for the purpose of providing a SBOM.

    2. ahtihn · · focus · HN ↗
      It's a contract, you're free to negotiate?

      If you agree to those terms knowing it's unrealistic, you're agreeing to give away your service for free.

      1. swiftcoder · · focus · HN ↗
        > It's a contract, you're free to negotiate?

        Well, someone on the business side of the house is free to negotiate. Whether engineering learns about the contract before sales has inked a 6-nines availability guarantee varies wildly by the company

        1. ahtihn · · focus · HN ↗
          That's a common issue in orgs where sales are incentivized based on size of contract without regard for the profitability of the contract.

          See also: selling features that don't exist and cost more to implement and maintain than the contract is worth.

    3. jerf · · focus · HN ↗
      Are you sure it wasn't prorated?

      If it wasn't prorated anyone who approved the contract needs training and/or firing. If it is prorated, that is generally not a problem. Small outages aren't even worth the effort of trying to get the money back, and if you have a large enough one to make it worthwhile it is likely the prorated refund is still going to be laughably small.

      1. stairlane · · focus · HN ↗
        I think it was prorated but am not 100% sure.
    4. toast0 · · focus · HN ↗
      Well, we don't know the application. Six nines is appropriate for some applications. If you're doing PSIP lookups to route 911 calls, that's reasonable. If you're sending out paper mailers, it's not.

      But six nines gives you 7.9 seconds a quarter. If you run a multihost system, that translates to ~1s dead host detection and switch and 3-4 switches per quarter. It's acheivable with reliable hardware and reasonable software. Otoh, it's very hard to hit if you need to move traffic to a different location to respond to a no notice location failure (failed automatic transfer switch, all fiber paths severed by construction because the redundant paths were in the same bundle, etc). If you have an out for 'cloud provider failure', that probably covers location failures.

      Often times a tight uptime promise like that also comes with maintenance windows. Depending on the application, degraded service or no service may be acceptable within the maintenance window.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.