‹ BackHN Continuity

Thread

Can we stop with the uptime percentages?

146 points · 115 comments · surprisetalk

  1. dmurray · · focus · HN ↗
    I usually tell people you don't need as much reliability as you think.

    Three nines reliability is great for most purposes. 8 hours downtime a year.

    If your system produces money at a constant rate, it captures 99.9% of the available money. Even two nines or one nine might be pretty good on that basis, when the alternative is spending 2x or 10x as much - let's build another unreliable system with that money that captures some other independent market opportunity.

    Poor reliability is a problem where you need to chain many systems together, or where the cost of a single failure is very large compared to a success. Or - as happens commonly because of load - if your periods of unreliability are correlated with periods of maximum opportunity, like an e-commerce site failing on Black Friday or a trading system failing when the market is most busy. But if you don't have one of those cases, evaluate whether investing in reliability is actually worth it to you.

    GitHub is an example where two nines of reliability ought to be OK. The argument against it is that it's bad marketing to have an unreliable service, especially one aimed at software engineers. And if GitHub is largely a marketing play by Microsoft anyway (do they really make back its cost in enterprise subscriptions?) then marketing considerations need to drive its reliability.

    1. prennert · · focus · HN ↗
      The thing is: its rarely constant rate anywhere.

      A short outage can snowball very easily in a lot of lost time. What I learned when working with enterprises is that above all they value reliability. This is for a reason.

      A short outage might at best trigger loads of paperwork for multiple hierarchies, big meetings etc. The org has no choice. It needs to evaluate if whatever happens is a threat to their business.

      In the worst case it is that, plus missing some crucial windows of delivery. This is because a system that is unavailable for a short time can cause backlogs that, like traffic jams, cascade as everyone has to slow down and then synchronously speed up again.

      Orgs have the option to create more resilience, but that is overhead similar to compliance. You need to drill all your backup plans all the time, otherwise they are worthless. The drills cost time and money. At scale it is infeasible to be robust to all failures. Therefore, enterprises (at least) often prefer reliable systems over sophisticated systems. Because this delegates the risk management to the vendors rather than adding an overhead to every employee. Because at some point the employee would just do drills all the time instead of work.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.