‹ BackHN Continuity

Thread

We're going to need default hard budget caps on pretty much everything

606 points · 302 comments · elffjs

  1. motionlessveloc · · focus · HN ↗
    I used to work on a support team of a well known backend type service that had hard budget caps.

    It was, unfortunately, a nightmare. There were tons of tickets and even threats of lawsuits from customers whose service got cut off hard at the worst possible time due to organic growth/going viral/big event/nobody knew about the limit/etc. Not only did they lose all the leads and revenue they would have gotten from that bump, but they also pissed off their own existing users who suddenly couldn't use the service either.

    Generally speaking, it's much better to use alerts instead of hard limit. Even in the worst case (hackers pwn your credentials and mine Bitcoin or whatever) the rest of your business is unaffected and you can negotiate with the billing department at comparative leisure.

    This is all assuming you have humans operating the service. If you're letting AI agents yolo infra in prod, you have a whole series of new problems.

    1. hypfer · · focus · HN ↗
      I'm of course lacking specifics, but it sounds like this take-away might not be the best solution?

      A better approach to that scenario might be to split into base load and peak load infra, similar to how we do it with the power grid.

      Base load could be an actual metal server that you pay x amount of money for and is fully yours, with peak load being handled by autoscaling dynamic stuff.

      Both things being on a fixed budget, if that budget would run out, the service would not degrade as in "grind to a halt" but as in "slows down", which is probably not a downtime as part of a communicated SLA.

      My point being that cloud and on-demand is useful, but a hybrid approach might in many cases make more sense. Though of course YMMV. I do not know what the requirements of your product are.

      1. kccqzy · · focus · HN ↗
        Depending on the size of the peak, overloading your base load could very well grind it to a halt. At times of overload, you need loadshedding to recover, which is exactly the opposite of what you are proposing here.
        1. hypfer · · focus · HN ↗
          I guess that depends a lot on the specific scenario and use-case.
          1. kccqzy · · focus · HN ↗
            What use case do you have in mind? In my experience every use case I can think of won’t work with your model. The diurnal variation between nighttime demand and daytime demand (in a given timezone) is too great.
            1. hypfer · · focus · HN ↗
              I'm kinda not interested in this bad faith debate, but I'd like to call out that I perceive it as such.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.