‹ BackHN Continuity

Thread

We're going to need default hard budget caps on pretty much everything

612 points · 304 comments · elffjs

  1. motionlessveloc · · focus · HN ↗
    I used to work on a support team of a well known backend type service that had hard budget caps.

    It was, unfortunately, a nightmare. There were tons of tickets and even threats of lawsuits from customers whose service got cut off hard at the worst possible time due to organic growth/going viral/big event/nobody knew about the limit/etc. Not only did they lose all the leads and revenue they would have gotten from that bump, but they also pissed off their own existing users who suddenly couldn't use the service either.

    Generally speaking, it's much better to use alerts instead of hard limit. Even in the worst case (hackers pwn your credentials and mine Bitcoin or whatever) the rest of your business is unaffected and you can negotiate with the billing department at comparative leisure.

    This is all assuming you have humans operating the service. If you're letting AI agents yolo infra in prod, you have a whole series of new problems.

    1. aenis · · focus · HN ↗
      I think it's not generally 'much better' to use alerts. People - end users - are by now quite used to seeing things go down for a while. No biggie. But a infra oopsie can kill a company in ways a short outage won't.

      And its of course not just people yolo'ing with AI. People were quite capable of causing such outages themselves just fine. Distributed, serverless systems are hard.

      1. traceroute66 · · focus · HN ↗
        > I think it's not generally 'much better' to use alerts.

        Anyone who says its 'much better' to use alerts instead of hard caps needs to Google 'alert fatigue'.

        Alerts are soft. You ignore them or miss them, nothing happens except you spending $$$$$$$$$$ more.

        Great if you're the cloud provider raking in the cash, but a poor way to run your infrastructure.

        Hard caps force you to implement correctly (to control costs in the first place) and have correct monitoring in place (to keep a healthy cap buffer).

        So it means you can't just vibecode some slop and blindly devops it via Github CI/CD. You actually need to think and reason about your infrastructure.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.