‹ BackHN Continuity

Thread

We're going to need default hard budget caps on pretty much everything

606 points · 302 comments · elffjs

  1. joshdavham · · focus · HN ↗
    It’s incredible that in 2026, AWS and GCP are only just now introducing this. It’s possibly one of the most obviously needed features for a cloud provider.

    Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.

    1. simonw · · focus · HN ↗
      It's definitely technically difficult. You can't easily estimate how much an operation is going to cost before you kick off that operation, which means as soon as you get close to the limit you are at risk of tripping it.

      Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.

      1. stuartaxelowen · · focus · HN ↗
        Advertising platforms have had this since their inception. They were just motivated because they could be left holding the bag.
      2. mcapodici · · focus · HN ↗
        Yes and then the choice is run it and forgive it, or, stop the process midway.

        If you stop then you have to decide whether to charge for uncompleted work.

        Interesting tradeoffs.

        For very small ops e.g. individual Lambda invocation you have similar concerns especially if lots are fired at once from a queue or schedule or fanout.

      3. nvader · · focus · HN ↗
        Yeah, we probably want some kind of traffic light system:

        Green means go Orange means finish what you're doing but don't start anything new Red means stop everything

        And probably a special rule to permit stable, critical spend through regardless, the same way we allow police and ambulance to run lights.

        1. Gigachad · · focus · HN ↗
          Stop everything is pretty damaging any real business though. Things were better in the era of VPSs. You paid for a fixed amount of compute, if you ran a stupidly expensive operation than it just maxed out your system for a certain amount of time and things slowed down. But you didn’t kill the service entirely and you didn’t have unlimited potential price
        2. eddythompson80 · · focus · HN ↗
          In this example, that would require the “big query” to have billing baked into its actual query runtime, which isn’t impossible, just not how one would design a query planner per se. Usually such services emit metrics of usage units, then the billing calculation happens in a completely different system taking into account discounts, promotions, contracts, regional and currency differences, etc.

          Suddenly a database, a storage service or a computer service needs to be aware of the billing situations and make behavioral decisions based on the billing status. Again, not impossible, but something that suddenly promotes billing from an async/non-crucial background service that can be paused, replayed, adjusted by account teams etc, into a crucial hot-path service.

          1. miki123211 · · focus · HN ↗
            The way you'd usually handle that AFAIK is to have the service ask the billing system for a "reservation" in its native units, likely with an attached TTL. Then, the service would translate those units to U.S. Dollars (or possibly Indian Rupees), taking your plan, discounts, vouchers, contracts, grandfathered pricing and all that into account. It would then "lock" the calculated amount of money, denying the reservation if total_spent + total_locked > spending_limit. After finishing the operation, the service would ask for actual billing and free the unused units.
            1. eddythompson80 · · focus · HN ↗
              Yes, thats one way to implement it. It still moves that global billing system into a ring 0 importance with 99.9999% availability and performance requirement, where before it wasn’t even in the picture. Not impossible, just suddenly every single AWS (or GCP or Azure etc) has a global single point of failure that gates all their performance and runtime actions.

              Not to mention having no way to network isolate such service as every single service in every single network boundary needs to be able to contact such service, and every service needs (more or less) same auth permissions on all user accounts it’s running.

              Again, solvable problems with enough code, but not simple problems by any means if you operate at a large scale.

      4. joshdavham · · focus · HN ↗
        > You can't easily estimate how much an operation is going to cost before you kick off that operation

        I recently had a debate with a colleague on this topic but concerning estimating the costs of AI agent work. For example, if you prompt an AI to refactor your codebase, the final cost can't be estimated perfectly, but I'm sure it can at least be estimated with some amount of precision! Like simply knowing that it will cost < $100 is actually great information even if the final work only ends up costing $5.

        I think there actually might be a business opportunity (or at least the opportunity to build something cool here) if anyone wants to work in the AI cost estimation space. It's not exactly an idea I want to pursue, but just thought I'd put it out there. AI cost estimation (even with wide confidence bands) would be very useful to a lot of people.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.