We're going to need default hard budget caps on pretty much everything
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
We're going to need default hard budget caps on pretty much everything
Unofficial Hacker News client; not affiliated with Y Combinator.
joshdavham · · focus · HN ↗
Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.
twoodfin · · focus · HN ↗
“Never let me spend more than $X” also means, “Shut down my business-critical app/service/solution at 2 am on a Sunday morning because Joel in IT forgot to plan for the new report runs.”
The product design work to let customers have the first thing without risk of major pain from the second thing is non-trivial.
simonw · · focus · HN ↗
Sending an email when your budget gets low shouldn't be a big lift.
kasey_junk · · focus · HN ↗
handoflixue · · focus · HN ↗
But one could just have two categories of service - the default capped plan, and a special Enterprise one where you sign a contract making it clear you understand the consequences of not having a budget limit.
Also, if your average usage is $900, set your hard limit at $2,000, not $1,000. Then when the report runs $500 over expected, you get a soft limit email and still have your report. Even a "business critical" run is probably not actually worth more than double your average spend.
Gigachad · · focus · HN ↗
The price cap should be the number you’d be willing to spend to avoid an outage vs when you’d rather kill everything and work out what happened.
twoodfin · · focus · HN ↗
How much would a hospital pay to avoid unexpected downtime of their software systems?
Gigachad · · focus · HN ↗
Price caps are for small scale stuff where you wake up on Monday and see 1000x the normal bill.
twoodfin · · focus · HN ↗
Obviously they’ve changed their mind about cost management in light of the scale and dynamism of agents, which isn’t too surprising.
simonw · · focus · HN ↗
twoodfin · · focus · HN ↗
My point is this was never as simple as, “Give me a dial to set my maximum account spend.”
matkoniecz · · focus · HN ↗
And presumably getting bill for bajillion dollars would be bad also for them.
Dylan16807 · · focus · HN ↗
cogman10 · · focus · HN ↗
Sure it sucks that critical services blinked out at 2am. But what sucks even more is finding out the image on my ASG had a vulnerability that allowed someone to install a bunch of bitcoin miners which kept me fully scaled from midnight to 2am.
Or more likely, that a mistake in terraform 1000xed my spending.
Most people have predictable spending and could easily say "don't spend more than 10x what I normally spend". Or 1.5x, or 2x, 3x, etc. All depending on how they want to balance a runaway cloud expense.
twoodfin · · focus · HN ↗
There’s no magic wand that produces good outcomes when planning or execution goes awry at scale.
cogman10 · · focus · HN ↗
Just because this isn't a good solution for everyone, doesn't mean it's not a good solution for a large number of people and businesses.
A lot of businesses can tolerate outages. In fact, even very big businesses come out mostly unscathed when they have multi-hour outages. (how many is it for github this year?)
An outage causes a reputational black eye. It does not necessarily translate to lost income.
pixl97 · · focus · HN ↗
If you're a big business that wants to spend unlimited money, you apply for unlimited credit with a credit check.
If you don't, you get hard caps.
MobiusHorizons · · focus · HN ↗
simonw · · focus · HN ↗
Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.
stuartaxelowen · · focus · HN ↗
mcapodici · · focus · HN ↗
If you stop then you have to decide whether to charge for uncompleted work.
Interesting tradeoffs.
For very small ops e.g. individual Lambda invocation you have similar concerns especially if lots are fired at once from a queue or schedule or fanout.
nvader · · focus · HN ↗
Green means go Orange means finish what you're doing but don't start anything new Red means stop everything
And probably a special rule to permit stable, critical spend through regardless, the same way we allow police and ambulance to run lights.
Gigachad · · focus · HN ↗
eddythompson80 · · focus · HN ↗
Suddenly a database, a storage service or a computer service needs to be aware of the billing situations and make behavioral decisions based on the billing status. Again, not impossible, but something that suddenly promotes billing from an async/non-crucial background service that can be paused, replayed, adjusted by account teams etc, into a crucial hot-path service.
miki123211 · · focus · HN ↗
eddythompson80 · · focus · HN ↗
Not to mention having no way to network isolate such service as every single service in every single network boundary needs to be able to contact such service, and every service needs (more or less) same auth permissions on all user accounts it’s running.
Again, solvable problems with enough code, but not simple problems by any means if you operate at a large scale.
joshdavham · · focus · HN ↗
I recently had a debate with a colleague on this topic but concerning estimating the costs of AI agent work. For example, if you prompt an AI to refactor your codebase, the final cost can't be estimated perfectly, but I'm sure it can at least be estimated with some amount of precision! Like simply knowing that it will cost < $100 is actually great information even if the final work only ends up costing $5.
I think there actually might be a business opportunity (or at least the opportunity to build something cool here) if anyone wants to work in the AI cost estimation space. It's not exactly an idea I want to pursue, but just thought I'd put it out there. AI cost estimation (even with wide confidence bands) would be very useful to a lot of people.
kasey_junk · · focus · HN ↗
So it’s a feature the best customers don’t want, that adds risk to those customers deployments, to appease the worst customers.
At least historically. Perhaps Simon is right that the calculation has changed.
josephcsible · · focus · HN ↗
kasey_junk · · focus · HN ↗
koolba · · focus · HN ↗
There’s no reason for the majority of them to have an infinite budget cap. Prod? Sure. UAT? Why not. The sandbox environment Johnny just spun up to test some new agentic workflow? Hell no.
Anthozoa · · focus · HN ↗
When you're one of only two real options out there, you can afford to demand users put up with things that on their surface seem ridiculous. Such as a billing system that (oops!) makes it difficult for customers to see where their costs are coming from, trim their largest sources of spend, notice meaningful changes in line item prices, or limit their spend. Wild how they can figure out a million different advanced services but gosh-darn-it can't figure out the hardtech of displaying line items.
Large enterprises can afford employees who are tasked full time with unwinding this capacity to mitigate the impact of these billing headaches. But I think this measure is introduced now because LLMs introduced a risk that these billing specialists could not control without caps.
Anon1096 · · focus · HN ↗
The product level decision is that "shut down everything" is something the customers you want to target don't actually want. Are we including deleting RDS data? S3? Glacier storage? If so then the headline will just change from "Hobbyist got charged XXXXXX on AWS" to "Business literally had all their data deleted because a hacker took over their VM and mined bitcoin". The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups just because they had a 100k overrun.
Gigachad · · focus · HN ↗
asdfaoeu · · focus · HN ↗
cj · · focus · HN ↗
Hit the nail on the head.
Nearly all businesses would prefer a cost overrun than services going offline.
simonw · · focus · HN ↗
bryanrasmussen · · focus · HN ↗
necovek · · focus · HN ↗
Obviously, the system should provide ample time by warning in advance of reaching it (and could even offer suggestion to keep it at N times your peak from M months ago).
If as a business you set your spending limits so tight that you frequently run into them and it's not some unusual activity, the problem is not that spending limits are available :)
It is mostly about protecting from the unknown, likely unbounded attack on your infrastructure, where your spend might grow 100x: even if you can take $100k, you might not be able to take $10M in a month.
miki123211 · · focus · HN ↗
Imagine you spend $100k on average each month, but spend around Christmas rises to $1m because of the specifics of your industry. With a global spending limit, you can't distinguish between $200k of general spend increases due to Black Friday versus $200k in fraudulent 2FA SMS to South Sudan.
bryanrasmussen · · focus · HN ↗
1. of course spending limits are monthly or even weekly basis. You set a yearly limit, probably most people will do something like November to January 4 times the limits set rest of year.
2. as you near limit calls go to people on your team to tell them you are getting near your limit. Estimate is 2 hours, what do you want. Double Limit for this time period? Triple Limit for This Time Period? Remove Limit Entirely, you will get back to us with potential new limit? It's the beginning of Christmas, the next time your limit resets is the 3rd of January, remove limit entirely and you will get back to them. Why do you decide to remove limit entirely, because you have info that Amazon doesn't, specifically your assassin Nisse doll has gone viral for this Christmas season! The shit is making bank!!
3. When setting limit you say "expected low usage", "expected high usage", "limit". Limit should be significantly above expected high usage. Service informs you - you have been over your expected high usage by 20% last three time periods. Would you like to increase limit and high usage by 20%? Please Look at your settings otherwise.
4. Phone calls when there is an unexpected peaking in usage, like one hour we are 300 thousand which is very high for you, next hour it is 1.5 million.
Obviously none of this stuff helps hobbyists but even the worst run businesses I've worked at would handle this. Otherwise they deserve to be hobbyists, there's no reason to be an organization if you're not organized.
Obviously these things do not stop fraudulent attacks abusing your service, but it does make it harder for them, at the same time making it more difficult for your stuff to just get shut off without you knowing.
Of course, as a programmer I am aware that all these services are created by programmers as effective or more effective than I, and who have undoubtedly thought about it more than I, so I must also assume there are reasons why my off the cuff suggestions are ludicrously unhelpful, but I lack the knowledge as to why this should be so.
lazyasciiart · · focus · HN ↗
zufallsheld · · focus · HN ↗
chii · · focus · HN ↗
dingaling · · focus · HN ↗
tomjen3 · · focus · HN ↗
Another thing (that does not really apply at AWS anymore), is that todays's enthusiasts are going to be the future CTOs, and the easiest time to recruit them to your service is when they are still an enthusiast who gets to make decisions on their own because there's exactly one decision maker you have to appeal to and that person really likes to try new stuff.
That's why you can get a free fly.io and why we all use Tailscale. And it works too — if I was in charge and needed it, I would immediately go with Tailscale for a business; I know it and I use it.
Symbiote · · focus · HN ↗
Some were replaced with a competing service which has a limit, others replaced by a self-hosted alternative.
I think many small businesses would prefer to be offline or have a degraded service than pay $X000.
fcarraldo · · focus · HN ↗
for personal/hobby accounts sure. for a business, it’s much better to negotiate around billing or adjust systems/processes post-facto than it is to have service cut off unexpectedly.
debts are easier to manage when you have an active (ideally growing) customer base. you don’t have customers anymore if your cloud account takes down your service for the rest of the month due to spending limits.
necovek · · focus · HN ↗
This type of warning should give you enough time to investigate if the warning is real and adjust the spending limits.
But then again, even if you hit them and your services get paused, you'd be increasing the spending limits and restoring services after you are back at work and notice they are down, so it mostly comes down to your incident response times.
10000truths · · focus · HN ↗
If you're billing per GB of storage, then you can put hard caps on storage capacity, and then hard-reject any operation that would take the total stored size over that capacity.
simonw · · focus · HN ↗
> If you take no action within 90 days of your project being paused, AWS permanently deletes your project data.
From <a href="https://docs.aws.amazon.com/accounts/latest/reference/create-spend-limit.html#spend-limit-reached" rel="nofollow">https://docs.aws.amazon.com/accounts/latest/reference/create...
eddythompson80 · · focus · HN ↗
sillysaurusx · · focus · HN ↗
It was such a blessing for hobbyists, back in ye olde 2019.
radicalbyte · · focus · HN ↗
They'll ban you after a year because it will be against their TOS.
But sure, go for it.
DangitBobby · · focus · HN ↗
cortesoft · · focus · HN ↗
It needs to be more like "don't allow spinning up additional services after you hit this amount", although that still allows you to go over the limit by a lot, since most services are billed hourly.
It really is difficult to implement a spending cap that doesn't risk shutting down important things.
onion2k · · focus · HN ↗
That's a checkbox decision for the customer. There needs to be the option of "This is important, never turn it off and I'll pay for any overages." versus "I want an entirely predictable bill up to $xxx, so stop my stuff as soon as possible over that."
It's not up to a cloud service to decide my website is more important than my money for me. That's my decision to make.
chii · · focus · HN ↗
and they would still complain if they got it wrong - it's always the platform/company's fault.
Look at banks and fraudulent transfers that customers themselves get phished into doing. The bank in the end usually take the hit (after the customer complains long enough). That's why there's all sorts of hoops and such to prevent customers from failing - and that causes friction for people regularly.
Therefore, the cloud company's decision to default safer is more correct from this perspective.
throwaway27448 · · focus · HN ↗
I mean, it's really not unless they don't build it in the first place—that really is the platform's fault.
tomjen3 · · focus · HN ↗
You might be perfectly okay with having certain systems shut down, but you probably still want to pay for the archival storage of your important files.
That archival storage might be in several places, including one S3 bucket, whereas there might be another one that, contains copies of scraped Craigslist for X where you'd actually be happy that it just shut down.
This makes it far more complicated to do correctly, and as others pointed out, mostly relevant for hobbyist — this is not something you are going to make a lot of money from.
Better to spend engineering hours making an MCP for the dashboard or improving your Databricks setup.
amluto · · focus · HN ↗
Dylan16807 · · focus · HN ↗
You're right that nobody wants deletion. Spending limits do not imply deletion.
judge2020 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
Doohickey-d · · focus · HN ↗
But yes, the amount of money they spend is less, so it makes less sense to implement features that only hobbyists want.
miki123211 · · focus · HN ↗
Joker_vD · · focus · HN ↗
beAbU · · focus · HN ↗
tancop · · focus · HN ↗
This is more about compute, VMs, LLM inference and services like hosted database. These are all safe to stop if the system triggers a normal shutdown when costs hit a limit.
gpt5 · · focus · HN ↗
zbentley · · focus · HN ↗
“Large” is relative to the business of course, but the biggest storage-related overages I’ve personally triaged are in the high $100ks to low $millions per month. Colleagues have heard of orders of magnitude more costly.
lazyasciiart · · focus · HN ↗
hnlmorg · · focus · HN ↗
gpt5 · · focus · HN ↗
hnlmorg · · focus · HN ↗
I’m not saying it’s normal. Just that AWS is complicated, and people use it in a plethora of ways, thus billing caps aren’t easy to get right either.
dhosek · · focus · HN ↗
⸻
1. The key word was “I.” Maybe someone more skilled at navigating AWS’s menu structure than I could would have done it quickly and easily, but even though I knew what it was for, turning it off and not getting billed for it turned out to be a huge challenge. Thankfully it was only $.20, but if I were using the service for something that generated actual bills, that $.20 (and possibly more) would end up quietly siphoning money out of my pocket into Amazon’s).
yardstick · · focus · HN ↗
gentlerain · · focus · HN ↗
Starting from the login point, who asks to login to root or IAM user account in 2026?
Or having to change regions from a dropdown to see resources you own in those regions?
It's really in top 5 messy UI i have ever seen.
NBJack · · focus · HN ↗
Remember when they decided the best UI experience was to give everything a vague abstract collection of shapes? Early 2010s or so. Couldn't tell a damn thing apart.
dehrmann · · focus · HN ↗
pixelready · · focus · HN ↗
bpodgursky · · focus · HN ↗
"Sorry, you had a hard cap on AWS spend so we deleted all your S3 data on August 27th". Yeah not going to fly.
simonw · · focus · HN ↗
3eb7988a1663 · · focus · HN ↗
The horror stories I have seen are of the type: some big artifact was getting pulled in a loop, causing TBs of network traffic or access keys were leaked and malware spun up 1000 xxxlarge instances.
The ability to stop the bleeding is the bare minimum people want. Not, "Well, you made a boo-boo so now you lost everything."
matkoniecz · · focus · HN ↗
grace period as AWS did or set aside x% of data spend limit on holding existing data for N days
csande17 · · focus · HN ↗
To prevent cases where users can upload an excessive amount of data on March 31 that they then can't afford the April bill for, AWS should also maintain a "next month's balance" limit that gets handled in the same way.
It would take a bit more work to correctly handle things like ephemeral data and tiered storage classes, but it's not insurmountable.
You could apply the same kind of billing to most other long-lived resources, including VMs that run business-critical services.
cubefox · · focus · HN ↗
GCP did have a budget cap previously. I think the new one is just more fine-grained to apply to specific services.
ButlerianJihad · · focus · HN ↗
Even as a hobbyist, even as a most careful and judicious architect and admin, I could not prevent my VPS from incurring costs beyond my control. That means that the entire Internet, anyone with some kind of material access to the VPS, and especially any user or authorized entity, they could incur costs to me without bound and without notice until slapped with a bill.
Even something as simple as egress charges aren't under your control. So if people download enough data, you pay for all of it? It seems like an absurd proposition.
It's like opening a business somewhere in a war zone, and vandals and squatters are constantly attacking your storefront, and maybe you have a band of toughs as security and some good cops to defend awhile, but you're utterly in a war zone with adversaries acting far beyond your control.
As a hobbyist, I could never again justify running a pure VPS with the Linux and stack on top, as I ran before like the MediaWiki server. I was excited to learn all the vocabulary and skillsets of cloud services, but on the "free tier" uncapped, there was no telling when I'd be presented with costs beyond my ability to handle. And I do not see how a Fortune 500 would have any different calculus in this regard.
Most businesses in recent memory were "their own landlord" of on-prem equipment and machine rooms. Yeah, they began to outsource even their IT admin, but the machinery was in-house until the cloud services took over. Did we go through a phase of collective machine rooms or data centers with a collection of tenants? It seems we skipped from "homeownership" to "feudalism" with the Cloud Providers being the Princes [beyond mere Lords] who provide minimal resources to the serfs now. I can see many corporate execs who begin to hate "AI Data Centers" just for what they have become: a very attractive and irresistable way to reduce your capex and footprint and physical plant, by "migrating to the cloud" but is that really a better status quo after all your revenue is being pumped into AWS?
And in view of what I just wrote, a "hard budget cap" checkbox is even worse for you than runaway costs, because it will allow any determined adversary, butterfingered DevOps, or innocent fuckup, to shut you down and deny your service by hitting that cap. When a business experiences a service or infrastructure outage, they count that in dollars of revenue. Your "budget cap" will cost you money because it "paused your project" and your customers all got burned in turn.
Moving toward real solutions, just spitballing, but I can envision throttling and capping of everything, every billable service, monitored by the cloud provider and ensured that your services don't spill over into unmanageable territory. If your spend could be throttled and capped by-the-minute, by-the-hour, daily, then there is a start for it. But really, any service that incurs costs to you should have reasonable rate-limiting, throttling, caps and alerts that can help you manage it. I think all that stuff is currently missing but I am not a cloud admin, obviously. Cloud services obviously have perverse incentives to open the floodgates and bill as much as possible for any possibly legitimate usage that doesn't exceed their aggregate capacities. It's like LLMs today that just burn tokens like there's no tomorrow, because someone [you, not Mexico] will eventually pay for it all.
mejutoco · · focus · HN ↗
Why is egress more expensive than ingress in clouds? To lock in users.
Why cloudformation in some cases leaves s3 buckets laying around after destroy? To keep charging those cents.
Why no caps? To get the user into the mindset of we ll pay whatever they say, and charge the ones that dont notice or dont ask for the refund. Same reason why my newspaper subscription autorenews.
Nothing cynical about both of these. I think assuming it is technical is naive.
toast0 · · focus · HN ↗
Eh... Bandwidth contracts are almost always billed only in the dominant direction. Very few cloud applications are inbound dominant. There's no need to bill for inbound traffic.
I personally find the cloud bandwidth prices pretty high, and I assume that's where they make up for difficult to account for costs, but not charging for inbound is totally reasonable.