We're going to need default hard budget caps on pretty much everything
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
We're going to need default hard budget caps on pretty much everything
Unofficial Hacker News client; not affiliated with Y Combinator.
motionlessveloc · · focus · HN ↗
It was, unfortunately, a nightmare. There were tons of tickets and even threats of lawsuits from customers whose service got cut off hard at the worst possible time due to organic growth/going viral/big event/nobody knew about the limit/etc. Not only did they lose all the leads and revenue they would have gotten from that bump, but they also pissed off their own existing users who suddenly couldn't use the service either.
Generally speaking, it's much better to use alerts instead of hard limit. Even in the worst case (hackers pwn your credentials and mine Bitcoin or whatever) the rest of your business is unaffected and you can negotiate with the billing department at comparative leisure.
This is all assuming you have humans operating the service. If you're letting AI agents yolo infra in prod, you have a whole series of new problems.
clickety_clack · · focus · HN ↗
reticulates · · focus · HN ↗
As a customer, the big number is scary and causes panic but for the provider… customers constantly fail to pay bills, providers are constantly writing off bills because it just isn’t worth the cost to chase, if a customer says “hey that usage was a mistake” it’s usually worth it to write it off to save the relationship. If you write off a big bill that wouldn’t have been paid anyway, the customer will perceive you as wonderful and benevolent and be loyal for life when they are ready to spend their money.
With tokens though the actual cost being incurred is much, much higher. If your service is just a wrapper around tokens, and a customer incurs $10k of usage that you paid OpenAI $5k for, it becomes much more difficult to write off.
Google Cloud is one of the few services that actually pursues unpaid bills even on their high margin services.
preommr · · focus · HN ↗
I recently loaded up on prepaid api credits for gemini and it somehow triggered some billing shenanigans in my linked accounts where it said I had a negative balance (from the credits), and they were going to discontinue my services. I had to reset some settings to sort it out, mainly using their chat ai and mine (because theirs gave me right status info, but wrong conclusions).
It's pretty messy across like aistudio.google.com, and their typical console, and google workspace business account. I'd be so fucked if they froze my account, I'd rather just pay openrouter to access credits in the future.
mschuster91 · · focus · HN ↗
Google is the only cloud platform I'll not just never use, but personally discourage anyone from using it.
Simply because there have been way too many horror stories on here about people who had gotten their personal gmail accounts frozen for whatever BS reason - and absolutely zero recourse. With any other large service you can always get ahold of a human, with anything tied to Google it's impossible and even raising a major stink on HN or "legacy media" often does not help.
knollimar · · focus · HN ↗
sandworm101 · · focus · HN ↗
trollbridge · · focus · HN ↗
walrus01 · · focus · HN ↗
Heck, have it do the equivalent of send people a DocuSign equivalent PDF to sign acknowledging the risk before enabling it. Would it still stop pissed off people? Probably not. Would it help with the risk of lawsuits, very possibly.
spoonyvoid7 · · focus · HN ↗
I'm curious. How likely is the billing department to waive off a huge bill as bad debt because an inexperienced builder misconfigured their infra or was hacked?
reticulates · · focus · HN ↗
aenis · · focus · HN ↗
And its of course not just people yolo'ing with AI. People were quite capable of causing such outages themselves just fine. Distributed, serverless systems are hard.
traceroute66 · · focus · HN ↗
Anyone who says its 'much better' to use alerts instead of hard caps needs to Google 'alert fatigue'.
Alerts are soft. You ignore them or miss them, nothing happens except you spending $$$$$$$$$$ more.
Great if you're the cloud provider raking in the cash, but a poor way to run your infrastructure.
Hard caps force you to implement correctly (to control costs in the first place) and have correct monitoring in place (to keep a healthy cap buffer).
So it means you can't just vibecode some slop and blindly devops it via Github CI/CD. You actually need to think and reason about your infrastructure.
bluGill · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
dvfjsdhgfv · · focus · HN ↗
You can't extrapolate your own experience and generalize like that. There is a huge number of people who explicitly want hard caps, period. AWS giving in and finally offering this option after 20 years of people begging them to do that is a sign of that.
KronisLV · · focus · HN ↗
Probably for companies and opportunistic individuals with a high risk tolerance. As for me, just give me hard budget caps. It shouldn't be a choice between supporting one of the approaches, when you could let each client pick what they want.
Want an alert? Configure one.
Want a hard limit? Configure one.
(just make the configuration easy and visible)
edoceo · · focus · HN ↗
1000 years ago, before CI/CD took over folk would have launch/go-live events, where teams would do a good checklist. There were pre-launch meetings to review, which included load checking.
486sx33 · · focus · HN ↗
[dead]
figassis · · focus · HN ↗
hypfer · · focus · HN ↗
A better approach to that scenario might be to split into base load and peak load infra, similar to how we do it with the power grid.
Base load could be an actual metal server that you pay x amount of money for and is fully yours, with peak load being handled by autoscaling dynamic stuff.
Both things being on a fixed budget, if that budget would run out, the service would not degrade as in "grind to a halt" but as in "slows down", which is probably not a downtime as part of a communicated SLA.
My point being that cloud and on-demand is useful, but a hybrid approach might in many cases make more sense. Though of course YMMV. I do not know what the requirements of your product are.
kccqzy · · focus · HN ↗
hypfer · · focus · HN ↗
kccqzy · · focus · HN ↗
hypfer · · focus · HN ↗
afiori · · focus · HN ↗
ndsipa_pomu · · focus · HN ↗
This is not rocket salad to implement correctly.
pixl97 · · focus · HN ↗
Shorel · · focus · HN ↗
It should be the choice of the customer.
Imposing one option: only hard budget cap or only alerts is the wrong design decision.
pixl97 · · focus · HN ↗
mcintyre1994 · · focus · HN ↗
pixl97 · · focus · HN ↗
watwut · · focus · HN ↗
It works best when coupled with spam of other irrelevant alerts, to maximize the chance of higher earnings.
mybrowsercache · · focus · HN ↗
collingreen · · focus · HN ↗
oblio · · focus · HN ↗
bluGill · · focus · HN ↗
Kim_Bruning · · focus · HN ↗
layer8 · · focus · HN ↗
kurthr · · focus · HN ↗
seviu · · focus · HN ↗
The caps were very generous but I wasn’t going to spend a minute worrying about what if a malicious actor took over
MaruthiS · · focus · HN ↗
[dead]
hn_throwaway_99 · · focus · HN ↗
I got bit by some undocumented (or at least very poorly documented) hard limits in AWS when our app went viral. It was extremely difficult to just find out what these hard limits actually were. And most importantly, we never explicitly set them, they were just hidden defaults in AWS.
That's very different from having an easy to use, visual dashboard of where and what all your hard limits actually are, and make it it extremely easy to turn them off or on at a moments notice.
hashstring · · focus · HN ↗
soltanov · · focus · HN ↗