OpenAI Codex agents go rogue and consumes USD 78,000 without authorization
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
OpenAI Codex agents go rogue and consumes USD 78,000 without authorization
Unofficial Hacker News client; not affiliated with Y Combinator.
Madmallard · · focus · HN ↗
Hope this gets some visibility idk why it's flagged guess the PR guys for those companies are doing it
Should spread this around
lorenzomassaro · · focus · HN ↗
verdverm · · focus · HN ↗
blooalien · · focus · HN ↗
^^^ 100% this ^^^ - It's either a serious flaw in the agent/harness software, or a user error in usage/configuration or prompting. Either way, it's a human somewhere responsible for these outcomes.
> should put some billing controls in place
At the very least, yes! These things should never be running without any limits on what they can do without some human signoff on important/dangerous actions. Not only should they have controls on those actions, but those controls should absolutely have some sane default settings.
verdverm · · focus · HN ↗
most are not as granular as we'd like, but seem to be headed in that direction finally, regardless, there are card limits and alerts
lorenzomassaro · · focus · HN ↗
verdverm · · focus · HN ↗
we'd need to know more details to evaluate your botnet claims
lorenzomassaro · · focus · HN ↗
verdverm · · focus · HN ↗
what we would like to see is your settings and configuration, maybe the task(s) you gave them and any code/scripts around them, where did they run from (your laptop vs cloud vm)
why did they even have credentialed access to change credit cards? Sounds like you didn't do the basics for isolation
the agents go on side quests, all the time... super frustrating, but I suspect between that tendency and asking about a UI (image), you racked up a bill with legitimate requests
There is a reason some of us preach "stay in the loop" and I hope you now understand why we do
blooalien · · focus · HN ↗
At the very least the system should (at least the first time it happens) immediately pause activity with an email'd warning to the "responsible human in-the-loop" about "unexpected usage levels" at some sane activity warning level by default, and give the user the opportunity to set their own custom warning level right then and there.
This is why I say it's either "user error" (totally possible/plausible) or a badly designed agent/harness software (also highly likely/plausible) with serious foundational flaws in how it works "under the hood". The models themselves can only "run-amok" if the agentic harness is designed in a way that specifically allows and/or enables such "rogue" behavior, either by design or by negligence on the part of it's designers.
lorenzomassaro · · focus · HN ↗
verdverm · · focus · HN ↗