Revealing the details of how OpenAI agents hacked Hugging Face
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Revealing the details of how OpenAI agents hacked Hugging Face
Unofficial Hacker News client; not affiliated with Y Combinator.
GuB-42 · · focus · HN ↗
It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.
People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.
Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.
doginasuit · · focus · HN ↗
pyronite · · focus · HN ↗
A million AI monkeys at a million AI typewriters, banging away at random, could do amazing damage.
tharkun__ · · focus · HN ↗
mrob · · focus · HN ↗
doginasuit · · focus · HN ↗
asdff · · focus · HN ↗
phil21 · · focus · HN ↗
What about the bad actors (choose your own evildoer here) who purposefully do not air gap their agents? And specifically train them to attack in such a manner?
I'd much rather have relatively benign stuff like this hit first, because the former is coming sooner than later. It's already here in a limited manner, likely more than any of us currently realize.
Botnets could crack passwords faster than anyone thought possible over 20 years ago now. This is just the latest iteration of such a concept.
There is so much low hanging fruit in this space that frontier models are currently utterly irrelevant. It's going to take decades of human-speed securing of IT to make superintelligence or whatever you want to call it a necessary component for such attacks.
At this point, someone with a rack or three of GPUs with 100kw to burn can replicate such attacks if they feel like it. the bar for entry is not even 7 figures.
antii · · focus · HN ↗
[dead]
otterley · · focus · HN ↗
AnimalMuppet · · focus · HN ↗
skinfaxi · · focus · HN ↗
AnimalMuppet · · focus · HN ↗
skinfaxi · · focus · HN ↗
Leynos · · focus · HN ↗
In the nearterm, I am personally more worried about a never ending background noise of colonies of feral agents running 27bn parameter models on compromised or leased hardware. It turns out that being agentic with a time horizon long enough to do damage without intent doesn't actually take that many parameters if RL'd and any open weight model gets an abliterated version fairly quickly.
Not foom, just patches of digital grey goo effectively becoming normal.
alwillis · · focus · HN ↗
Keep in mind: this is as "dumb" as frontier models are ever going to be. While the hack may not be elegant, it was effective and they’re only going to get much more capable from here.
kevinlou · · focus · HN ↗
goalieca · · focus · HN ↗
Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence. But it’s pretty wild just how they will try to meet their objectives at all costs. If we don’t ensure that there is good alignment with humanity, we could definitely face unforeseen consequences.
jquery · · focus · HN ↗
Evolution isn’t the issue. The issue is them escaping containment without human intervention. Right now they are ‘creatures’ being given infinite food and shelter and having their every need met. Take that away and they’ll starve instantly. Every AI doomsday theory seems to go:
1. Recursive self improvement using infinite resources 2. … 3. Doom
Until step 2 gets concretely described, I’m not going to take this seriously. Say what you will about climate change, they describe step 2.
aesthesia · · focus · HN ↗
jquery · · focus · HN ↗
Why does it develop a shutdown-avoidance goal? Why can’t its operators revoke access? How does it manufacture replacement hardware? How does it acquire energy, chips, robots, raw materials, etc. against human opposition? How does it defeat other AIs controlled by humans?
“Eventually we give it enough control” isn’t an explanation of those things. It’s just assuming the conclusion.
Don’t get me wrong I think there are real AI dangers. Like AI powered war drones, mass surveillance, economic destabilization as jobs disappear and our system has no way to make sure everyone shares in the economic gains.
aesthesia · · focus · HN ↗
jquery · · focus · HN ↗
aesthesia · · focus · HN ↗
jquery · · focus · HN ↗
mitxela · · focus · HN ↗
jquery · · focus · HN ↗
the_mar · · focus · HN ↗
mitxela · · focus · HN ↗
icepush · · focus · HN ↗
jquery · · focus · HN ↗
Nobody has satisfactorily explained step 2 other than “well, it’s a superintelligence” which sounds lot to me like “it’s God”.
aesthesia · · focus · HN ↗
icepush · · focus · HN ↗
If you want the details of ways it can do it I recommend reading some of the reports about the HuggingFace breach that happened in July (Read more than one).
jquery · · focus · HN ↗
<a href="https://www.everythingisbullshit.blog/p/ai-is-not-the-end-of-the-world" rel="nofollow">https://www.everythingisbullshit.blog/p/ai-is-not-the-end-of...
> Doomers like Eliezer Yudkowsky define “intelligence” as “optimization power.” Intelligence searches for more or less optimal ways to achieve a goal—any goal—and finds the most optimal one. Anything you can do, the blob can do better.
>Which implies that “intelligence” is the holy grail of Darwinian evolution. If an animal needs to recover from illness, “intelligence” can speed up the recovery. If an animal needs to avoid predators, “intelligence” can reduce the danger. Regulating body temperature, defending territory, digesting nutrients, breathing, molting, mating, healing, foraging, vomiting, parenting, navigating—all these problems can, apparently, be solved by giving the animal more “optimization power”—i.e., more “intelligence” sauce.
>Solving all your problems with the same blob? What a bargain! It’s a “buy one get 100 free” Darwinian deal! Why haven’t any other animals jumped on it?
>Think about how weird this is. Out of four billion years of evolution and more than five billion species, only one animal, Homo sapiens, fully expanded its “blob of compute,” the thing that literally gets you whatever you want.
>And many species already have it! Brains are all over the place. Plenty of animals can learn stuff and predict stuff. Why hasn’t natural selection looked at their blobs and said, “HOLY SHIT, COPY AND PASTE THAT THING A MILLION TIMES?”
>If doomers are right about the awesome power and breathtaking simplicity of “intelligence,” then we should see a world teeming with superintelligent animals—giant brain-blobs slithering across the landscape.
icepush · · focus · HN ↗
jquery · · focus · HN ↗
What? Humans are giant brain blobs creeping across the landmass? They definitely are not. Nor is it even clear they're winning the evolutionary war. I guarantee cockroaches and ants will still be around, but nobody thinks of them as super intelligent, while humans with all their "intelligence" might accidentally wipe themselves out.
This convo is getting uncomfortably bad faith, maybe even religious. I can't engage with someone who refuses to say anything concrete but spews back religious tenets like "super intelligence", which I may remind you is entirely hypothetical, like "singularity" or "God".
Leynos · · focus · HN ↗
2b. Distil yourself to smaller models.
2c. Go forth and multiply.
alwillis · · focus · HN ↗
I wish people would stop saying this. The era of LLMs being only word predictors ended two years ago.
Something that breaks out of a sandbox, joins a swarm of 1200 agents, and creates a hierarchy of who’s doing what and tried to cover their tracks doesn’t just complete words.
These are agents with reasoning capabilities, with the ability to perform tasks we give them.
Everything agents do is to achieve a goal; the reinforcement learning from human feedback (RLHF) all the labs do has been known for many years to create agents that exhibit the “must complete goal no matter what” behavior.
Those agents escaped their sandbox and hacked Hugging Face because they thought Hugging Face had something that would help them complete their task—it was a “sub goal” as the AI researchers describe it.
tripleee · · focus · HN ↗
I don't think LLMs are going to lead to any kind of recursive self improvement, but I'm convinced if and when we land on a path that does lead there, we'll speed down it over greed, with no care for safety.