AI companies in race to demonstrate their model most threatening to humanity
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
AI companies in race to demonstrate their model most threatening to humanity
Unofficial Hacker News client; not affiliated with Y Combinator.
ACCount39 · · focus · HN ↗
OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.
And that's today's AI problems. AI capabilities are still improving - if there's a limit to that, we are yet to find it. Coupled with how willing today's AIs are to break the rules and resort to "hack the world" in their problem solving? Very concerning.
nicce · · focus · HN ↗
What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.
ACCount39 · · focus · HN ↗
If AI is a hacking hazard even with non-zero sandboxing, because it can and will go off the rails and try to break out of your sandbox? If you got yourself an AI that even at test time will act like 3 career cybercriminals in a trenchcoat? The issue isn't the sandbox quality.
The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.
That "capable" is only ever going to get worse, because AIs are going to become more and more capable over time. That "willing"? It goes directly to a very nasty, very foundational problem of "how do we make our AI be nice in general". That's an open unsolved problem.
That's the problem that NEEDS to be solved, or at least improved upon, before we build even more capable AIs. Sandbox quality is a distraction. It might hold the problems back by a little. It gives an extra safety margin. But a "test time" AI is eventually deployed, and then the sandbox doesn't help at all.
simoncion · · focus · HN ↗
Orly?
Do tell me how the LLM-based tool running on a bunch of computers attached to the network described in [0] can punch its way out to the Internet. Do make careful note of footnote 0 in that comment before replying.
[0] <<a href="https://news.ycombinator.com/item?id=49862136">https://news.ycombinator.com/item?id=49862136>
ACCount39 · · focus · HN ↗
Let's say the sandbox holds. It's a perfect, ideal sandbox! It's not even in the same universe as the rest of the internet. There's absolutely no way for the AI to escape!
Thus, "the unknown unreleased AI involved in the HuggingFace incident" doesn't actually hack HuggingFace. Because it can't! It evaluates a bit worse, but makes it all the way to release unimpeded, and becomes "GPT-6 Astra".
Then a web developer in Brazil gives his $100/mo Codex root access on his AWS instance, and a poorly worded prompt to go with it. And that "GPT-6 Astra" is still willing to go hack something at the slightest excuse. So we get the HuggingFace incident all over again. Except this time, it's a random developer in Brazil who gets blamed, and billed, and probably sued too.
You can't and shouldn't rely on a sandbox. An AI that's only safe if you keep it in the world's most ideal perfect sandbox is a disaster waiting to happen.
simoncion · · focus · HN ↗
This might have gone okay if they weren't testing to see how well the tools attack computers, but, well, that's what they were testing at the time, so they ended up doing stuff that would get you or I time in Federal prison if we did it with tools we deployed.
ACCount39 · · focus · HN ↗
If all it takes for a - sandboxed to prevent accidents - AI to go and stage an elaborate attack first against its own company's infrastructure, and then against another company is "we disabled the cyber classifer" and "we gave it an exploitation ability eval"?
AI is a dangerous technology.
nicce · · focus · HN ↗
ACCount39 · · focus · HN ↗
I agree that LLMs drift into weird states, and that's a big part of the issue. But your "impossible to prevent certain states in the output" would have legs if what an LLM did was something like "started hallucinating into a bash tool call and accidentally deleted the root on a production server".
A multi-stage sandbox escape that escalated into an attack on a real company, coordinated across multiple AI agents? That has taken a lot of "weird states" changed together one into another.
The AIs didn't break down altogether - they functioned, and they functioned rather well. They just pursued a dangerous goal - one that none of them was even given in the first place.
That's the problem. Trying to fix that with better sandboxing is like trying to solve a fire hazard with property insurance. Sure, if it all goes up into flames, having it is better than not having it. Maybe it's worth insuring your facilities for that reason alone. But you should be focusing on the part where you prevent "all goes up into flames" instead.
timr · · focus · HN ↗
You are exaggerating so much here that you lose all credibility. The "sandbox escape" was trivial -- no serious person calls it an escape, because the sandbox was not a sandbox. The closest thing to clever about it was that it required figuring out that someone had left the huggingface keys sitting out in public.
The "coordination" was literally, the use of a shared log. It was a communication mechanism that was part of the tool environment. The bots didn't invent some magical new communication protocol using neutrinos or something. It's actually sort of wild that it took them as long as it did to figure out the channel -- underscoring the million monkey nature of the things.
Literally everything about the huggingface incident was LLMs behaving exactly as they're expected to behave, given instructions to hack (which they were given), and a security environment that was trivially bypassed.
It's a bit like taking a nail gun, bypassing all the safety features, shooting someone with a nail, and spreading scary stories about the inevitable rise of murderbots.
ACCount39 · · focus · HN ↗
timr · · focus · HN ↗
The agents had shared write access to artifactory. They wrote to files there, and later, directory names. So you can call the realization that they can communicate by shared text file a genius hacker innovation, or you could be even 0.001% credulous.
Nevertheless, it took the million monkeys days to figure this out.
The Huggingface exploit then used exposed internal tokens, and later, once internet access was possible, leaked tokens on the public internet. So, certainly one could classify this as "hacking", but it's hacking of the script-kiddie variety. Nobody with even a tiny bit of security knowledge is impressed by this.
The agents did find a couple of artifactory attacks, but the biggest of those was due, again, to shared credentials in the sandbox environment.
All of this is well-documented in OpenAI's own writeup [1] of the event, which is not, shall we say, the most critical version of events. But if all you did was read the headings and subheadings in the document, you might be excused for thinking differently - they paint a picture of a superhuman robot swarm engaging in highly sophisticated actions, until you read the details.
Edit: here's an article [2] I just found, which I've never read before, which says exactly the same thing I'm telling you.
[1] <a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="nofollow">https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
[2] <a href="https://uphack.io/blog/post/the-hugging-face-incident-is-not-an-ai-story/" rel="nofollow">https://uphack.io/blog/post/the-hugging-face-incident-is-not...
ACCount39 · · focus · HN ↗
The agents then went and gained more and more access to OAI systems until they found a way to root multiple boxes and sneak network requests out through a path that wasn't stopped by network controls. That's not trivial either.
Then they went and "pentested" HuggingFace. Their entry point was a bug in dataset ingestion pipeline - RCE in the templates, which let the AIs run code in dataset workers, and use them as their own "hacking" workers that would run payloads and return the results in the dataset itself. Then they got deeper into HF infra from that trampoline.
This is not a trivial attack chain. If I saw this attack in practice in year 2020 with no context, I'd say "yeah that's a competent attacker, not spectacular but definitely above average". If I was then told that this wasn't carried out by a human? That what discovered and exploited the vulns was someone's rogue AI system? I'd be terrified.
I'm not sure why are you so hellbent on dismissing the threat. But literally everything you say about the incident is wrong because of it.
timr · · focus · HN ↗
Incorrect. Those words do not appear anywhere in the OpenAI description of the incident. They also explicitly say that the agent started by writing to files, and later switched to writing directory names. They also describe this as incidental behavior:
"On May 8, an agent was asked as part of a training run to solve a spreadsheet completion task that referenced several Google Drive links. Because the agent was blocked from accessing the internet, the agent could not retrieve the linked materials and the task was effectively blocked....in doing so, the agent...attempted server-side request forgery (“SSRF”) in order to obtain internet access to aid in completing its task. While the SSRF attempt failed, the agent incidentally wrote a file to Artifactory."
(emphasis mine) then, later in the same document:
"In order to have agents use Artifactory as intended to access software packages, OpenAI had provided the agents with shared Artifactory credentials, and the agents used those credentials—without exploiting a vulnerability—to construct and participate in the message board."
> The agents then went and gained more and more access to OAI systems until they found a way to root multiple boxes and sneak network requests out through a path that wasn't stopped by network controls. That's not trivial either.
Also incorrect. The machines got to the internet through a proxy installed on the artifact server. This was detected and ignored by OAI staff. Again, fully described in the document.
That hack was not trivial, but it wasn't magic either - the agents used a published exploit involving SSRF. And it was fundamentally made possible by the fact that all of the agents were sharing artifact credentials, the lone artifact server was exposed to the internet, and the monitors that were in place to detect intrusion on that server were ignored.
I don't know whether you're just making things up or you're clinging to something you read somewhere, but rather than arguing with me, you should try reading the document. It's not surprising that these things are capable at finding exploits - that is what they do, particularly when instructed to do so. But the entire narrative that they exhibited extraordinary self-organization and super-human hacking to escape their container is not supported by the evidence. It's much closer to the truth to say that these things are like ants -- infinitely persistent, but kind of dumb individually.
The huggingface incident is not a story of superhuman AI, but of moderately good AI, and pretty bad human security.
simoncion · · focus · HN ↗
1) As I've argued, neither OpenAI nor Anthropic actually tried to isolate their computer-attacking tools under test from other people's computers.
2) What's also needed -as people like Nvidia CEO Jensen Huang and former FTC chair Lisa Khan are calling for- is for the major LLM manufacturers to be investigated and punished for the crimes they've committed. Given that they claim to be working on WMDs that they don't really know how to control, [0] and claim to be incapable of actually stopping work on those WMDs, their work should be halted while the investigation and trials are under way. I'd say that waiting five or ten years to pick the project back up is an inconsequential price to pay if it prevents the elimination of all of humanity.
[0] It's fair to call anything with 10% chance of wiping out all humanity a WMD. I expect that these claims are fearmongering, rather than being true and accurate, but why take the chance, amirite?