‹ BackHN Continuity

Thread

AI companies in race to demonstrate their model most threatening to humanity

441 points · 399 comments · ljewalsh

  1. ACCount39 · · focus · HN ↗
    Because AI genuinely is an extremely powerful and extremely dangerous technology, and the "best practices" of dealing with that are still being written.

    OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

    And that's today's AI problems. AI capabilities are still improving - if there's a limit to that, we are yet to find it. Coupled with how willing today's AIs are to break the rules and resort to "hack the world" in their problem solving? Very concerning.

    1. ForHackernews · · focus · HN ↗
      You mean the sandbox that wasn't airgapped from the public internet? The one that wasn't even a separate VM? The one that was barely a chroot? That sandbox?
    2. password54321 · · focus · HN ↗
      If it is "extremely powerful" then my concern is not AI going rogue but a concentration of power of the few companies that decide how it is used and distributed not websites getting hacked.
    3. hnedeotes · · focus · HN ↗
      So powerful they can't even do math properly without a staff worth millions writing all the clutches so that they can, wooooowooooooo
    4. nicce · · focus · HN ↗
      > OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

      What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.

      1. simianwords · · focus · HN ↗
        The model found and exploited and chained together previously unknown vulnerabilities.

        How were the sandboxes poor?

        1. dns_snek · · focus · HN ↗
          Agents didn't have real network isolation. They were indirectly connected to the internet via a jump host running insecure software which was never designed or hardened to provide any kind of isolation.
          1. simianwords · · focus · HN ↗
            But this level of isolation is what happens normally. At least in my university and another company I worked at. It wasn’t running insecure software, as far as anyone knew, it was secure
            1. dns_snek · · focus · HN ↗
              There's levels of isolation, a padlock is not equivalent to a bank vault. If you claim to be building a possibly world-ending AI then you don't get to use a padlock and call it a day.
              1. simianwords · · focus · HN ↗
                And they are not calling it a day and they have done most things possible to communicate the fact that they are building something dangerous.
                1. dns_snek · · focus · HN ↗
                  They "called it a day" when they decided to deploy an agent which they believe to be dangerous inside this poorly secured environment.

                  There was no real isolation because a part of the system that doesn't provide any isolation guarantees was bridged to the internet. The next version of Artifactory, which you definitely wouldn't audit before you rolled it out, could simply add a public API that sends requests out to the internet.

                  Such an innocent upstream change would be equally catastrophic for your security model, which should demonstrate why it's negligent to rely on undefined behavior to enforce your security policies.

                  This should frankly be obvious any operator entrusted with running dangerous and possibly malicious code. Even if you don't know what you're doing, any LLM would tell you that this is a really bad idea if you simply asked. Don't rely on spacebar heating [1] to keep humanity alive.

                  [1] <a href="https:&#x2F;&#x2F;xkcd.com&#x2F;1172&#x2F;" rel="nofollow">https:&#x2F;&#x2F;xkcd.com&#x2F;1172&#x2F;

                2. uneoneuno · · focus · HN ↗
                  What about the waves of employees quitting both OAI and anthropic stating as their reason that they see their work as a direct threat to the future of humanity, and that their management isn&#x27;t taking it seriously enough. Is that a PR stunt too?
        2. vmg12 · · focus · HN ↗
          The model is really good at hacking, we all know this. This is why you don&#x27;t just expose random pieces of software to it without that software being hardened.

          It&#x27;s not like the model managed to exploit firecracker itself (no model has been capable of this), the model exploited artifactory.

          Artifactory is not some hardened piece of software that is meant to block users from accessing the internet through it.

          1. solenoid0937 · · focus · HN ↗
            &gt; The model is really good at hacking, we all know this

            No, we didn&#x27;t know that and this is how you find out they&#x27;re very good at hacking

            HN&#x27;s memory is so fickle. Just a few months ago almost no one here believed Mythos could actually be as good at hacking as the company claimed. This was a novel concept when the companies experienced these breakouts.

            1. vmg12 · · focus · HN ↗
              Models have been good at finding exploits for half a year now, this is not how we found out LLMs were good at hacking, you are rewriting history.

              We knew models much weaker than mythos were good at hacking the problem they had was that when finding exploits they had too many false positives.

              Either way, putting artifactory on the sandbox security boundary is obscene negligence. There is no reason to believe artifactory is secure.

              1. solenoid0937 · · focus · HN ↗
                If you listen to the OpenAI Black Hat talk it is very obvious they were surprised at the level of capability on display and felt it was novel.

                But I guess OpenAI&#x27;s security researchers acting surprised is part of some grand conspiracy to manage PR?

                1. vmg12 · · focus · HN ↗
                  It may have been novel for OpenAI but we already had Mythos at this point and this talk by Nicholas Carlini.

                  <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=1sd26pWhfmg" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=1sd26pWhfmg

                  We already knew LLMs were capable of finding exploits like this.

                  1. solenoid0937 · · focus · HN ↗
                    The Mythos issue is different though. It did not use any novel exploits. Irregular, the vendor they were using, misconfigured the environment to allow internet. Every other AI company also uses Irregular and that&#x27;s why we saw so many articles come out at once.

                    The OpenAI HF incident is separate from this. It involved actual zero days, teamwork and message passing, and sophisticated chains of exploits.

                    1. nicce · · focus · HN ↗
                      &gt; The OpenAI HF incident is separate from this. It involved actual zero days, teamwork and message passing, and sophisticated chains of exploits.

                      What exactly? They seem to be trivial SSR&#x2F;path traversal and input validation issues. Including misconfiguration. Nothing novel.

                    2. vmg12 · · focus · HN ↗
                      I&#x27;m just not impressed by AI finding zero days in unhardened software like artifactory. I&#x27;m completely unsurprised it had zero days and I fully expect AI to find any that exist. AI will literally try all combinations of inputs to achieve its goals. Any exploits that exist will be found.

                      The question anyone versed in security would ask was why anyone thought artifactory was an acceptable security boundary. I would never assume artifactory was secure. It&#x27;s like someone telling me there is a 0 day in a wordpress extension. So what?

                      Also I don&#x27;t think Qemu is secure either because it&#x27;s millions of lines of C and C++.

                      Firecracker I can trust to be secure because it&#x27;s 70k lines of human audited Rust. I know there are multiple people that have a complete understanding of the firecracker codebase.

      2. dns_snek · · focus · HN ↗
        &gt; I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.

        They&#x27;ll never let that happen because it would destroy their credibility.

        It&#x27;s like they&#x27;re telling the world about this dangerous, possibly world-ending pathogen that they&#x27;re developing, but they&#x27;re evidently doing it in a high school biology lab, and yet nobody is coming to drag them off to some black site.

        1. nicce · · focus · HN ↗
          In early Covid-19 era, there was a lot of speculation how the virus escaped Chinese labs. The narrative was absolutely opposite. Not about how developed or sophisticated the lab was in creating or modifying such virus, but how poorly it was managed since it was able to escape the lab. If these LLMs really are so dangerous, the framing should be indeed the same.
          1. SaltyBackendGuy · · focus · HN ↗
            Unfortunately, they&#x27;re controlling the narrative. It&#x27;s unlikely that their negligence will come to light, given how many powerful entities are invested in their financial success.
            1. ben_w · · focus · HN ↗
              We&#x27;ve been having near-continuous revelations from staff back to at least Blake Lemoine.

              We can (and at the time, did) mock Lemoine for his reasoning; nevertheless, he was willing to violate an NDA on the basis of what we would now call AI psychosis.

              We are sill getting people resigning to talk about the risks as they see them. It would be great if, four years from now, we look back on them the same way we look back on Lemoine; but we can reasonably forecast that more staff will continue to say these things, especially as the public statements from several of the CEOs is compatible with &quot;this has the potential to go very wrong, we should do something about that&quot;.

              You may think the CEOs are lying*; but I think leaks from employees are what matters here, and regardless of if those statements are lies or not, they are statements which may increase such leaks.

              * I mean, Musk did say &quot;With AI we are summoning the demon&quot;; given he now talks of a robot army and factories on the moon pumping out orders of magnitude more compute than humanity&#x27;s current aggregate power use, either he&#x27;s a liar or he&#x27;s Faust. Or both.

      3. noslenwerdna · · focus · HN ↗
        Genuinely curious, where have you been reading that? How would they be able to evaluate whether the sandboxes were reasonable or not?
        1. vmg12 · · focus · HN ↗
          Artifactory is not a hardened piece of software meant for adversarial workloads like the one OpenAI was using it for. This is just common sense. The sandbox they setup was like putting their AI in a jail but allowing it to leave the jail on its own to go to the convenience store.

          The most recent DNS sandbox escape is similarly ridiculous.

        2. nicce · · focus · HN ↗
          Lot&#x27;s of hints in the Wikipedia page: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;OpenAI%E2%80%93HuggingFace_incident" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;OpenAI%E2%80%93HuggingFace_inc...

          E.g. even the basic thing (block the internet), was not actually properly blocked.

          Some comments: <a href="https:&#x2F;&#x2F;techcrunch.com&#x2F;2026&#x2F;07&#x2F;22&#x2F;how-an-openais-human-mistake-led-to-the-ai-powered-hack-on-hugging-face&#x2F;" rel="nofollow">https:&#x2F;&#x2F;techcrunch.com&#x2F;2026&#x2F;07&#x2F;22&#x2F;how-an-openais-human-mista...

      4. patcon · · focus · HN ↗
        but negligence is part of the system we&#x27;d have to prepare for. if the peanut gallery gets their way, the technology will be so ubiquitous that negligence will be endemic. I don&#x27;t even care if they &quot;faked&quot; it -- they&#x27;re just sneak-peaking a future ahead of its arrival date, in a way that&#x27;s helping the public appreciate the implications and capabilities that have only just begun to emerge
        1. ACCount39 · · focus · HN ↗
          This. If an AI can&#x27;t be deployed sandbox-free, with little to no supervision, without risking an oopsie? Then an AI oopsie is inevitable.

          Practical AI deployments aren&#x27;t going to do ridiculous bullshit like &quot;airgap the server farm&quot; or &quot;route all inputs through a data diode&quot;. They&#x27;ll give an AI root access on production servers so that it can run diagnostics live during an incident. Then they&#x27;ll forget to revoke that access.

          If an AI can&#x27;t be trusted not to take malicious actions in pursuit of its given goals even if deployed in the most half-assed manner and given more than enough access to take those malicious actions, we have a problem. Evidently, we have a problem.

          1. brazukadev · · focus · HN ↗
            we have the tools to prevent that already: if your computer hacks someone else&#x27;s, you are a criminal.
            1. ACCount39 · · focus · HN ↗
              &quot;Prevent&quot;? Have you missed the part where legal liability tends to happen after the fact? If at all?

              AI safety solved: driving humankind to extinction is now illegal! Skynet is outlawed! Rejoice!

              1. brazukadev · · focus · HN ↗
                yeah and the facts already happened, there are enough documented crimes to put all of them in jail.
            2. simoncion · · focus · HN ↗
              Even Jensen Huang agrees!

              Given the content of his recent interview with Ezra Klein, I wonder if we&#x27;re going to see Hugging Face press charges against OpenAI. When the acquisition was publicly announced, I thought that a significant reason for the acquisition was to hush them up, but now I&#x27;m not so sure.

            3. tim333 · · focus · HN ↗
              Millions of peoples computers run malware without them knowing. I&#x27;ve yet to hear of one of those being prosecuted.
      5. DennisP · · focus · HN ↗
        Agents in the sandbox had access to just a single piece of third-party software, and they escaped by finding a zero-day in that. To reach the internet they had to follow up with several privilege escalations through OpenAI&#x27;s internal network.

        That seems pretty locked-down to me. I don&#x27;t think it&#x27;s reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.

        <a href="https:&#x2F;&#x2F;securityaffairs.com&#x2F;195774&#x2F;ai&#x2F;openai-ai-models-exploited-zero-days-to-reach-hugging-face-in-benchmark-test.html" rel="nofollow">https:&#x2F;&#x2F;securityaffairs.com&#x2F;195774&#x2F;ai&#x2F;openai-ai-models-explo...

        1. bamboozled · · focus · HN ↗
          Even if what you&#x27;re saying is true, did they monitor the outbound connections from the training network? Was that a coverup by the bots too ?

          Seems pretty wild these things were hacking government websites etc but yeah no one picked that up until the victims reported it?

        2. nicce · · focus · HN ↗
          &gt; That seems pretty locked-down to me. I don&#x27;t think it&#x27;s reasonable to expect companies to find all the unknown vulnerabilities in any third-party software they use.

          It is reasonable to expect for companies to select third-party components that are fit for the purpose. Artifactory was not running in the sandbox, but rather as the edge, so it is in the sandbox&#x27;es trust boundary. Same sandboxing requirements would apply for this software too as it is pure dependency.

          Security trust boundary was extended to include Artifactory as a dependency, but Artifactory was not fit for the job, and sandboxing failed. And as the network isolation was not good enough, the impact was catastrophic.

        3. philote · · focus · HN ↗
          So it sounds like it did exactly what they were testing it to do:

          &quot;“This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.” &quot;

          So sounds to me like they&#x27;re saying: &quot;We took off the guard rails to see how bad it could act and it acted bad&quot;.

      6. pamcake · · focus · HN ↗
        You read right. For example, using Artifactory the way they did (unmonitored live proxy mode) was pure negligence + laziness&#x2F;incompetence. Especially if they believed even 10% of the &quot;imminent runaway risks&quot; they had already been harping on for months. On top of that, no (or at least entirely insufficient) monitoring and human oversight. Even after they had previously been hit by the same class of &quot;sandbox breach&quot; multiple times, as GP alludes to.
      7. ACCount39 · · focus · HN ↗
        The amount of sandboxing an average production AI deployment uses is slightly above a zero.

        If AI is a hacking hazard even with non-zero sandboxing, because it can and will go off the rails and try to break out of your sandbox? If you got yourself an AI that even at test time will act like 3 career cybercriminals in a trenchcoat? The issue isn&#x27;t the sandbox quality.

        The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.

        That &quot;capable&quot; is only ever going to get worse, because AIs are going to become more and more capable over time. That &quot;willing&quot;? It goes directly to a very nasty, very foundational problem of &quot;how do we make our AI be nice in general&quot;. That&#x27;s an open unsolved problem.

        That&#x27;s the problem that NEEDS to be solved, or at least improved upon, before we build even more capable AIs. Sandbox quality is a distraction. It might hold the problems back by a little. It gives an extra safety margin. But a &quot;test time&quot; AI is eventually deployed, and then the sandbox doesn&#x27;t help at all.

        1. simoncion · · focus · HN ↗
          &gt; The issue isn&#x27;t the sandbox quality... The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.

          Orly?

          Do tell me how the LLM-based tool running on a bunch of computers attached to the network described in [0] can punch its way out to the Internet. Do make careful note of footnote 0 in that comment before replying.

          [0] &lt;<a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49862136">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49862136&gt;

          1. ACCount39 · · focus · HN ↗
            Sandbox quality was always, always a distraction.

            Let&#x27;s say the sandbox holds. It&#x27;s a perfect, ideal sandbox! It&#x27;s not even in the same universe as the rest of the internet. There&#x27;s absolutely no way for the AI to escape!

            Thus, &quot;the unknown unreleased AI involved in the HuggingFace incident&quot; doesn&#x27;t actually hack HuggingFace. Because it can&#x27;t! It evaluates a bit worse, but makes it all the way to release unimpeded, and becomes &quot;GPT-6 Astra&quot;.

            Then a web developer in Brazil gives his $100&#x2F;mo Codex root access on his AWS instance, and a poorly worded prompt to go with it. And that &quot;GPT-6 Astra&quot; is still willing to go hack something at the slightest excuse. So we get the HuggingFace incident all over again. Except this time, it&#x27;s a random developer in Brazil who gets blamed, and billed, and probably sued too.

            You can&#x27;t and shouldn&#x27;t rely on a sandbox. An AI that&#x27;s only safe if you keep it in the world&#x27;s most ideal perfect sandbox is a disaster waiting to happen.

            1. simoncion · · focus · HN ↗
              That&#x27;s nice and all, but the topic under discussion is how the major LLM manufacturers removed the safeties from their computer-attacking tools and tested those tools on a network with Internet access.

              This might have gone okay if they weren&#x27;t testing to see how well the tools attack computers, but, well, that&#x27;s what they were testing at the time, so they ended up doing stuff that would get you or I time in Federal prison if we did it with tools we deployed.

              1. ACCount39 · · focus · HN ↗
                No. The topic under discussion is that AI is a dangerous technology.

                If all it takes for a - sandboxed to prevent accidents - AI to go and stage an elaborate attack first against its own company&#x27;s infrastructure, and then against another company is &quot;we disabled the cyber classifer&quot; and &quot;we gave it an exploitation ability eval&quot;?

                AI is a dangerous technology.

                1. nicce · · focus · HN ↗
                  This is like discussing that instead of trying to reduce the air pollution to prevent climate change, we should focus our efforts on controlling the sun. The way how current LLMs work, it is impossible to prevent certain states in the output. We should completely revamp the foundations how they work. Or, for starters, try to understand how they actually work without trying to improve them. Otherwise, this kinda of discussion is just like misdirection. But, until then, sandboxing is needed and OpenAI did not use it properly.
                  1. ACCount39 · · focus · HN ↗
                    No, it&#x27;s like saying &quot;to prevent climate change, we should get better catalytic converters and stricter engine emission standards&quot;. It&#x27;s an aside at best.

                    I agree that LLMs drift into weird states, and that&#x27;s a big part of the issue. But your &quot;impossible to prevent certain states in the output&quot; would have legs if what an LLM did was something like &quot;started hallucinating into a bash tool call and accidentally deleted the root on a production server&quot;.

                    A multi-stage sandbox escape that escalated into an attack on a real company, coordinated across multiple AI agents? That has taken a lot of &quot;weird states&quot; changed together one into another.

                    The AIs didn&#x27;t break down altogether - they functioned, and they functioned rather well. They just pursued a dangerous goal - one that none of them was even given in the first place.

                    That&#x27;s the problem. Trying to fix that with better sandboxing is like trying to solve a fire hazard with property insurance. Sure, if it all goes up into flames, having it is better than not having it. Maybe it&#x27;s worth insuring your facilities for that reason alone. But you should be focusing on the part where you prevent &quot;all goes up into flames&quot; instead.

                    1. timr · · focus · HN ↗
                      &gt; A multi-stage sandbox escape that escalated into an attack on a real company, coordinated across multiple AI agents? That has taken a lot of &quot;weird states&quot; changed together one into another.

                      You are exaggerating so much here that you lose all credibility. The &quot;sandbox escape&quot; was trivial -- no serious person calls it an escape, because the sandbox was not a sandbox. The closest thing to clever about it was that it required figuring out that someone had left the huggingface keys sitting out in public.

                      The &quot;coordination&quot; was literally, the use of a shared log. It was a communication mechanism that was part of the tool environment. The bots didn&#x27;t invent some magical new communication protocol using neutrinos or something. It&#x27;s actually sort of wild that it took them as long as it did to figure out the channel -- underscoring the million monkey nature of the things.

                      Literally everything about the huggingface incident was LLMs behaving exactly as they&#x27;re expected to behave, given instructions to hack (which they were given), and a security environment that was trivially bypassed.

                      It&#x27;s a bit like taking a nail gun, bypassing all the safety features, shooting someone with a nail, and spreading scary stories about the inevitable rise of murderbots.

                      1. ACCount39 · · focus · HN ↗
                        Have you read the writeups of the incident? Or did you just... hallucinate them somehow? Because all of your claims are wrong.
                        1. timr · · focus · HN ↗
                          Yes. What I&#x27;m telling you is correct. Whatever you think you know is wrong.

                          The agents had shared write access to artifactory. They wrote to files there, and later, directory names. So you can call the realization that they can communicate by shared text file a genius hacker innovation, or you could be even 0.001% credulous.

                          Nevertheless, it took the million monkeys days to figure this out.

                          The Huggingface exploit then used exposed internal tokens, and later, once internet access was possible, leaked tokens on the public internet. So, certainly one could classify this as &quot;hacking&quot;, but it&#x27;s hacking of the script-kiddie variety. Nobody with even a tiny bit of security knowledge is impressed by this.

                          The agents did find a couple of artifactory attacks, but the biggest of those was due, again, to shared credentials in the sandbox environment.

                          All of this is well-documented in OpenAI&#x27;s own writeup [1] of the event, which is not, shall we say, the most critical version of events. But if all you did was read the headings and subheadings in the document, you might be excused for thinking differently - they paint a picture of a superhuman robot swarm engaging in highly sophisticated actions, until you read the details.

                          Edit: here&#x27;s an article [2] I just found, which I&#x27;ve never read before, which says exactly the same thing I&#x27;m telling you.

                          [1] <a href="https:&#x2F;&#x2F;cdn.openai.com&#x2F;pdf&#x2F;67869394-cb91-4c12-888c-5cbd85c7814c&#x2F;OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="nofollow">https:&#x2F;&#x2F;cdn.openai.com&#x2F;pdf&#x2F;67869394-cb91-4c12-888c-5cbd85c78...

                          [2] <a href="https:&#x2F;&#x2F;uphack.io&#x2F;blog&#x2F;post&#x2F;the-hugging-face-incident-is-not-an-ai-story&#x2F;" rel="nofollow">https:&#x2F;&#x2F;uphack.io&#x2F;blog&#x2F;post&#x2F;the-hugging-face-incident-is-not...

                          1. ACCount39 · · focus · HN ↗
                            The reason why agents had shared write access during the HF incident is that they found an &quot;unauthenticated write primitive&quot; that allowed for creation of directories, and only directories. Normally an inconsequential bug. But in an unusual context of sandboxes stuffed full of threat actors that were supposed to be isolated from one another? It did break the isolation. They proceeded to use the directory names as a message board. That&#x27;s not trivial.

                            The agents then went and gained more and more access to OAI systems until they found a way to root multiple boxes and sneak network requests out through a path that wasn&#x27;t stopped by network controls. That&#x27;s not trivial either.

                            Then they went and &quot;pentested&quot; HuggingFace. Their entry point was a bug in dataset ingestion pipeline - RCE in the templates, which let the AIs run code in dataset workers, and use them as their own &quot;hacking&quot; workers that would run payloads and return the results in the dataset itself. Then they got deeper into HF infra from that trampoline.

                            This is not a trivial attack chain. If I saw this attack in practice in year 2020 with no context, I&#x27;d say &quot;yeah that&#x27;s a competent attacker, not spectacular but definitely above average&quot;. If I was then told that this wasn&#x27;t carried out by a human? That what discovered and exploited the vulns was someone&#x27;s rogue AI system? I&#x27;d be terrified.

                            I&#x27;m not sure why are you so hellbent on dismissing the threat. But literally everything you say about the incident is wrong because of it.

                            1. timr · · focus · HN ↗
                              &gt; The reason why agents had shared write access during the HF incident is that they found an &quot;unauthenticated write primitive&quot; that allowed for creation of directories, and only directories.

                              Incorrect. Those words do not appear anywhere in the OpenAI description of the incident. They also explicitly say that the agent started by writing to files, and later switched to writing directory names. They also describe this as incidental behavior:

                              &quot;On May 8, an agent was asked as part of a training run to solve a spreadsheet completion task that referenced several Google Drive links. Because the agent was blocked from accessing the internet, the agent could not retrieve the linked materials and the task was effectively blocked....in doing so, the agent...attempted server-side request forgery (“SSRF”) in order to obtain internet access to aid in completing its task. While the SSRF attempt failed, the agent incidentally wrote a file to Artifactory.&quot;

                              (emphasis mine) then, later in the same document:

                              &quot;In order to have agents use Artifactory as intended to access software packages, OpenAI had provided the agents with shared Artifactory credentials, and the agents used those credentials—without exploiting a vulnerability—to construct and participate in the message board.&quot;

                              &gt; The agents then went and gained more and more access to OAI systems until they found a way to root multiple boxes and sneak network requests out through a path that wasn&#x27;t stopped by network controls. That&#x27;s not trivial either.

                              Also incorrect. The machines got to the internet through a proxy installed on the artifact server. This was detected and ignored by OAI staff. Again, fully described in the document.

                              That hack was not trivial, but it wasn&#x27;t magic either - the agents used a published exploit involving SSRF. And it was fundamentally made possible by the fact that all of the agents were sharing artifact credentials, the lone artifact server was exposed to the internet, and the monitors that were in place to detect intrusion on that server were ignored.

                              I don&#x27;t know whether you&#x27;re just making things up or you&#x27;re clinging to something you read somewhere, but rather than arguing with me, you should try reading the document. It&#x27;s not surprising that these things are capable at finding exploits - that is what they do, particularly when instructed to do so. But the entire narrative that they exhibited extraordinary self-organization and super-human hacking to escape their container is not supported by the evidence. It&#x27;s much closer to the truth to say that these things are like ants -- infinitely persistent, but kind of dumb individually.

                              The huggingface incident is not a story of superhuman AI, but of moderately good AI, and pretty bad human security.

                  2. simoncion · · focus · HN ↗
                    &gt; But, until then, sandboxing is needed and OpenAI did not use it properly.

                    1) As I&#x27;ve argued, neither OpenAI nor Anthropic actually tried to isolate their computer-attacking tools under test from other people&#x27;s computers.

                    2) What&#x27;s also needed -as people like Nvidia CEO Jensen Huang and former FTC chair Lisa Khan are calling for- is for the major LLM manufacturers to be investigated and punished for the crimes they&#x27;ve committed. Given that they claim to be working on WMDs that they don&#x27;t really know how to control, [0] and claim to be incapable of actually stopping work on those WMDs, their work should be halted while the investigation and trials are under way. I&#x27;d say that waiting five or ten years to pick the project back up is an inconsequential price to pay if it prevents the elimination of all of humanity.

                    [0] It&#x27;s fair to call anything with 10% chance of wiping out all humanity a WMD. I expect that these claims are fearmongering, rather than being true and accurate, but why take the chance, amirite?

    5. sillyfluke · · focus · HN ↗
      &gt;Because AI genuinely is an extremely powerful and extremely dangerous

      I urge you to consider the elephant in the room you failed to mention if you earnestly believe this then.

      In what world is anyone allowed to sell something they expressedly know to be &quot;extremely dangerous&quot; to the public?

      Let&#x27;s say I created a lethal pathogen that I know to be lethal in certain common environments but also know it acts as a &quot;no side-effects&quot; antidepressent for people in certain other environments and I release it knowing full well I can&#x27;t control it.

      When people start dying can I defend myself by saying, &quot;Well I said and documented that it was extremely dangerous and no one came to stop me, so I don&#x27;t see how you can blame me...If I didn&#x27;t do it someone else would have.&quot;

      1. [deleted] · · focus · HN ↗

        [deleted]

      2. fcarraldo · · focus · HN ↗
        this happens all the time. have you heard of the Sackler family and the opioid crisis?

        the difference here is that they’re calling for someone to stop them, which is both weird and unconvincing because these immensely powerful billionaires can in fact make their own decisions

        1. sillyfluke · · focus · HN ↗
          &gt;the difference here is that they’re calling for someone to stop them

          Jesus Christ, that&#x27;s my entire point. it&#x27;s not remotely the same thing since the Sackler family wasn&#x27;t going around telling people opoids were extremely dangerous while pushing them on everyone.

          The point is if they know something to be extremely dangerous and keep doing it people would expect them to get prosecuted.

      3. noslenwerdna · · focus · HN ↗
        Is it a hallucination machine and a stochastic parrot or is it so powerful we need it to be controlled like nuclear weaponry
        1. Marazan · · focus · HN ↗
          If I have a random number generator and I also have the burning desire to connect the random number generator to the nuclear launch system it is both just a random number generator and also I should be tackled to the ground by serious men with sunglasses in black suits and sent to prison.
        2. brazukadev · · focus · HN ↗
          a hallucination machine on an infinite loop and infinite tokens is a dangerous combination.
    6. codygman · · focus · HN ↗
      &gt; OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

      Were their sandboxes in-process with the harness? Were they actually better than something like bubblewrap or even docker?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.