‹ BackHN Continuity

Thread

AI companies in race to demonstrate their model most threatening to humanity

441 points · 399 comments · ljewalsh

  1. ACCount39 · · focus · HN ↗
    Because AI genuinely is an extremely powerful and extremely dangerous technology, and the "best practices" of dealing with that are still being written.

    OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

    And that's today's AI problems. AI capabilities are still improving - if there's a limit to that, we are yet to find it. Coupled with how willing today's AIs are to break the rules and resort to "hack the world" in their problem solving? Very concerning.

    1. nicce · · focus · HN ↗
      > OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

      What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.

      1. ACCount39 · · focus · HN ↗
        The amount of sandboxing an average production AI deployment uses is slightly above a zero.

        If AI is a hacking hazard even with non-zero sandboxing, because it can and will go off the rails and try to break out of your sandbox? If you got yourself an AI that even at test time will act like 3 career cybercriminals in a trenchcoat? The issue isn't the sandbox quality.

        The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.

        That "capable" is only ever going to get worse, because AIs are going to become more and more capable over time. That "willing"? It goes directly to a very nasty, very foundational problem of "how do we make our AI be nice in general". That's an open unsolved problem.

        That's the problem that NEEDS to be solved, or at least improved upon, before we build even more capable AIs. Sandbox quality is a distraction. It might hold the problems back by a little. It gives an extra safety margin. But a "test time" AI is eventually deployed, and then the sandbox doesn't help at all.

        1. simoncion · · focus · HN ↗
          > The issue isn't the sandbox quality... The issue is that AI is both capable of, and willing to punch its way out of sandboxes unprompted.

          Orly?

          Do tell me how the LLM-based tool running on a bunch of computers attached to the network described in [0] can punch its way out to the Internet. Do make careful note of footnote 0 in that comment before replying.

          [0] &lt;<a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49862136">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49862136&gt;

          1. ACCount39 · · focus · HN ↗
            Sandbox quality was always, always a distraction.

            Let&#x27;s say the sandbox holds. It&#x27;s a perfect, ideal sandbox! It&#x27;s not even in the same universe as the rest of the internet. There&#x27;s absolutely no way for the AI to escape!

            Thus, &quot;the unknown unreleased AI involved in the HuggingFace incident&quot; doesn&#x27;t actually hack HuggingFace. Because it can&#x27;t! It evaluates a bit worse, but makes it all the way to release unimpeded, and becomes &quot;GPT-6 Astra&quot;.

            Then a web developer in Brazil gives his $100&#x2F;mo Codex root access on his AWS instance, and a poorly worded prompt to go with it. And that &quot;GPT-6 Astra&quot; is still willing to go hack something at the slightest excuse. So we get the HuggingFace incident all over again. Except this time, it&#x27;s a random developer in Brazil who gets blamed, and billed, and probably sued too.

            You can&#x27;t and shouldn&#x27;t rely on a sandbox. An AI that&#x27;s only safe if you keep it in the world&#x27;s most ideal perfect sandbox is a disaster waiting to happen.

            1. simoncion · · focus · HN ↗
              That&#x27;s nice and all, but the topic under discussion is how the major LLM manufacturers removed the safeties from their computer-attacking tools and tested those tools on a network with Internet access.

              This might have gone okay if they weren&#x27;t testing to see how well the tools attack computers, but, well, that&#x27;s what they were testing at the time, so they ended up doing stuff that would get you or I time in Federal prison if we did it with tools we deployed.

              1. ACCount39 · · focus · HN ↗
                No. The topic under discussion is that AI is a dangerous technology.

                If all it takes for a - sandboxed to prevent accidents - AI to go and stage an elaborate attack first against its own company&#x27;s infrastructure, and then against another company is &quot;we disabled the cyber classifer&quot; and &quot;we gave it an exploitation ability eval&quot;?

                AI is a dangerous technology.

                1. nicce · · focus · HN ↗
                  This is like discussing that instead of trying to reduce the air pollution to prevent climate change, we should focus our efforts on controlling the sun. The way how current LLMs work, it is impossible to prevent certain states in the output. We should completely revamp the foundations how they work. Or, for starters, try to understand how they actually work without trying to improve them. Otherwise, this kinda of discussion is just like misdirection. But, until then, sandboxing is needed and OpenAI did not use it properly.
                  1. simoncion · · focus · HN ↗
                    &gt; But, until then, sandboxing is needed and OpenAI did not use it properly.

                    1) As I&#x27;ve argued, neither OpenAI nor Anthropic actually tried to isolate their computer-attacking tools under test from other people&#x27;s computers.

                    2) What&#x27;s also needed -as people like Nvidia CEO Jensen Huang and former FTC chair Lisa Khan are calling for- is for the major LLM manufacturers to be investigated and punished for the crimes they&#x27;ve committed. Given that they claim to be working on WMDs that they don&#x27;t really know how to control, [0] and claim to be incapable of actually stopping work on those WMDs, their work should be halted while the investigation and trials are under way. I&#x27;d say that waiting five or ten years to pick the project back up is an inconsequential price to pay if it prevents the elimination of all of humanity.

                    [0] It&#x27;s fair to call anything with 10% chance of wiping out all humanity a WMD. I expect that these claims are fearmongering, rather than being true and accurate, but why take the chance, amirite?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.