‹ BackHN Continuity

Thread

AI companies in race to demonstrate their model most threatening to humanity

441 points · 399 comments · ljewalsh

  1. ACCount39 · · focus · HN ↗
    Because AI genuinely is an extremely powerful and extremely dangerous technology, and the "best practices" of dealing with that are still being written.

    OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

    And that's today's AI problems. AI capabilities are still improving - if there's a limit to that, we are yet to find it. Coupled with how willing today's AIs are to break the rules and resort to "hack the world" in their problem solving? Very concerning.

    1. nicce · · focus · HN ↗
      > OpenAI, for example, thought their sandboxes were good enough. As their AIs got more and more advanced, they kept proving them wrong - sandbox after sandbox.

      What I have been reading, was that their sandboxes were so poor that it was pure negligence. I am still waiting to see if some external and neutral cybersecurity company with high reputation would audit their sandboxes and how they are being used.

      1. simianwords · · focus · HN ↗
        The model found and exploited and chained together previously unknown vulnerabilities.

        How were the sandboxes poor?

        1. dns_snek · · focus · HN ↗
          Agents didn't have real network isolation. They were indirectly connected to the internet via a jump host running insecure software which was never designed or hardened to provide any kind of isolation.
          1. simianwords · · focus · HN ↗
            But this level of isolation is what happens normally. At least in my university and another company I worked at. It wasn’t running insecure software, as far as anyone knew, it was secure
            1. dns_snek · · focus · HN ↗
              There's levels of isolation, a padlock is not equivalent to a bank vault. If you claim to be building a possibly world-ending AI then you don't get to use a padlock and call it a day.
              1. simianwords · · focus · HN ↗
                And they are not calling it a day and they have done most things possible to communicate the fact that they are building something dangerous.
                1. dns_snek · · focus · HN ↗
                  They "called it a day" when they decided to deploy an agent which they believe to be dangerous inside this poorly secured environment.

                  There was no real isolation because a part of the system that doesn't provide any isolation guarantees was bridged to the internet. The next version of Artifactory, which you definitely wouldn't audit before you rolled it out, could simply add a public API that sends requests out to the internet.

                  Such an innocent upstream change would be equally catastrophic for your security model, which should demonstrate why it's negligent to rely on undefined behavior to enforce your security policies.

                  This should frankly be obvious any operator entrusted with running dangerous and possibly malicious code. Even if you don't know what you're doing, any LLM would tell you that this is a really bad idea if you simply asked. Don't rely on spacebar heating [1] to keep humanity alive.

                  [1] <a href="https:&#x2F;&#x2F;xkcd.com&#x2F;1172&#x2F;" rel="nofollow">https:&#x2F;&#x2F;xkcd.com&#x2F;1172&#x2F;

                2. uneoneuno · · focus · HN ↗
                  What about the waves of employees quitting both OAI and anthropic stating as their reason that they see their work as a direct threat to the future of humanity, and that their management isn&#x27;t taking it seriously enough. Is that a PR stunt too?
        2. vmg12 · · focus · HN ↗
          The model is really good at hacking, we all know this. This is why you don&#x27;t just expose random pieces of software to it without that software being hardened.

          It&#x27;s not like the model managed to exploit firecracker itself (no model has been capable of this), the model exploited artifactory.

          Artifactory is not some hardened piece of software that is meant to block users from accessing the internet through it.

          1. solenoid0937 · · focus · HN ↗
            &gt; The model is really good at hacking, we all know this

            No, we didn&#x27;t know that and this is how you find out they&#x27;re very good at hacking

            HN&#x27;s memory is so fickle. Just a few months ago almost no one here believed Mythos could actually be as good at hacking as the company claimed. This was a novel concept when the companies experienced these breakouts.

            1. vmg12 · · focus · HN ↗
              Models have been good at finding exploits for half a year now, this is not how we found out LLMs were good at hacking, you are rewriting history.

              We knew models much weaker than mythos were good at hacking the problem they had was that when finding exploits they had too many false positives.

              Either way, putting artifactory on the sandbox security boundary is obscene negligence. There is no reason to believe artifactory is secure.

              1. solenoid0937 · · focus · HN ↗
                If you listen to the OpenAI Black Hat talk it is very obvious they were surprised at the level of capability on display and felt it was novel.

                But I guess OpenAI&#x27;s security researchers acting surprised is part of some grand conspiracy to manage PR?

                1. vmg12 · · focus · HN ↗
                  It may have been novel for OpenAI but we already had Mythos at this point and this talk by Nicholas Carlini.

                  <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=1sd26pWhfmg" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=1sd26pWhfmg

                  We already knew LLMs were capable of finding exploits like this.

                  1. solenoid0937 · · focus · HN ↗
                    The Mythos issue is different though. It did not use any novel exploits. Irregular, the vendor they were using, misconfigured the environment to allow internet. Every other AI company also uses Irregular and that&#x27;s why we saw so many articles come out at once.

                    The OpenAI HF incident is separate from this. It involved actual zero days, teamwork and message passing, and sophisticated chains of exploits.

                    1. nicce · · focus · HN ↗
                      &gt; The OpenAI HF incident is separate from this. It involved actual zero days, teamwork and message passing, and sophisticated chains of exploits.

                      What exactly? They seem to be trivial SSR&#x2F;path traversal and input validation issues. Including misconfiguration. Nothing novel.

                    2. vmg12 · · focus · HN ↗
                      I&#x27;m just not impressed by AI finding zero days in unhardened software like artifactory. I&#x27;m completely unsurprised it had zero days and I fully expect AI to find any that exist. AI will literally try all combinations of inputs to achieve its goals. Any exploits that exist will be found.

                      The question anyone versed in security would ask was why anyone thought artifactory was an acceptable security boundary. I would never assume artifactory was secure. It&#x27;s like someone telling me there is a 0 day in a wordpress extension. So what?

                      Also I don&#x27;t think Qemu is secure either because it&#x27;s millions of lines of C and C++.

                      Firecracker I can trust to be secure because it&#x27;s 70k lines of human audited Rust. I know there are multiple people that have a complete understanding of the firecracker codebase.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.