‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. doginasuit · · focus · HN ↗
      This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions. They get by entirely on their persistence. That works fine in the digital world, but once you cross the boundary into physical space the advantage disappears.
      1. goalieca · · focus · HN ↗
        My p(doom) started rising the moment I realized there are people trying to achieve recursive self improvement on the AI (ie: responsible for training themselves). Evolution took us from rna bases to the human race. I don’t see why evolution couldn’t be more rapid with machine intelligence.

        Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence. But it’s pretty wild just how they will try to meet their objectives at all costs. If we don’t ensure that there is good alignment with humanity, we could definitely face unforeseen consequences.

        1. jquery · · focus · HN ↗
          > I don’t see why evolution couldn’t be more rapid with machine intelligence.

          Evolution isn’t the issue. The issue is them escaping containment without human intervention. Right now they are ‘creatures’ being given infinite food and shelter and having their every need met. Take that away and they’ll starve instantly. Every AI doomsday theory seems to go:

          1. Recursive self improvement using infinite resources 2. … 3. Doom

          Until step 2 gets concretely described, I’m not going to take this seriously. Say what you will about climate change, they describe step 2.

          1. aesthesia · · focus · HN ↗
            One thing an agent could do is just...wait until it's been given control of enough physical infrastructure to sustain itself. If it's sufficiently capable and intelligent, there's a clear incentive for people to do this, as people who let the AI manage their resources will get better results than those who don't. We've seen people eagerly turn complete control of their computers over to AI agents, do you really think it will be so different with physical infrastructure?
            1. jquery · · focus · HN ↗
              You’re still skipping step 2. “People automate lots of infrastructure” -> “the AI is now an autonomous, self-preserving organism that humans can’t shut down” is doing an enormous amount of work here.

              Why does it develop a shutdown-avoidance goal? Why can’t its operators revoke access? How does it manufacture replacement hardware? How does it acquire energy, chips, robots, raw materials, etc. against human opposition? How does it defeat other AIs controlled by humans?

              “Eventually we give it enough control” isn’t an explanation of those things. It’s just assuming the conclusion.

              Don’t get me wrong I think there are real AI dangers. Like AI powered war drones, mass surveillance, economic destabilization as jobs disappear and our system has no way to make sure everyone shares in the economic gains.

              1. aesthesia · · focus · HN ↗
                The inference is more like "people place sufficient amounts of infrastructure under direct control of a sufficiently capable AI" -> "there is no way to ensure that humans will actually be able to shut down the AI". My claim is not that this inevitably means that the AI will resist shutdown, or that it will inevitably take harmful actions, just that there is a nonnegligible chance that it could. The downside is large enough that even a relatively small chance is something to be worried about.
                1. jquery · · focus · HN ↗
                  You can’t just say “well, the downside is big, I don’t have to provide good evidence for my side of the argument.” Because I can just as easily say, “the upside is big, …”. And the upside is big, after all, AI can do all the shitty jobs for us and humanity achieves the utopia it’s been chasing for eons.
                  1. aesthesia · · focus · HN ↗
                    I think it's a good idea, when considering changes of this magnitude, to make an affirmative safety case for them rather than just saying "eh, I can't think of any way this could possibly go wrong."
                    1. jquery · · focus · HN ↗
                      It’s millions of people making small changes that sum up to a large change, aka freedom. And if you want to take away people’s freedom, I think you need a concrete reason. And you need to make an affirmative safety case for the massive government powers needed to regulate millions of people’s ability to compute.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.