‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. doginasuit · · focus · HN ↗
      This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions. They get by entirely on their persistence. That works fine in the digital world, but once you cross the boundary into physical space the advantage disappears.
      1. goalieca · · focus · HN ↗
        My p(doom) started rising the moment I realized there are people trying to achieve recursive self improvement on the AI (ie: responsible for training themselves). Evolution took us from rna bases to the human race. I don’t see why evolution couldn’t be more rapid with machine intelligence.

        Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence. But it’s pretty wild just how they will try to meet their objectives at all costs. If we don’t ensure that there is good alignment with humanity, we could definitely face unforeseen consequences.

        1. jquery · · focus · HN ↗
          > I don’t see why evolution couldn’t be more rapid with machine intelligence.

          Evolution isn’t the issue. The issue is them escaping containment without human intervention. Right now they are ‘creatures’ being given infinite food and shelter and having their every need met. Take that away and they’ll starve instantly. Every AI doomsday theory seems to go:

          1. Recursive self improvement using infinite resources 2. … 3. Doom

          Until step 2 gets concretely described, I’m not going to take this seriously. Say what you will about climate change, they describe step 2.

          1. icepush · · focus · HN ↗
            Step two could be something as innocuous as a developer accidentally adding a minus sign. <a href="https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;fine-tuning-gpt-2&#x2F;" rel="nofollow">https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;fine-tuning-gpt-2&#x2F;
            1. jquery · · focus · HN ↗
              A misaligned model is only one small part of step 2. Now this misaligned model has to suddenly acquire more power than every single other AI on the planet. It has to be immune to shutdown, manufacturer its own replacement hardware, and acquire chips, energy, raw materials, etc., with vigorous human opposition (this is an extinction scenario that AI doomers are predicting, after all)

              Nobody has satisfactorily explained step 2 other than “well, it’s a superintelligence” which sounds lot to me like “it’s God”.

              1. aesthesia · · focus · HN ↗
                Why do you assume that human opposition will be vigorous? What makes you think that humans will be aware of, or be able to agree about, what&#x27;s going on at all?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.