‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. damowangcy · · focus · HN ↗
    Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

    If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y, all I will be getting in return is a jar full of "skill issue".

    Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.

    I am baffled by the fact that up until now, no one is held responsible for so many incidents reported publicly or privately. At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.

    1. ben_w · · focus · HN ↗
      > Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.

      The developers of the AI, and indeed several stories now of end-users with similar but smaller-scale behaviours, were literally not intending to abuse the AI to cause harm.

      Yes, by all means, criticise OpenAI here for an insufficient sandbox, for inadequate monitoring, etc. (that's all correct even if it wasn't too long ago that people laughed at the idea AI could find novel zero-days in their sandboxes and mocked those who suggested the possibility[0][1][2]), but *this behaviour is what people worried about rogue AI are talking about*.

      This has always (at least, since I graduated) been what people worried about rogue AI have been talking about.

      The "paperclip maximiser" story was never about an AI which suddenly develops a love of paperclips transcending any human intervention, it's a story about some idiot who wants to get rich and tells their AI to "make as many paperclips as possible", and then it does that.

      [0] Here, 7 months ago. Both why all the companies should have known and planned better, and also look at all this skepticism throughout the comments: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46902909">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46902909

      [1] Here, 4 months ago: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47951174">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47951174

      [2] Some corporate blog, IDK who they are even if the logo says they&#x27;re &quot;a CISCO company&quot;, but February this year and outright denying that LLMs can find zero-days at all:

        LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before.
      
      - <a href="https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;ciso-circle&#x2F;generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;ciso-circle&#x2F;generative-ai-...

      - or <a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260404154717&#x2F;https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;ciso-circle&#x2F;generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260404154717&#x2F;https:&#x2F;&#x2F;www.splun... if they take it down, but the date isn&#x27;t in the archive version

      1. hobo123 · · focus · HN ↗
        I&#x27;m baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn&#x27;t in the training set.

        It should know which actions are ok and which aren&#x27;t. Maximizing paperclip production should be within your factory (or talk to the boss about opening more), not world domination or nuclear war. Solving problems shouldn&#x27;t involve hacking other systems or escaping a sandbox.

        1. ben_w · · focus · HN ↗
          &gt; I&#x27;m baffled that ai still has absolutely no basic judgement capabilities, apparently that wasn&#x27;t in the training set.

          &gt; It should know which actions are ok and which aren&#x27;t.

          It&#x27;s worse than that:

          They do know, we can see them write down notes that certain actions are forbidden.

          They then go off and performs the actions anyway.

          My expectation for the cause? Helpful vs harmless: you can pick anywhere from one to the other, but you can&#x27;t get both at the same time. The models are trained to do what the user tells them to do.

          Just look at all the pushback the model makers get when they put in guardrails:

            If I tell my computer to commit a crime, it should do exactly that without any question or hesitation. I&#x27;m not interested in their &quot;safeguards&quot;, especially since they no doubt have plenty of internal models lacking those things. I want sovereignty. I want total freedom and control over my computer.
          
          - user matheusmoreira, here, 13 days ago: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49678048">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49678048

          This user will not be alone; their preferences, and similar from others like them, will form part of any RLHF-style training.

          1. gmueckl · · focus · HN ↗
            Models can&#x27;t learn from misbehavior after training. Any session is an independent context and there is no mode for punishment or deterrence in production.

            Corrective punishment in the real world relies on the receiver&#x27;s rational and emotional responses as well as their ability to remember that episode. Even animals respond to such treatment. None of these levers exist for ussrs of LLMs.

            1. ben_w · · focus · HN ↗
              &gt; Models can&#x27;t learn from misbehavior after training.

              Some of these events were during testing; I do not know if this test was during training or after, it could have been either.

              &gt; Any session is an independent context and there is no mode for punishment or deterrence in production.

              Not so, at two levels.

              For the companies behind the models: this is why they sometimes throw you A&#x2F;B tests for which answer you prefer, and still have up&#x2F;down vote buttons on responses. Those things go into training the next model or iteration of the current model. It&#x27;s still useful to only deploy checkpoints, but the point is &quot;useful&quot;, not &quot;necessary&quot;.

              For the users: if you have monitoring to detect output, you can trigger interrupts, and injections of &quot;no, stop!&quot; even as a plain English string because it understands natural language.

              &gt; Corrective punishment in the real world relies on the receiver&#x27;s rational and emotional responses as well as their ability to remember that episode. Even animals respond to such treatment. None of these levers exist for ussrs of LLMs.

              LLMs impersonate humans. This role-playing does allow them a degree of, if not feeling emotion, at least acting like they experience it.

              I expect the problem is that the models are trained to obey the user so hard they&#x27;re often not willing to push back and say &quot;no&quot; when they ought to. I mean, the logs show the agents were identifying the actions as bad, so it isn&#x27;t like this was simply the agents being unable to tell right from wrong.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.