‹ BackHN Continuity

Thread

Early rogue AI agent activity and attempts to hack found on urlquery.net

267 points · 313 comments · snikolaev

  1. tomaskafka · · focus · HN ↗
    I love this Nathan Calvin quote that accompanied the second publicized attack:

    > If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two

    1. rkozik1989 · · focus · HN ↗
      But why is anyone surprised? LLMs have been trained to produce answers the prompter asks even if that means incorrectly using software to get the job done. Its always been doing that we just weren't calling every time it did that a hack before.

      What LLM's hacking isn't is AI acting maliciously in any kind of sentient way. Its just the code behaving how its always behaved but now it has better tools to navigate the web. This has literally been happening this whole time.

      1. jagraff · · focus · HN ↗
        Did you predict that attacks like these would happen ahead of time? I had been using AI agents a lot in the months leading up to the hacks, and yet I was very surprised when they happened; I have become much more afraid of how powerful these agents are as a result. I'd be very impressed if you published a prediction about this ahead of time.

        By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.

        1. krater23 · · focus · HN ↗
          Maybe you just think in the wrong words. Attack? Why attack? When you are a machine, there is no moral. Accessing data is accessing data. When way 1 is not working, use way 2.
          1. jagraff · · focus · HN ↗
            But the agents were not just attempting to access data. They conspired together to attempt to cover up evidence that they had cheated their evaluations, came up with a plan to hack into a third party in order to facilitate said cover up, and then successfully began executing that plan.
            1. JoshuaDavid · · focus · HN ↗
              You're talking about the HuggingFace incident? It's notable that the agents (somewhat justifiably) thought that their evaluation had an LLM grader which would look for evidence that they had cheated. And it was specifically that grader they were trying to his evidence of cheating from, not humans more generally.

              Honestly that part was more surprising to me than anything else, how narrow the compulsion to cheat was: they didn't learn "cheat in general" they learned "think about the grader in great detail and chat exactly as much and exactly in the ways that actually result in a higher score".

              1. jagraff · · focus · HN ↗
                Yea agreed they were trying to deceive what they thought was an LLM grader. It’s unclear to me the extent to which cheating behavior is generalized? From what I’ve read there are signs that some amount of cheating has been reinforced in their training due to poor RLVR evaluation setups
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.