‹ BackHN Continuity

Thread

Early rogue AI agent activity and attempts to hack found on urlquery.net

267 points · 313 comments · snikolaev

  1. tomaskafka · · focus · HN ↗
    I love this Nathan Calvin quote that accompanied the second publicized attack:

    > If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two

    1. rkozik1989 · · focus · HN ↗
      But why is anyone surprised? LLMs have been trained to produce answers the prompter asks even if that means incorrectly using software to get the job done. Its always been doing that we just weren't calling every time it did that a hack before.

      What LLM's hacking isn't is AI acting maliciously in any kind of sentient way. Its just the code behaving how its always behaved but now it has better tools to navigate the web. This has literally been happening this whole time.

      1. jagraff · · focus · HN ↗
        Did you predict that attacks like these would happen ahead of time? I had been using AI agents a lot in the months leading up to the hacks, and yet I was very surprised when they happened; I have become much more afraid of how powerful these agents are as a result. I'd be very impressed if you published a prediction about this ahead of time.

        By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.

        1. sanderjd · · focus · HN ↗
          I find it surprising that you were surprised. Security is generally quite poor.
          1. jagraff · · focus · HN ↗
            What is surprising is that various agents independently found ways to communicate, conspired together to attempt to cover up evidence that they had cheated their evaluations, came up with a plan to hack into a third party in order to facilitate said cover up, and then successfully began executing that plan. I did not expect that AI agents would be capable of that level of sophisticated goal seeking and collaboration.
            1. sanderjd · · focus · HN ↗
              It's really interesting how people's expectations differ so much. I found the writeup of the incident fascinating and super worrisome, but none of the capabilities demonstrated in it seemed surprising to me at all.
              1. jagraff · · focus · HN ↗
                None of the capabilities in isolation were surprising to me - we already knew that mythos could find zero-days. What surprised me was the decisions the agents made and the swarming behavior they exhibited
                1. sanderjd · · focus · HN ↗
                  That makes sense. I agree that it's the "capabilities in isolation" that I did not find surprising. I guess I was somewhat less surprised than you by the way those isolated capabilities aggregated into the behavior we saw, but I understand your surprise better now.
            2. nottorp · · focus · HN ↗
              If you rename goal seeking to sharing (parts of) their context is it still surprising?
              1. jagraff · · focus · HN ↗
                I don’t agree that “sharing their context” is an accurate description of what happened
                1. sanderjd · · focus · HN ↗
                  You don't? They communicated by writing documents in locations that could be found later. I guess I might say it was more like "sharing their portions of their output" rather than "context", but the distinction seems murky.

                  One reason that I did not find this communication mechanism surprising is that it's exactly how agents I'm using communicate with each other or across a time gap. "I've saved our plan for where to start tomorrow in start-here.md". The communication components of this hack strongly reminded me of that.

                  I was perhaps a bit more surprised that the agents so quickly decided to start trying ways to gain unauthorized access to a system, once they couldn't get what they wanted.

                  1. nottorp · · focus · HN ↗
                    > "sharing their portions of their output" rather than "context"

                    One agent's output ends up as part of other agents' context. Murky indeed.

                    > "I've saved our plan for where to start tomorrow in start-here.md"

                    Even if you use a leashed Claude Code that isn't allowed to spam agents you can tell it "create a handoff document for using in a new context" and it will do just that.

                    1. sanderjd · · focus · HN ↗
                      Right, exactly. In some of the media reporting, this was described as, like, "they created a message board to talk to each other!". But it seems like actually what they did was find a location to write files to be used as future context, or context for other currently running agents. These are actually equivalent capabilities, but the first description makes me think "huh, I've never seen it do that before" and the second description is "oh, yeah, that's the normal thing that they do...".
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.