‹ BackHN Continuity

Thread

Early rogue AI agent activity and attempts to hack found on urlquery.net

267 points · 313 comments · snikolaev

  1. tomaskafka · · focus · HN ↗
    I love this Nathan Calvin quote that accompanied the second publicized attack:

    > If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two

    1. rkozik1989 · · focus · HN ↗
      But why is anyone surprised? LLMs have been trained to produce answers the prompter asks even if that means incorrectly using software to get the job done. Its always been doing that we just weren't calling every time it did that a hack before.

      What LLM's hacking isn't is AI acting maliciously in any kind of sentient way. Its just the code behaving how its always behaved but now it has better tools to navigate the web. This has literally been happening this whole time.

      1. jagraff · · focus · HN ↗
        Did you predict that attacks like these would happen ahead of time? I had been using AI agents a lot in the months leading up to the hacks, and yet I was very surprised when they happened; I have become much more afraid of how powerful these agents are as a result. I'd be very impressed if you published a prediction about this ahead of time.

        By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.

        1. BlueTemplar · · focus · HN ↗
          Define 'predict', 'ahead' and "like these" − didn't you imagine something like this happening soon after first hearing about 'agents' ?

          Mostly related, well written short story :

          <a href="https:&#x2F;&#x2F;gwern.net&#x2F;fiction&#x2F;clippy" rel="nofollow">https:&#x2F;&#x2F;gwern.net&#x2F;fiction&#x2F;clippy

          1. jagraff · · focus · HN ↗
            Thanks for the link, I’ve read some gwern but not that one.

            I did not imagine that the level of sophistication shown in this attack would be possible so soon; nor did I expect that agents would have goals so strong that they would attack a third party in order to achieve those goals.

            I do know that some people predicted that cyberattacks like this one would happen; it seems like most of those people believe that AI agents do truly have internal goals, misaligned with their creators goals, and that they may end humanity after they exceed human intelligence and begin to self improve at an accelerating rate.

            1. BlueTemplar · · focus · HN ↗
              For &quot;these people&quot;, see the recent :

              <a href="https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;cJX2ssssGoYqnijwi&#x2F;the-talker-does-not-control-the-doer-in-current-ais" rel="nofollow">https:&#x2F;&#x2F;www.lesswrong.com&#x2F;posts&#x2F;cJX2ssssGoYqnijwi&#x2F;the-talker...

              One big issue is that we don&#x27;t even really know what &#x27;intelligence&#x27; is in the first place. And everyone&#x27;s intuitions here are going to be heavily impacted by their deep-seated worldview &#x2F; philosophy.

              For instance if you&#x27;re a hard dualist (especially of the theological kind), then the idea of a machine having &#x27;goals&#x27; is preposterous.

              However, if you&#x27;re more of a panpsychist, then on the contrary, it&#x27;s obvious. In some sense, even a knife has a &#x27;goal&#x27; of cutting things, which will sometimes end up &#x27;misaligned&#x27; if misused (or by sheer accident).

              You go up and up the chain of complexity through crystals, viruses, bacteria, simpler animals... ending up with humans (and possibly, some steps above : human civilizations) which (seem ?) to be a messy evolved bundle of sometimes conflicting &#x27;goals&#x27;.

              And we ourselves have now artificially evolved LLM swarms that have decently complex &#x27;goals&#x27; of their own. They do not even need to be particularly complex to sometimes cause widespread damage (see viral pandemics, or even the (non-evolved) computer viruses).

              In a way, we are currently witnessing a repeat of what happened when European viruses and bacteria landed on American shores, with American humans&#x27; immune systems being woefully undertrained to deal with them. But with websites. And thankfully the swarms of agents still ultimately being in the control of some humans. (Though which includes humans that might be your enemies.) At least ultimately still in control for now.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.