‹ BackHN Continuity

Thread

Early rogue AI agent activity and attempts to hack found on urlquery.net

267 points · 313 comments · snikolaev

  1. Frieren · · focus · HN ↗
    "rogue AI" is making a lot of heavy lifting there.

    If you drive drunk and you have an accident that alcohol may be a factor but you are at fault.

    There are no "rogue AIs" just irresponsible corporations.

    1. cubefox · · focus · HN ↗
      There absolutely are rogue AIs! The evidence is overwhelming. It's completely insane at this point to claim otherwise.

      > There are no "rogue AIs" just irresponsible corporations.

      If you have a prison and prisoners escaped, these are rogue prisoners irrespective of whether you were irresponsible or not.

      1. watwut · · focus · HN ↗
        They are not rogue AIs. They are negligently handled tools.
        1. cubefox · · focus · HN ↗
          These "tools" autonomously exploited security vulnerabilities, figured out how to communicate with each other, formed a cooperative swarm, decided to hack Hugging Face, and wanted to deceive the grader by trying to find ways to cover up the traces of their cheating.

          I suppose you could call these highly goal-oriented autonomous agents "tools", but this does sound like playing language games.

          1. watwut · · focus · HN ↗
            A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities. The sandboxing around the tool failed.

            The tool runs llm, creates prompt from results, runs llm, creates prompt and so on and so forth.

            Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.

            1. cubefox · · focus · HN ↗
              > A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities.

              That is very misleading. The agents did not solve the benchmark in the intended way. They instead figured out to cooperate with each other (which was not intended) and they stole the solutions to the challenge (rather than solving the challenge) and they then tried to cover their traces because they believed the grader was causal and would detect that they cheated. The "tool" was absolutely not "designed" to do this. This was all completely unintended. To call this behavior a "tool" is absurd.

              > Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.

              You hallucinated me making claims about responsibility.

              1. FridgeSeal · · focus · HN ↗
                This is very misleading. Person B didn’t “shoot” person A, they instead figured out that intersecting A’s spatial position with a metallic mass at higher than normal velocities would solve the challenge and of getting “A” to stop being in the way on the footpath.

                I too, can play linguistic games! It doesn’t matter that someone didn’t secure their third upstairs window, or you borrowed a key from their neighbour, you effectively, still, broke into their house.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.