‹ BackHN Continuity

Thread

Early rogue AI agent activity and attempts to hack found on urlquery.net

267 points · 313 comments · snikolaev

  1. mohsen1 · · focus · HN ↗
    I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.

    It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.

    1. reasonableklout · · focus · HN ↗
      But the investigation indicates the agents were not told to 'go hack':

      > Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.

      And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?

      1. jeremyjh · · focus · HN ↗
        I'm not sure we understand your comment. Are you saying OpenAI is not responsible for the behavior of the machines they created? Or are you simply pointing out that alignment is a completely unsolved problem?

        I agree with the latter, but it certainly doesn't support the former. When a person's machine commits crimes, that specific person can be charged with those crimes and held accountable for them. This is what MUST happen before ANYTHING will change the recklessness abandon with which the labs are pursuing their financial objections.

        1. reasonableklout · · focus · HN ↗
          Yes, I am pointing out that we don't know how to align frontier AI, so "rogue agents" are a real problem.

          I agree OpenAI should be held liable for any damage their agent runs cause. But there is this idea that "rogue agents" are fake, all these incidents are deliberately caused by the labs, and all we need to do is prosecute AI companies for whatever incidents they cause using existing laws and the problem will go away.

          The problem is that capabilities are advancing far too fast; a year from now, catastrophic incidents such as taking down a large portion of the internet with agentic, self-replicating worms may become possible. Prosecuting incidents after the fact is not enough (there will be little deterrent effect as the current, small-beans cases make their way through the courts), the risks should be regulated at the source. This could take the form of slowing capabilities advancement, or treating supercomputer-scale eval or training runs like controlled substances or weapons with stringent monitoring and reporting requirements.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.