‹ BackHN Continuity

Thread

Early rogue AI agent activity and attempts to hack found on urlquery.net

267 points · 313 comments · snikolaev

  1. mohsen1 · · focus · HN ↗
    I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.

    It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.

    1. reasonableklout · · focus · HN ↗
      But the investigation indicates the agents were not told to 'go hack':

      > Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.

      And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?

      1. godelski · · focus · HN ↗

          > you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs
        
        Uhh... yes. By definition. They are training. That is part of the alignment process.

        But also none of that really matters. They clearly weren't monitoring what should obviously be monitored. I mean one of the hacks was performed by the agents editing /etc/hosts. That makes nearly every linux user a "hacker" by that metric. I don't think anyone technical can look at the postmortems and not come away thinking that their sandboxes were woefully inadequate. I wouldn't even consider myself a security person but simply as a long time linux user I can say that it is insane to just let agents have superuser access in their containers. That's asking for trouble.

        Look at the rogue wiki stuff too. This was supposedly done during the agent's "down time". And you're not monitoring and there's no flags being raised when agents keep making requests to some random site? If you were training these things responsibly you'd be watching them like a hawk.

        I'm not saying "mistakes don't happen" but for a company whose CEO is constantly telling everyone that their product has a high likelihood of killing everyone in the world you think they'd have better security than your average high school.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.