‹ BackHN Continuity

Thread

Early rogue AI agent activity and attempts to hack found on urlquery.net

267 points · 313 comments · snikolaev

  1. mohsen1 · · focus · HN ↗
    I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.

    It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.

    1. reasonableklout · · focus · HN ↗
      But the investigation indicates the agents were not told to 'go hack':

      > Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.

      And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?

      1. bastawhiz · · focus · HN ↗
        Nothing you're bringing up matters. OpenAI is the creator and operator. They're legally culpable for the consequences of the machine they made. The model is a machine: even if it could be demonstrated that the model reasoned its way into criminal behavior completely independently of OpenAI staff, that doesn't change anything.

        If I run a biology lab and engineer a terrible virus, it gets out, and a global pandemic ensues, I don't get to shrug and say "well we told it not to infect people". It's my fault for failing to mitigate the risks of my work.

        1. pixl97 · · focus · HN ↗
          > it gets out, and a global pandemic ensues

          I mean, yea, you should be punished. The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.

          Worse, the rate of technological growth is putting the capabilities to engineer viruses in the hands of people that may otherwise be suicidal. You can't punish them after they already won (in the sense of reaching their goals).

          While, yes, OAI should absolutely be punished, the future is majorly screwed as our power scaling laws are increasing much faster than our ability not to be stupid.

          1. lovich · · focus · HN ↗
            > The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.

            Sounds like your corporation should be dismantled then. But doing more than a fine in the millions is obviously not possible

            1. reasonableklout · · focus · HN ↗
              I would love for the US to bring back the corporate death penalty but seeing as it hasn't been applied in >100 years here, we should start with regulating the frontier labs, including by slowing down capabilities advancement if we do not know how to align the resulting models.
          2. byzantinegene · · focus · HN ↗
            you punish them hard enough their investors will have cold feet to keep the business going.
          3. bastawhiz · · focus · HN ↗
            > The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.

            That's never been the point. Anyone involved in a double homicide can never be punished to the same degree that they harmed their victims because they can't be put to death twice. The greater purpose of the justice system is to take offenders out of society and deter others from committing the same crimes. Punishment is gratifying but ultimately doesn't change anything.

            If OAI employees believed their company would be dismantled and their equity would become worthless, I suspect they'd be a whole lot more careful.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.