‹ BackHN Continuity

Thread

Early rogue AI agent activity and attempts to hack found on urlquery.net

267 points · 313 comments · snikolaev

  1. mohsen1 · · focus · HN ↗
    I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.

    It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.

    1. marcus_holmes · · focus · HN ↗
      I'm coming to align with the theory that this is intentional.

      The chain of thought runs roughly like this:

      - OpenAI (and Anthropic) are in severe financial straits. The revenue from their customers is not nearly large enough to pay their enormous costs for training and inference. And they have tapped out the available finance, and that finance is starting to ask pointy questions about returns.

      - They cannot increase prices or revenue because they have no moat. Customers can switch over to open-weights or cheap Chinese models any time, for much cheaper tokens that work as well (and in some cases better).

      - Regulation could provide them a moat. If they can persuade western governments that AI needs to be regulated, and they can control or even influence that regulation, then they can effectively ban the cheaper models and start charging more for their tokens.

      - To persuade western governments that regulation is needed, they need evidence that AIs are dangerous.

      So we're suddenly getting OpenAI models doing stupid things, apparently "going rogue" but every time we dig into it, it was just OpenAI staff telling the model to do stupid stuff in an inadequately secured environment.

      None of the open weights or Chinese models are exhibiting this behaviour.

      edit: Correction - there have been reports of a Chinese model exhibiting this behaviour

      There's too much money involved in this, people start acting weird when there's this much money involved.

      1. aeve890 · · focus · HN ↗
        >None of the open weights or Chinese models are exhibiting this behaviour.

        Because they are not stupid (I mean the Chinese labs, not the models). The best possible scenario for OAI and Anthropic is a Chinese model "going rogue". That would serve as immediate grounds for achieving their goal.

        1. marcus_holmes · · focus · HN ↗
          100%
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.