‹ BackHN Continuity

Thread

There are no "rogue" AI agents

396 points · 269 comments · zzzeek

  1. silverFork · · focus · HN ↗
    From what I understand, in one case, they had physically disconnected the sandbox from internet and asked it to do something and it had used connections through (import routines) that they had allowed, to pseudo escape the sandbox. Yes it wasn't obviously trying escape the sandbox but it escaped it because it doesn't understand the boundaries and neither do most humans other than the ones that provided the instructions that it had used. So it wasn't a rogue attempt but the fact that boundaries may be not be that easy to set despite what people think.
    1. verdverm · · focus · HN ↗
      You mean the proxy to package registries from one of the early incidents?

      I have not heard about any instances where physical disconnect has happened, would appreciate any links to update my priors

      other non hacking cases of negligence include suicide and school shootings, which I have heard they were aware of and monitoring, but did not contact authorities

      1. silverFork · · focus · HN ↗
        There are references about escaping offline sandbox.. I dont know about shootings!!

        <a href="https:&#x2F;&#x2F;www.primeintellect.ai&#x2F;blog&#x2F;universal-offline-sandbox-escape" rel="nofollow">https:&#x2F;&#x2F;www.primeintellect.ai&#x2F;blog&#x2F;universal-offline-sandbox...

        1. verdverm · · focus · HN ↗
          <a href="https:&#x2F;&#x2F;www.theguardian.com&#x2F;technology&#x2F;2026&#x2F;sep&#x2F;22&#x2F;british-columbia-sues-openai-sam-altman-tumbler-ridge-school-shooting" rel="nofollow">https:&#x2F;&#x2F;www.theguardian.com&#x2F;technology&#x2F;2026&#x2F;sep&#x2F;22&#x2F;british-c...

          re sandbox, I mean with actual OAI incidents, not theoretical

          one can mirror dependencies internally, rather than putting a simple proxy in place, I&#x27;ve built auth a thing, 100 lines of stdlib only Go and scripts for the mirroring process, our rationale was reliability b&#x2F;c upstream providers go down, and also only allowing approved images and packages, so devs cannot bring in random stuff

          1. lkjdsklf · · focus · HN ↗
            This has been common place for at least 20 years at this point.

            We did it as my very first startup and we were stupid children back then.

            Kind of telling that OpenAI didn’t.

        2. hn8726 · · focus · HN ↗
          If you read the article it states clearly that the sandbox wasn&#x27;t offline though? There were API calls to certain endpoint(s) allowed, and the model simply used that endpoint&#x27;s feature to query data from the internet
    2. kylebyte · · focus · HN ↗
      From that it sounds to me like the sandbox wasn&#x27;t physically disconnected from the internet.
    3. zugi · · focus · HN ↗
      &gt; physically disconnected the sandbox from internet ... used connections ... that they had allowed.

      That&#x27;s not &quot;physically disconnected the internet&quot;, that&#x27;s &quot;disabled some connections but enabled others.&quot;

      So the agent found and used the non-blocked connections.

      1. silverFork · · focus · HN ↗
        For the Ai code to execute, it needs the import functions... So that firewall between executing the code vs processing doesn&#x27;t really work.
    4. wat10000 · · focus · HN ↗
      If it was physically disconnected from the internet then it wouldn’t have been able to escape.

      This is so easy to do. Get a computer without wireless stuff. Don’t plug it into a network. If it needs access to other computers, make sure none of them have wireless stuff and make sure none of them have access to an internet connection. No matter how smart your AI is, it won’t be able to escape this.

      This clearly is not what they did.

    5. rafaelasor · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.