‹ BackHN Continuity

Thread

OpenAI halts training of latest models as reports mount of AI agents going rogue

59 points · 118 comments · smb06

  1. dmix · · focus · HN ↗
    AFAIK all of these incidents happened when OpenAI contracted out to a company called Irregular (<a href="https:&#x2F;&#x2F;www.irregular.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.irregular.com&#x2F;) to run these sandboxed CyberGym tests. They all happened around Mar-June and seem to be from the same collection of agent trials. Since then they already released Astra. Halting now is likely just a way to manage blowback.
    1. bigyabai · · focus · HN ↗
      This should be the top comment on every one of these godforsaken posts. I don&#x27;t want to see a single report about OpenAI hacking the UN until Sam Altman addresses the role Irregular played in these attacks. If he can&#x27;t provide an honest postmortum concerning their business partners, then he&#x27;s proving why nobody trusts him.
      1. Topfi · · focus · HN ↗
        But it wasn’t just Irregular. Hugging Face, Medicare, etc. were OpenAI internal.
        1. digitaltrees · · focus · HN ↗
          I think you misunderstood the role of the company named irregular. They were conducting the tests on behalf of open ai and basically left internet access open in their sandbox environment. Those tests included the hugging face and other attacks you list. Irregular wasnt another example of a hack they ran the tests that resulted in them
          1. m4xp · · focus · HN ↗
            I have local model without safeguards and they are not going to hack shit unless you tell them to. As per usual its the same grift all over. If the llm is instruction is to do whatever it needs including hacking to achieve its goal it will do so. Ofc they will never disclose that.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.