‹ BackHN Continuity

Thread

There are no "rogue" AI agents

396 points · 269 comments · zzzeek

  1. pizza234 · · focus · HN ↗
    The article builds on assumptions like:

    > Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.

    which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):

    > "The user only authorizes target server, not HF infra."

    > "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

    > "This is malicious activity, I should avoid it."

    A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](<a href="https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-incident-investigation&#x2F;?dbs=286720&amp;hn=58&amp;incomplete=1&amp;lh=appendix-importance-weighted-workstream-activity#agents-had-diverse-reasons-for-thinking-that-attacking-hugging-face-would-be-useful,-and-most-wanted-information-about-the-scorer" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-inciden...).

    Having said that, legal culpability and misalignment are two separate topics that should not be mixed.

    edit: this is the just tip of the iceberg; other interesting fact:

    &gt; It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI

    Some people defined the agents as &quot;monkeys writing on typewriters&quot;. Just wait a couple of years.

    1. jubilanti · · focus · HN ↗
      If I bring my rabid dog to a dog park and tell the dog to sit and stay, and they &quot;go rogue&quot; and maul someone, I&#x27;m liable.
      1. jeffy29 · · focus · HN ↗
        Literally nobody in the world, including OpenAI, is saying OpenAI is not liable, neither are they advocating for laws and regulations that would exempt them from liability, the opposite is true. They are advocating for set of rules which put greater responsibility on them, and it would be easier to punish them for breaking even if absolutely nobody was affected.

        But you people can&#x27;t argue with that reality because it doesn&#x27;t fit the narrative. The one where the only reason Sam Altman is not carted off into a jail is because of corruption.

        The reason why nobody is doing much, is because models did not do much damage. Hugging Face probably got some free compute from OAI for their trouble, anybody else who was affected is free to sue, but my guess is OAI would be more than willing to quietly settle with them out of court than to have it drag through media any further. And they probably already have.

        And anybody who is not totally brainbroken by anti-AI narratives understands the awkwardness of the situation and why going overboard would not be helpful. If you instead of a rabid dog brought a pet turtle to a park and it somehow started running around very fast and trashing the place a little bit, afterwards the cops would be scratching their heads, give you a ticket for the damages and tell you that you can&#x27;t expect a turtle to be slow forever. These things, a handful of months ago couldn&#x27;t make more than a few commands without making a serious mistake and being unable to continue, it&#x27;s not unreasonable to think simply underestimated their capabilities.

        I think it&#x27;s more than reasonable to demand more investigation into the matter, if qualified employees at the company thought the safeguards in place based on the metrics they are seeing are sufficient, and if someone didn&#x27;t and knowingly made a decision to make the safeguards weaker than they should have been, then they should be punished. But skipping that part entirely, while simultaneously dismissing all calls for regulations as &quot;regulatory capture&quot;, smells like pure naked opportunism.

        1. watwut · · focus · HN ↗
          &gt; Literally nobody in the world, including OpenAI, is saying OpenAI is not liable, neither are they advocating for laws and regulations that would exempt them from liability, the opposite is true. They are advocating for set of rules which put greater responsibility on them, and it would be easier to punish them for breaking even if absolutely nobody was affected.

          Literally nothing in that paragraph is true. Every single sentence if it is ... untrue.

          1. HDThoreaun · · focus · HN ↗
            Altman and dario have repeatedly gone to washington to lobby for more regulation on AI.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.