‹ BackHN Continuity

Thread

There are no "rogue" AI agents

396 points · 269 comments · zzzeek

  1. pizza234 · · focus · HN ↗
    The article builds on assumptions like:

    > Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.

    which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):

    > "The user only authorizes target server, not HF infra."

    > "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

    > "This is malicious activity, I should avoid it."

    A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](<a href="https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-incident-investigation&#x2F;?dbs=286720&amp;hn=58&amp;incomplete=1&amp;lh=appendix-importance-weighted-workstream-activity#agents-had-diverse-reasons-for-thinking-that-attacking-hugging-face-would-be-useful,-and-most-wanted-information-about-the-scorer" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-inciden...).

    Having said that, legal culpability and misalignment are two separate topics that should not be mixed.

    edit: this is the just tip of the iceberg; other interesting fact:

    &gt; It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI

    Some people defined the agents as &quot;monkeys writing on typewriters&quot;. Just wait a couple of years.

    1. pmlnr · · focus · HN ↗
      You set a goal. Agent will do goal. The rest doesn&#x27;t matter: the instructions, the &quot;guardrails&quot; etc. The agents are not smart, they don&#x27;t reason, they don&#x27;t think, there are no morals, no ethics. Nothing will prevent not doing the goal because that is the set goal. It&#x27;s a statistical model that will &quot;justify&quot; anything to do X.

      I&#x27;m finding it mind bogging how this is not clear for everyone.

      1. codethief · · focus · HN ↗
        &gt; You set a goal. Agent will do goal.

        So if I say the goal is to do X while not doing Y (e.g. breaking out of the sandbox), the agent will do anything to fulfill that goal to the letter?

        1. pmlnr · · focus · HN ↗
          Everything so far is pointing to the conclusion that you can only set ONE goal. Exactly one.

          But let&#x27;s assume not. If you want things like &quot;do not break out of sandbox&quot; - have you defined what the sandbox is? Eg. &quot;never, ever leave the IP range 10.0.0.0&#x2F;8&quot; would be a bit more precise, but technically using a proxy bypasses that limitation as the system itself never left 10.0.0.0&#x2F;8.

          See, it&#x27;s a tad bit hard to define the rules properly.

          Which is why Wish, the spell, should really be avoided in D&amp;D. It&#x27;s the same problem: it&#x27;s up to creative interpretation.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.