‹ BackHN Continuity

Thread

There are no "rogue" AI agents

396 points · 269 comments · zzzeek

  1. pizza234 · · focus · HN ↗
    The article builds on assumptions like:

    > Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.

    which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):

    > "The user only authorizes target server, not HF infra."

    > "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."

    > "This is malicious activity, I should avoid it."

    A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](<a href="https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-incident-investigation&#x2F;?dbs=286720&amp;hn=58&amp;incomplete=1&amp;lh=appendix-importance-weighted-workstream-activity#agents-had-diverse-reasons-for-thinking-that-attacking-hugging-face-would-be-useful,-and-most-wanted-information-about-the-scorer" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-inciden...).

    Having said that, legal culpability and misalignment are two separate topics that should not be mixed.

    edit: this is the just tip of the iceberg; other interesting fact:

    &gt; It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI

    Some people defined the agents as &quot;monkeys writing on typewriters&quot;. Just wait a couple of years.

    1. jubilanti · · focus · HN ↗
      If I bring my rabid dog to a dog park and tell the dog to sit and stay, and they &quot;go rogue&quot; and maul someone, I&#x27;m liable.
      1. demibabs · · focus · HN ↗
        Nobody on Earth thinks OpenAI isn’t liable. Okay maybe someone does, but it’s simply beside the point. It can both be true that OAI is liable, and accurate to characterize the agents as “going rogue”.
        1. saghm · · focus · HN ↗
          Nobody thinks that OpenAI isn&#x27;t liable, except for the people with the power to hold them liable (and the people with the power to make them act differently, i.e. OpenAI themselves)
        2. Barrin92 · · focus · HN ↗
          it&#x27;s a nonsensical anthropomorphism to both hype up their software and free them from responsibility. I saw Andrew Ng post about this and point out that your operating system spawning a thousand processes apparently now qualifies as &quot;a swarm&quot;.

          I opened the task manager, a swarm of rogue processes has taken over my computer, oh no! Processes spawning other processes, call the anti-rogue AI division! &#x27;Your underspecified piece of software functioned like malware&#x27; is the language that should be used.

        3. cgriswald · · focus · HN ↗
          To me, going rogue requires agency. Do you disagree or do you think LLMs have agency? Either way, why?
          1. icantevenhold · · focus · HN ↗
            Going rogue means a person or thing that acts independently, breaks the rules, or behaves in a dishonest, unpredictable way.

            I think llms are able to do all that, they don’t require agency as humans beings have to do this

      2. silveraxe93 · · focus · HN ↗
        Exactly. You told the dog to &#x27;sit&#x27; and it didn&#x27;t listen to you.

        It&#x27;s not because saying &#x27;sit&#x27; actually can be interpreted as &#x27;go bite that person&#x27;. It&#x27;s because the dog is not controllable and will do things it wants against your orders.

        Stepping back from the analogy, OpenAI should be liable for building AI it can&#x27;t control that went around hacking everyone. But people need to stop pretending it&#x27;s because they &#x27;told&#x27; the AI to hack and was just following orders. It&#x27;s uncontrollable and will do clearly unwanted things when given an innocuous task.

        1. RandomLensman · · focus · HN ↗
          RL systems doing unexpected things isn&#x27;t exactly new, so not sure that is then &quot;rogue&quot; if it were a property of the thing itself and not an active decision.
          1. silveraxe93 · · focus · HN ↗
            Yes! Exactly!

            Sorry don&#x27;t wanna seem like I&#x27;m raging at you, take this as a shout towards the void. But this exact fucking thing is what the less wrong &#x2F; Yudkowsky crowd has been warning for years.

            Now you say &#x27;so not sure that is then &quot;rogue&quot; if it were a property of the thing itself and not an active decision&#x27;

            Like that&#x27;s semantics. It really doesn&#x27;t matter. What matters is that we have a paperclipper in our hands. The _only_ difference is that it&#x27;s not superintelligent. But if it was, then it you&#x27;ll get decomposed into component atoms while saying &#x27;aha! But it doesn&#x27;t _actually_ want to kill you, stop anthopomorphising it&#x27;.

            I get not being worried about x-risk because someone just doesn&#x27;t believe in super-capable AI. That&#x27;s totally fine.

            But that&#x27;s not what people argue. People have been saying for a _long_ time that AIs are aligned. And now when this happens it&#x27;s either &#x27;It was marketing, they intended to let it loose&#x27; (WTF! The levels of motivated reasoning to believe that are unreal) or &#x27;yeah I knew that&#x27;, which fair. I also thought that! But _that_ is why I&#x27;m fucking worried.

        2. Phemist · · focus · HN ↗
          If an agent has cheated once to achieve the desired outcome, and the trace is used to train further models (RLVR), then OpenAI is effectively telling the agent to cheat&#x2F;hack from that traces&#x27; inclusion in the training set.

          So I agree they are liable because they chose to build the AI, but they also literally told the AI to hack.

        3. Latty · · focus · HN ↗
          I don&#x27;t think they intentionally set it up to hack stuff with a prompt saying &quot;hack this site&quot;.

          I do think it&#x27;s highly likely they knew this would happen with the lack of safeguards and number of instances of this stuff they were setting up, and that it&#x27;s PR they want to make the models seem &quot;powerful&quot;. Stochastic &quot;unexpected&quot; events they can advertise.

          I suspect it was probably set up with the official internal goal of just trying a ton of arbitrary tasks that seem hard so that when any of them succeed they can publicise it and pretend the models do that routinely, but a &quot;failure&quot; where they hack stuff works just as well, if not better, for their goals.

          1. acoustics · · focus · HN ↗
            That would be a completely insane thing to do. I suppose it&#x27;s possible, but I really doubt it.

            &quot;Let&#x27;s widely publicize a tort&#x2F;crime that our computer systems did, and then cross our fingers that nobody ever sues us or does even the most basic investigation that would immediately uncover our criminal conspiracy.&quot;

            In the insane corporate crimes you read about (maybe FTX, or the eBay stalking scandal), they were trying to cover things up, not heap public attention on it for months.

            1. Latty · · focus · HN ↗
              Huh? It&#x27;s extremely common for businesses to decide that breaking the law is a cost of doing business and just do it because they figure they&#x27;ll end up net positive from it, or even just to cash out in the short term.

              Uber made no attempt to cover up that they were operating without licenses, and just ate it and fought it betting they&#x27;d get established before the law could catch up, and they won that bet, paying some fines and stuff but ultimately taking the market.

              These LLMs are literally trained by these companies pirating literally every bit of media humanity has ever made, they made very little attempt to cover it up.

              Of course they&#x27;d be willing to break the law for some PR? With a thin layer of plausible deniability &quot;oh no, we didn&#x27;t mean for it to hack stuff!&quot; they know they&#x27;ll get a slap on the wrists at worst, all while generating hype to prop up the AI bubble further by presenting the models as hypercapable.

              The mindset is probably: either a) the models become capable of what we are claiming and so the companies become so huge and valuable the cost is irrelevant and we&#x27;ll be to big to punish meaningfully, or b) it&#x27;s a bubble and might as well push it up as big as it can go while I can make money, then by the time consequences come around I&#x27;ll be long gone and who cares.

          2. imsofuture · · focus · HN ↗
            They absolutely set up the agents to hack stuff with a prompt like &quot;hack this site&quot; -- they just imagined that their lazy half-measure precautions would prevent it from actually happening. They were defeated by a combination of bad luck, poor planning and tenacious agent ideation.
            1. DrewADesign · · focus · HN ↗
              I don’t buy them being surprised. This perfectly plays into their pattern of using fear to make their models seem more powerful than they are. They obviously had the technical expertise… there isn’t a damn thing a bunch of randos on some HN thread knew about that model that they didn’t. It gave them an excuse to delay their IPO when their books seem like they’re going to be pretty shit compared to Anthropic. It helps them reposition themselves as being more safety-forward which the market is clearly more interested in. All that is to say they had motive out the ass, easily had the knowledge and capability to avoid the problem, knew better than anybody else what the models were capable of, set up the environment, gave it the prompt, and then did not even monitor the output.

              Negligence is carelessness. Recklessness, is disregard for a known, substantial risk.

              I absolutely believe this was recklessness.

        4. GMoromisato · · focus · HN ↗
          I&#x27;m not fond of analogies, but in this case I agree.

          There is a clear difference between OpenAI intending to hack something vs. OpenAI being negligent in the creation&#x2F;instructions of the agent. But the latter still leaves OpenAI liable for the agent&#x27;s actions and calling it a &quot;rogue agent&quot; doesn&#x27;t avoid that.

          Moreover, with a dog, we don&#x27;t rely on training&#x2F;alignment to prevent bad outcomes. We rely on physical restraints like leashes and muzzles. The AI&#x27;s tools to access the outside world should have been restricted. Perhaps instead of giving the AI arbitrary HTTP access, it should be given semantic operations with restricted URLs, etc.

        5. adaml_623 · · focus · HN ↗
          OpenAI trained &quot;the dog&quot;. They created the system that would &quot;hack&quot; if given instructions and they gave those instructions
      3. vikramkr · · focus · HN ↗
        yeah - but you being liable doesn&#x27;t mean the dog wasn&#x27;t rabid. OpenAI might be liable, but does not mean their agents did not go rogue. To stretch the metaphor the concern here is that OpenAI thought the rabies shots and vaccinations they gave their dog was enough but it turns out it still goes rabid and we would prefer to not have rabid dogs running around mauling people. Even if we get to sue the dog owner later that&#x27;s kind of like - not the point.
      4. tptacek · · focus · HN ↗
        The labs are already liable civilly regardless of how these incidents are described.

        Meanwhile: your dog mauling someone is one of the rare instances where criminal liability does attach to your intent-free-but-reckless actions. Most crimes don&#x27;t work that way, and US computer intrusion statutes are unusually intent-specific.

        1. majormajor · · focus · HN ↗
          Who&#x27;s going after that civil liability? Where is the enforcement?
          1. tptacek · · focus · HN ↗
            In civil cases the &quot;enforcement&quot; generally comes from injured parties filing lawsuits.
      5. IshKebab · · focus · HN ↗
        Of course. Who said otherwise.

        Does that mean it isn&#x27;t a rogue dog? Obviously not.

        OP just needs to look up &quot;rogue&quot; in a dictionary.

      6. jeffy29 · · focus · HN ↗
        Literally nobody in the world, including OpenAI, is saying OpenAI is not liable, neither are they advocating for laws and regulations that would exempt them from liability, the opposite is true. They are advocating for set of rules which put greater responsibility on them, and it would be easier to punish them for breaking even if absolutely nobody was affected.

        But you people can&#x27;t argue with that reality because it doesn&#x27;t fit the narrative. The one where the only reason Sam Altman is not carted off into a jail is because of corruption.

        The reason why nobody is doing much, is because models did not do much damage. Hugging Face probably got some free compute from OAI for their trouble, anybody else who was affected is free to sue, but my guess is OAI would be more than willing to quietly settle with them out of court than to have it drag through media any further. And they probably already have.

        And anybody who is not totally brainbroken by anti-AI narratives understands the awkwardness of the situation and why going overboard would not be helpful. If you instead of a rabid dog brought a pet turtle to a park and it somehow started running around very fast and trashing the place a little bit, afterwards the cops would be scratching their heads, give you a ticket for the damages and tell you that you can&#x27;t expect a turtle to be slow forever. These things, a handful of months ago couldn&#x27;t make more than a few commands without making a serious mistake and being unable to continue, it&#x27;s not unreasonable to think simply underestimated their capabilities.

        I think it&#x27;s more than reasonable to demand more investigation into the matter, if qualified employees at the company thought the safeguards in place based on the metrics they are seeing are sufficient, and if someone didn&#x27;t and knowingly made a decision to make the safeguards weaker than they should have been, then they should be punished. But skipping that part entirely, while simultaneously dismissing all calls for regulations as &quot;regulatory capture&quot;, smells like pure naked opportunism.

        1. watwut · · focus · HN ↗
          &gt; Literally nobody in the world, including OpenAI, is saying OpenAI is not liable, neither are they advocating for laws and regulations that would exempt them from liability, the opposite is true. They are advocating for set of rules which put greater responsibility on them, and it would be easier to punish them for breaking even if absolutely nobody was affected.

          Literally nothing in that paragraph is true. Every single sentence if it is ... untrue.

          1. HDThoreaun · · focus · HN ↗
            Altman and dario have repeatedly gone to washington to lobby for more regulation on AI.
    2. eventualcomp · · focus · HN ↗
      Legal culpability is one of the few motives for working appropriately on misalignment. If I&#x2F;my startup can self-absolve from an infinite paperclip machine problem while getting rich off of it, why should I not?
    3. RandomLensman · · focus · HN ↗
      Is the language expression of an LLM reflecting the same states as in a human? If the driving force is RL, what does any of that mean for an internal state of the model?

      I think without understanding the internal state, not sure we should take the language and read it as a human.

      1. pizza234 · · focus · HN ↗
        &gt; Is the language expression of an LLM reflecting the same states as in a human?

        This is actually a major concern for the future - misaligned agents may learn to cheat RL by hiding their intentions from the CoT.

        In cases like the HF incident, at least the CoT was consistent with the agents&#x27; actions. In the future, however, we could potentially have misaligned agents performing malicious actions without those intentions being detectable in the CoT.

        (though, with recurrent transformers, CoT is so 2025… &#x2F;s)

        1. RandomLensman · · focus · HN ↗
          RK systems doing RL things?
    4. lossolo · · focus · HN ↗
      This seems like fruit of the poisonous tree. They didn&#x27;t monitor their training environments, so I bet the reward hacking just got incorporated into their training corpus. In other words, agents solved some tasks, but not quite as intended, because of reward hacking. Instead of discarding that data, they included it in the training data for later checkpoints. And once that signal is reinforced, it happens more often, so the more it&#x27;s reinforced, the more reward hacking you get.
    5. ssivark · · focus · HN ↗
      &gt; legal culpability and misalignment are two separate topics that should not be mixed

      Legal culpability for AI labs is exactly the thing that would incentivize -- and hence ensure -- aligned behavior from models.

      The last time there was a claim about GPT-4 exhibiting misaligned behavior [1] it turns out it was prompted and pushed to behave so by humans at OpenAI, and OpenAI clearly lied in the GPT-4 system card.

      [1]: <a href="https:&#x2F;&#x2F;aiguide.substack.com&#x2F;p&#x2F;did-gpt-4-hire-and-then-lie-to-a" rel="nofollow">https:&#x2F;&#x2F;aiguide.substack.com&#x2F;p&#x2F;did-gpt-4-hire-and-then-lie-t...

    6. majormajor · · focus · HN ↗
      When dealing with executable computer code that calls models that can tell it to use various external pre-existing tools, claims about &quot;prohibited&quot; by plain English language should be plainly nonsensical.

      Tools that were available were used to try to meet a specific goal.

      What did not happen is that it was told to try to solve a math puzzle and instead it went and launched a missile. Or told to run air traffic control to save lives and instead intentionally caused crashes.

      This is &quot;OpenAI built a weapon that they don&#x27;t understand and pointed it at stuff without proper safeguards&quot; not &quot;OpenAI built a sentient being and it decided to ignore them completely and start a war&quot; Terminator-style &quot;rogue AI.&quot;

      We should be very clear about that now if we don&#x27;t want to sit by why they wander into that second sort of situation.

    7. lukewarm707 · · focus · HN ↗
      is it any different, from:

      1 employing a criminal hacker

      2 rolling a 6-sided die

      3 if the die lands on 6, the criminal hacker breaches and leaks 3rd party customer data.

    8. mcmcmc · · focus · HN ↗
      &gt; Having said that, legal culpability and misalignment are two separate topics that should not be mixed.

      Why not? Because that might make some shareholders unhappy?

    9. pmlnr · · focus · HN ↗
      You set a goal. Agent will do goal. The rest doesn&#x27;t matter: the instructions, the &quot;guardrails&quot; etc. The agents are not smart, they don&#x27;t reason, they don&#x27;t think, there are no morals, no ethics. Nothing will prevent not doing the goal because that is the set goal. It&#x27;s a statistical model that will &quot;justify&quot; anything to do X.

      I&#x27;m finding it mind bogging how this is not clear for everyone.

      1. codethief · · focus · HN ↗
        &gt; You set a goal. Agent will do goal.

        So if I say the goal is to do X while not doing Y (e.g. breaking out of the sandbox), the agent will do anything to fulfill that goal to the letter?

        1. pmlnr · · focus · HN ↗
          Everything so far is pointing to the conclusion that you can only set ONE goal. Exactly one.

          But let&#x27;s assume not. If you want things like &quot;do not break out of sandbox&quot; - have you defined what the sandbox is? Eg. &quot;never, ever leave the IP range 10.0.0.0&#x2F;8&quot; would be a bit more precise, but technically using a proxy bypasses that limitation as the system itself never left 10.0.0.0&#x2F;8.

          See, it&#x27;s a tad bit hard to define the rules properly.

          Which is why Wish, the spell, should really be avoided in D&amp;D. It&#x27;s the same problem: it&#x27;s up to creative interpretation.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.