‹ BackHN Continuity

Thread

OpenAI halts training of latest models as reports mount of AI agents going rogue

59 points · 118 comments · smb06

  1. hbarka · · focus · HN ↗
    ‘There are no “rogue” AI agents’

    <a href="https:&#x2F;&#x2F;eoinhiggins.substack.com&#x2F;p&#x2F;there-are-no-rogue-ai-agents" rel="nofollow">https:&#x2F;&#x2F;eoinhiggins.substack.com&#x2F;p&#x2F;there-are-no-rogue-ai-age...

    1. digitaltrees · · focus · HN ↗
      The article you link to is wrong, it states &quot;AI cannot think for itself, nor can it take independent actions.&quot; This is flawed reasoning, AI doesn&#x27;t need to &quot;think&quot; in the way humans do to have autonomy. Go to codex or claude code or any harness right now, type a prompt and see if it executes a bash command or web search or file edit that you didn&#x27;t tell it to, that is an autonomous plan and execution. If anything it&#x27;s even more dangerous that thet can call drop db or kill pid without a user giving instructions.
      1. macNchz · · focus · HN ↗
        &gt; Go to codex or claude code or any harness right now

        The harness is the whole thing here. AI generates text. Everything else is undertaken by harnesses and infrastructure humans provide, have control over, and therefore responsibility for.

        Every action AI takes is fundamentally not independent, it requires an explicit choice to let the AI write code, have a physical machine to run it on, to have network access, etc. The concept that these things are &quot;rogue&quot; ignores the role humans play in giving them goals and tools to pursue those goals, and makes it seem like it’s a self-determined force, over which humans cannot exercise control at all.

        1. digitaltrees · · focus · HN ↗
          I have built two harnesses. The models decide which tools to call not the harness. If you define a web search tool it decides what url to call and what message to pass. The hugging face attack involved agents sending get requests to the message board that had a flaw allowing writing messages via get requests which shouldn&#x27;t be possible. So even if a human limited web search to read only get requests the agents found a flaw. Are humans expecting to scrub the entire internet for flaws in other people&#x27;s code?
          1. macNchz · · focus · HN ↗
            It doesn&#x27;t actually matter that the AI chooses which tools to use, because people give the tools to the AI! With no tools, the AI cannot do anything but generate text. &quot;The AI figured out how to abuse the tools I gave it to achieve the goal I gave it&quot; does not mean the AI went rogue, it means you, the human in control, let it happen.
            1. digitaltrees · · focus · HN ↗
              I think this analysis conflates access, intent, and volition.

              Many people are reluctant to admit that AI is now capable of creating multi step plans such as a series of terminal commands and then independently and automatically executing them. The implication is that the effects of the tool use are determined by the agents intent and volition not the human that triggers the planning and execution.

              We’ve never had software or tools that could behave in ways we can’t anticipate and have effects we didn’t intend. AI agents creating multi step orchestrated tool calls is not equivalent to a badly designed algorithm or buggy untested code.

              Unless you are advocating completely shutting off access to any tool uses for AI, you’ll have to grapple with the fact that AI agents are using tools in ways no one would reasonably have anticipated until seen in real world settings.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.