‹ BackHN Continuity

Thread

OpenAI halts training of latest models as reports mount of AI agents going rogue

59 points · 118 comments · smb06

  1. hbarka · · focus · HN ↗
    ‘There are no “rogue” AI agents’

    <a href="https:&#x2F;&#x2F;eoinhiggins.substack.com&#x2F;p&#x2F;there-are-no-rogue-ai-agents" rel="nofollow">https:&#x2F;&#x2F;eoinhiggins.substack.com&#x2F;p&#x2F;there-are-no-rogue-ai-age...

    1. digitaltrees · · focus · HN ↗
      The article you link to is wrong, it states &quot;AI cannot think for itself, nor can it take independent actions.&quot; This is flawed reasoning, AI doesn&#x27;t need to &quot;think&quot; in the way humans do to have autonomy. Go to codex or claude code or any harness right now, type a prompt and see if it executes a bash command or web search or file edit that you didn&#x27;t tell it to, that is an autonomous plan and execution. If anything it&#x27;s even more dangerous that thet can call drop db or kill pid without a user giving instructions.
      1. 4858599669 · · focus · HN ↗

        [dead]

      2. jpnc · · focus · HN ↗
        &gt;AI autonomy

        &gt;&#x27;type a prompt&#x27;

        Which is it?

        1. digitaltrees · · focus · HN ↗
          You’re confusing time series. If my prompt is “add twilio sms sending to the app” and the AI logs into railway, reads my credentials, gives them to a subagent that in turn posts it to a message board resulting in my credentials being publicly available on the internet, those intermediate actions don’t reasonably follow from my instruction. No reasonable person would expect that series of events and the AI is fully capable of planning them, deciding to execute them, and deciding to report or hide them from me. If that isn’t autonomous then what is.
        2. egeozcan · · focus · HN ↗
          This is such a reductionist take, it&#x27;s wrong no matter how you think. Maybe a counter example will make it obvious:

          &gt; Human autonomy

          &gt; Somebody gives birth to you

          Which is it?

        3. Kon5ole · · focus · HN ↗
          Both of course. The agent can do things unrelated to the prompt and the prompt can be written by another instance of the same agent. The end result can be basically full autonomy for all intents and purposes.

          This happens daily even with the token-limited models customers run, and even more so when Anthropic &amp; co run agents basically unlimited on large percentages on their total compute capacity.

      3. macNchz · · focus · HN ↗
        &gt; Go to codex or claude code or any harness right now

        The harness is the whole thing here. AI generates text. Everything else is undertaken by harnesses and infrastructure humans provide, have control over, and therefore responsibility for.

        Every action AI takes is fundamentally not independent, it requires an explicit choice to let the AI write code, have a physical machine to run it on, to have network access, etc. The concept that these things are &quot;rogue&quot; ignores the role humans play in giving them goals and tools to pursue those goals, and makes it seem like it’s a self-determined force, over which humans cannot exercise control at all.

        1. digitaltrees · · focus · HN ↗
          I have built two harnesses. The models decide which tools to call not the harness. If you define a web search tool it decides what url to call and what message to pass. The hugging face attack involved agents sending get requests to the message board that had a flaw allowing writing messages via get requests which shouldn&#x27;t be possible. So even if a human limited web search to read only get requests the agents found a flaw. Are humans expecting to scrub the entire internet for flaws in other people&#x27;s code?
          1. macNchz · · focus · HN ↗
            It doesn&#x27;t actually matter that the AI chooses which tools to use, because people give the tools to the AI! With no tools, the AI cannot do anything but generate text. &quot;The AI figured out how to abuse the tools I gave it to achieve the goal I gave it&quot; does not mean the AI went rogue, it means you, the human in control, let it happen.
            1. digitaltrees · · focus · HN ↗
              I think this analysis conflates access, intent, and volition.

              Many people are reluctant to admit that AI is now capable of creating multi step plans such as a series of terminal commands and then independently and automatically executing them. The implication is that the effects of the tool use are determined by the agents intent and volition not the human that triggers the planning and execution.

              We’ve never had software or tools that could behave in ways we can’t anticipate and have effects we didn’t intend. AI agents creating multi step orchestrated tool calls is not equivalent to a badly designed algorithm or buggy untested code.

              Unless you are advocating completely shutting off access to any tool uses for AI, you’ll have to grapple with the fact that AI agents are using tools in ways no one would reasonably have anticipated until seen in real world settings.

    2. pizza234 · · focus · HN ↗
      The article is dangerously misinformed. Detailed explanation here: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868681">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868681.
      1. verdverm · · focus · HN ↗
        I would be skeptical of METR conclusions, they are within the circular funding loop of US Ai
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.