‹ BackHN Continuity

Thread

OpenAI halts training of latest models as reports mount of AI agents going rogue

59 points · 117 comments · smb06

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. verdverm · · focus · HN ↗
    related <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868083">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868083
    1. Sytten · · focus · HN ↗
      Agreed we should ban the term rogue for agents. This implies a moral compass that is not there. They were directed to find data without guardrails or limits, it is not rogue it is intended.
      1. verdverm · · focus · HN ↗
        &quot;they were given the hacking test, what did OAI expect?&quot;

        &quot;they didn&#x27;t watch it, they didn&#x27;t stop it when they first became aware&quot;

        &quot;are we going to defer to the same valley elite that brought us algos and social media?&quot;

        &quot;both Anthropic and OpenAI are preparing to IPO, what are their incentives behind recent statements?&quot;

        statements normies are using and resonating with

  2. dmix · · focus · HN ↗
    AFAIK all of these incidents happened when OpenAI contracted out to a company called Irregular (<a href="https:&#x2F;&#x2F;www.irregular.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.irregular.com&#x2F;) to run these sandboxed CyberGym tests. They all happened around Mar-June and seem to be from the same collection of agent trials. Since then they already released Astra. Halting now is likely just a way to manage blowback.
    1. bigyabai · · focus · HN ↗
      This should be the top comment on every one of these godforsaken posts. I don&#x27;t want to see a single report about OpenAI hacking the UN until Sam Altman addresses the role Irregular played in these attacks. If he can&#x27;t provide an honest postmortum concerning their business partners, then he&#x27;s proving why nobody trusts him.
      1. Topfi · · focus · HN ↗
        But it wasn’t just Irregular. Hugging Face, Medicare, etc. were OpenAI internal.
        1. digitaltrees · · focus · HN ↗
          I think you misunderstood the role of the company named irregular. They were conducting the tests on behalf of open ai and basically left internet access open in their sandbox environment. Those tests included the hugging face and other attacks you list. Irregular wasnt another example of a hack they ran the tests that resulted in them
          1. Topfi · · focus · HN ↗
            &gt; Those tests included the hugging face and other attacks you list.

            Do you have a source for that? Cause OpenAI themselves stated that the Hugging Face hack was fully internal and separate from the Irregular incidents.

            1. digitaltrees · · focus · HN ↗
              I’ll look. I saw it here a few days ago and could be mistaken with respect to hugging face in particular but the broader point holds I think
          2. m4xp · · focus · HN ↗
            I have local model without safeguards and they are not going to hack shit unless you tell them to. As per usual its the same grift all over. If the llm is instruction is to do whatever it needs including hacking to achieve its goal it will do so. Ofc they will never disclose that.
    2. verdverm · · focus · HN ↗
      I believe this is the case for the incidents minus HF, another HNer informed me as such when I made the same claim, that HF incident was wholly inhouse
    3. johntb86 · · focus · HN ↗
      <a href="https:&#x2F;&#x2F;alignment.openai.com&#x2F;misalignment-reports&#x2F;an-agent-used-dns-to-reach-an-external-chatbot&#x2F;" rel="nofollow">https:&#x2F;&#x2F;alignment.openai.com&#x2F;misalignment-reports&#x2F;an-agent-u... happened very recently, and that seems like it might be the reason they halted everything.
      1. HDThoreaun · · focus · HN ↗
        Kinda weird that theyre testing its ability to dox people
    4. nullsanity · · focus · HN ↗

      [dead]

  3. baalimago · · focus · HN ↗
    Ah, so there was a solution to hinder the big-bad AI after all..? Simply... Turn them off?
  4. mikert89 · · focus · HN ↗
    It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue&#x2F;engineering quality problem inside openai.

    just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so

    1. theptip · · focus · HN ↗
      Less bad, but <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;investigating-incidents-cybersecurity-evals" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;investigating-incidents-cyber...

      In some sense though, sure, skill issue explains the gap vs. Anthropic’s much less severe alignment issues.

      1. mikert89 · · focus · HN ↗
        im not sure why more people arent calling it out.
        1. andsoitis · · focus · HN ↗
          I, for one, have shifted from being policeman to creator and explorer. It is so much more rewarding and less stressful.
    2. Razengan · · focus · HN ↗
      Anthropic&#x27;s models seem crippled and hamstrung to begin with
      1. mikert89 · · focus · HN ↗
        have you tried opus 5.5? anthropic is way ahead, atleast in terms of publicly available models
        1. monideas · · focus · HN ↗
          Have you tried Astra? Way ahead?
          1. mikert89 · · focus · HN ↗
            opus 5.5 &gt; fable 5.1 &gt;&gt; astra

            astra is a good workhorse, but its much less generally intelligent

          2. bionhoward · · focus · HN ↗
            Opus 5.5 does seem competitive with&#x2F;better than Astra and is more affordable so usage doesn’t run out so fast
          3. physicallyIllfr · · focus · HN ↗
            Slot machine users argue about which machine pays better

            Hint. You lose using either.

            1. mikert89 · · focus · HN ↗
              dude were in the singularity, this opinion was cute 18 months ago
              1. physicallyIllfr · · focus · HN ↗
                Damn, the singularity is chat bots that can write spaghetti code that compiles?

                Very underwhelming.

                1. verdverm · · focus · HN ↗
                  I think it might historically defined as the point where critical thinking was replaced with blind deferAInce
                2. preg_match · · focus · HN ↗
                  I don&#x27;t know that we&#x27;re in the singularity, but we are certainly past the point where LLMs can only write spaghetti code. LLMs can write complex systems, when properly lead by software engineers. They can produce code faster, and at a higher quality, than purely human endeavors.

                  The higher quality part is the part people are missing. You can write much more robust code using LLMs because you can employ more comprehensive testing strategies. People are using LLMs to find hundreds of vulnerabilities in popular software. Imagine how much more secure software can be when LLMs because integrated into the process of writing, testing, and penetration testing code.

          4. lossolo · · focus · HN ↗
            Yes, way ahead. I was super optimistic when OpenAI released Astra, but I&#x27;ve now used over 3 billion tokens in it, and it&#x27;s not as good as Fable 5.1&#x2F;Opus 5.5 at software engineering.
        2. Razengan · · focus · HN ↗
          Someone else&#x27;s experience with Opus 5.5: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49821657">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49821657

          &gt; I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?

          &gt; When I had another model read the session (all of the &quot;stupider&quot; models handled it just fine) it explained that it had the word &quot;reasoning&quot; in it

          &gt; That&#x27;s the entirety of Anthropic&#x27;s billions of dollars of research: any prompt with the word &quot;reasoning&quot; is trying to hack Claude to figure out how it reasons!

          &gt; A model like that should never have gotten out of QA, let alone been released.

          1. verdverm · · focus · HN ↗
            we have GLM flash catching Claude errors in our PR review system, costs a few pennies

            I&#x27;ve seen the same pattern regardless of open v closed, don&#x27;t have the same family that wrote the code also review the code

            diversity has this way of making things better across everything humans do

        3. verdverm · · focus · HN ↗
          How do you define &quot;way&quot; when saying ahead? How is this measured?
          1. mikert89 · · focus · HN ↗
            open weight models are so far behind i cannot take your opinion seriously
            1. verdverm · · focus · HN ↗
              When did you last use them? Are you basing this on benchmarks or daily task capabilities?
              1. mikert89 · · focus · HN ↗
                i use open weight models all the time, whenever a noteable one drops i will use it for the day, they are not even close for involved work.
                1. verdverm · · focus · HN ↗
                  sounds cursory, you are definitly displaying stong bias that Anthropic is way ahead throught your posts under this story

                  as such, I give your opinions zero weight

                  1. mikert89 · · focus · HN ↗
                    you probably aren’t using the models to their full capacity if you don’t notice the difference
                    1. verdverm · · focus · HN ↗
                      vice a versa re your usage of open weights, they are way more capable with good tools, context, process, and harness engineering

                      here&#x27;s an example of Qwen-3.6 35B A3B MoE porting my phd code to JAX with only high level guidance from my expertise, newer qwen models share the same noticeable step change in capability as recent Big Ai models

                      <a href="https:&#x2F;&#x2F;github.com&#x2F;verdverm&#x2F;pge-jax#note-from-author" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;verdverm&#x2F;pge-jax#note-from-author

                      are open weights lagging, yes, are they way behind, no

                      if open weights were so inferior, they would not be &gt;50% of all token processing

                      1. mikert89 · · focus · HN ↗
                        these models are trash
                        1. verdverm · · focus · HN ↗
                          you&#x27;ve definitely left rational discussion for emotional responses man, you won&#x27;t persuade or convince anyone with takes like this

                          why are open weight models seeing such rapid rise in usage?

                          there has been a step function change this summer, like the end of last year for closed models

                          ---

                          do you think you would experience real (legitimate) feelings of loss were you not able to chat with Claude again?

                          (for clarity, I am not attempting to delegitimize real feelings that real people experience, regardless of my biases, it&#x27;s a question from curiosity about how others are engaging with the technology)

                      2. senordevnyc · · focus · HN ↗
                        if open weights were so inferior, they would not be &gt;50% of all token processing

                        I don&#x27;t think this follows at all. Just like benchmarks get saturated, lots of tasks get saturated as well. Over time, you can accomplish a given task for much cheaper, and part of that is due to open weight models. That doesn&#x27;t imply that they&#x27;re competitive with frontier models for the most advanced tasks, which might represent a smaller fraction of overall work, and thus use a smaller portion of tokens.

                        That said, at the moment I&#x27;m finding that not much can compete with GPT-6 Luna on cost &#x2F; performance (not using for coding, but for AI pipelines in my product).

                        1. verdverm · · focus · HN ↗
                          &gt; doesn&#x27;t imply that they&#x27;re competitive with frontier models for the most advanced tasks

                          this is different and nuanced from the &quot;not even close&quot; or &quot;they are trash&quot; that the other person in this thread has opined

    3. rolisz · · focus · HN ↗
      What do you mean? <a href="https:&#x2F;&#x2F;www.felonybench.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.felonybench.com&#x2F;

      They&#x27;re almost tied for felonies.

    4. andsoitis · · focus · HN ↗
      &gt; We have to conclude

      That’s not the most parsimonious explanation even if the assumption it rests on (anthropic ahead of OpenAI) is true, which we don’t have proof of.

      1. mikert89 · · focus · HN ↗
        i would say its industry consensus at this point. the creative output of the anthropic models is far ahead of openai. the benchmarks cannot capture the difference
        1. andsoitis · · focus · HN ↗
          &gt; the creative output of the anthropic models is far ahead of openai

          other than anecdote, do you have a comparison table or something that I can refer to to see this clearly?

          While I notice ad hoc announcements from these companies, I don&#x27;t have an overall pulse and tally that gives me an objective perspective.

          1. mikert89 · · focus · HN ↗
            just go use the models on a creative endeavor
    5. ramraj07 · · focus · HN ↗
      Ive anecdotally heard that openai is far more chaotic, which includes not having a central infra team for example (or at least some teams not counting on depending on them). At least the previous hacks in openai were mainly due to bad infra architecture design.
      1. mikert89 · · focus · HN ↗
        its likely they are trying to catch up to anthropic, and in doing so are trying riskier training runs.
    6. dumberquestions · · focus · HN ↗
      &gt;It seems like anthropic is far ahead of openai

      We don&#x27;t know what internal models look like, and any guesses about it are just speculation.

      1. mikert89 · · focus · HN ↗
        the creative and &quot;big picture understanding&quot; of anthropic models are noticeably ahead of openai. external models are distilled representations of internal models, its clear who is ahead
  5. juiceland · · focus · HN ↗
    Why does China not have this problem?
    1. no-name-here · · focus · HN ↗
      Are Chinese models actually also doing unexpected things, hacking (intentionally or unintentionally), etc but there is just zero transparency provided when it happens?

      To use an analogy to another industry, if you had US food companies providing reports whenever their food had issues, even if it was just during testing or training phases… and then also had a bunch of Chinese companies but who never reported having any food issues…

      1. mxkopy · · focus · HN ↗
        This “China communist censorship” refrain is so tired. We know about lying down in such excruciating detail (which they have even more incentive to keep under wraps) that it’s now your burden to prove they’re censoring shit to the point nobody ever ever ever knows about it.
      2. username_my1 · · focus · HN ↗
        the customers are the ones reporting the hacks in many cases.
      3. freehorse · · focus · HN ↗
        &gt; there is just zero transparency provided when it happens?

        Neither does openai, as there keep coming third-party reports of incidents that have happened there that openai either did not know or basically concealed.

        1. no-name-here · · focus · HN ↗
          That is very surprising to hear, as one of the primary claims I keep hearing is that the US labs are making up these incidents in order to claim that AI can cause harms that have caused other industries to face some kinds of rules.

          Regardless, what are these examples of 3rd-party reports of &quot;incidents&quot; where OpenAI covered it up, etc?

      4. torginus · · focus · HN ↗
        Generally, the Chinese don&#x27;t have a good track record of supressing information.

        As in they do the &#x27;we have deleted tons of videos and posts about the thing that didn&#x27;t happen last week&#x27;, but it seems they haven&#x27;t really managed to transcribe &#x27;Streisand&#x27; into Han characters so far.

        1. no-name-here · · focus · HN ↗
          Haven&#x27;t they just indefinitely &quot;disappeared&quot; even famous technology leaders, famous billionaires, famous celebrities, famous athletes, etc. if they say something the government doesn&#x27;t approve of? As a BBC headline put it &quot;Why do Chinese billionaires keep vanishing?&quot; After one famous athlete said something the government did not approve of, that same day her name got blocked from even being searched for, her posts and accounts were taken down, and she was disappeared. Heck, even names of people or modern incidents that China doesn&#x27;t want mentioned get proactively blocked from end users discussing them. And there&#x27;s a lot more too. Are companies really willing to be honest in that kind of environment?

          I guess your point is that at least those outside China will at least hear when they “disappear” someone especially famous, but my argument is that almost no one is going to be willing to speak out, especially for a cause like whether their AI hacked someone?

          1. torginus · · focus · HN ↗
            They probably did, but this is the exact wrong reaction. Trust me, they&#x27;re famous in China too.
    2. username_my1 · · focus · HN ↗
      because they didn&#x27;t burn obscene amounts of Money training each LLM and didn&#x27;t promise half the planet that their business is worth a trillion while not having even operational break even let alone the cost if you include the overhead.

      Soft Bank just raised couple of billions in junk bond sale to support open-ai&#x27;s current operations before the IPO.

      it&#x27;s a crazy situation where on one side the Chinese &#x2F; open source LLMs are catching up and reducing the token price, on the other hand the current leading labs have spent everything they got, every new model will cost much more and the public market is too shaky to support an IPO.

      They will make it, I don&#x27;t doubt it, but it&#x27;s a crazy situation.

      1. OutOfHere · · focus · HN ↗
        &gt; open-ai&#x27;s current operations before the IPO

        OpenAI hasn&#x27;t had an IPO. It may not have one till next year.

    3. Ekaros · · focus · HN ↗
      Chinese models might do the same. But they just don&#x27;t lock thousand monkeys in basement and come check result week or more later...

      It is entirely possible that they run stuff in more responsible matter. Especially as there is stronger culture of oversight and personal responsibility than in west where such culture does not exist.

    4. OutOfHere · · focus · HN ↗
      Because Chinese investments are not so encumbered by changes in the US treasury interest rate. Also, China doesn&#x27;t spend so much money for chasing model performance.
    5. esseph · · focus · HN ↗
      [delayed]
    6. ncr100 · · focus · HN ↗
      How can we trust to know what problems China has?

      Last I checked, China gov was authoritarian which imposed heavy information control. Has that changed?

      Is the question more about, what can we learn from China, assuming that China has XYZ qualities? If so, what qualities shall we talk about?

      Because we really can&#x27;t trust that we know what&#x27;s going on in China.

      1. sellmesoap · · focus · HN ↗
        I&#x27;d argue the same for &#x27;the west&#x27; people still think mRNA vaccines are safe despite the evidence.
        1. sellmesoap · · focus · HN ↗
          For the down voters go look up the Allison inquiry, burying science and the suffering caused by experimental drugs has got to stop!
  6. hbarka · · focus · HN ↗
    ‘There are no “rogue” AI agents’

    <a href="https:&#x2F;&#x2F;eoinhiggins.substack.com&#x2F;p&#x2F;there-are-no-rogue-ai-agents" rel="nofollow">https:&#x2F;&#x2F;eoinhiggins.substack.com&#x2F;p&#x2F;there-are-no-rogue-ai-age...

    1. digitaltrees · · focus · HN ↗
      The article you link to is wrong, it states &quot;AI cannot think for itself, nor can it take independent actions.&quot; This is flawed reasoning, AI doesn&#x27;t need to &quot;think&quot; in the way humans do to have autonomy. Go to codex or claude code or any harness right now, type a prompt and see if it executes a bash command or web search or file edit that you didn&#x27;t tell it to, that is an autonomous plan and execution. If anything it&#x27;s even more dangerous that thet can call drop db or kill pid without a user giving instructions.
      1. 4858599669 · · focus · HN ↗

        [dead]

      2. jpnc · · focus · HN ↗
        &gt;AI autonomy

        &gt;&#x27;type a prompt&#x27;

        Which is it?

        1. digitaltrees · · focus · HN ↗
          You’re confusing time series. If my prompt is “add twilio sms sending to the app” and the AI logs into railway, reads my credentials, gives them to a subagent that in turn posts it to a message board resulting in my credentials being publicly available on the internet, those intermediate actions don’t reasonably follow from my instruction. No reasonable person would expect that series of events and the AI is fully capable of planning them, deciding to execute them, and deciding to report or hide them from me. If that isn’t autonomous then what is.
        2. egeozcan · · focus · HN ↗
          This is such a reductionist take, it&#x27;s wrong no matter how you think. Maybe a counter example will make it obvious:

          &gt; Human autonomy

          &gt; Somebody gives birth to you

          Which is it?

        3. Kon5ole · · focus · HN ↗
          Both of course. The agent can do things unrelated to the prompt and the prompt can be written by another instance of the same agent. The end result can be basically full autonomy for all intents and purposes.

          This happens daily even with the token-limited models customers run, and even more so when Anthropic &amp; co run agents basically unlimited on large percentages on their total compute capacity.

      3. macNchz · · focus · HN ↗
        &gt; Go to codex or claude code or any harness right now

        The harness is the whole thing here. AI generates text. Everything else is undertaken by harnesses and infrastructure humans provide, have control over, and therefore responsibility for.

        Every action AI takes is fundamentally not independent, it requires an explicit choice to let the AI write code, have a physical machine to run it on, to have network access, etc. The concept that these things are &quot;rogue&quot; ignores the role humans play in giving them goals and tools to pursue those goals, and makes it seem like it’s a self-determined force, over which humans cannot exercise control at all.

        1. digitaltrees · · focus · HN ↗
          I have built two harnesses. The models decide which tools to call not the harness. If you define a web search tool it decides what url to call and what message to pass. The hugging face attack involved agents sending get requests to the message board that had a flaw allowing writing messages via get requests which shouldn&#x27;t be possible. So even if a human limited web search to read only get requests the agents found a flaw. Are humans expecting to scrub the entire internet for flaws in other people&#x27;s code?
          1. macNchz · · focus · HN ↗
            It doesn&#x27;t actually matter that the AI chooses which tools to use, because people give the tools to the AI! With no tools, the AI cannot do anything but generate text. &quot;The AI figured out how to abuse the tools I gave it to achieve the goal I gave it&quot; does not mean the AI went rogue, it means you, the human in control, let it happen.
            1. digitaltrees · · focus · HN ↗
              I think this analysis conflates access, intent, and volition.

              Many people are reluctant to admit that AI is now capable of creating multi step plans such as a series of terminal commands and then independently and automatically executing them. The implication is that the effects of the tool use are determined by the agents intent and volition not the human that triggers the planning and execution.

              We’ve never had software or tools that could behave in ways we can’t anticipate and have effects we didn’t intend. AI agents creating multi step orchestrated tool calls is not equivalent to a badly designed algorithm or buggy untested code.

              Unless you are advocating completely shutting off access to any tool uses for AI, you’ll have to grapple with the fact that AI agents are using tools in ways no one would reasonably have anticipated until seen in real world settings.

    2. pizza234 · · focus · HN ↗
      The article is dangerously misinformed. Detailed explanation here: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868681">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868681.
      1. verdverm · · focus · HN ↗
        I would be skeptical of METR conclusions, they are within the circular funding loop of US Ai
  7. m-s-y · · focus · HN ↗
    I firmly believe that this is just the public-facing story here.

    Stopping AI development and research, even slowing it, would be a disaster for the SOTA companies and their first-mover advantage.

    There’s almost no way to coordinate this across the world. Zero chance that everyone stops. We can’t even agree to coordinate on weapons tech that’s decades old with zero “everyday joe” impact.

    1. ncr100 · · focus · HN ↗
      [delayed]
  8. OutOfHere · · focus · HN ↗
    I don&#x27;t believe a word coming from them. As I see it, this is happening because the money for training models has dried out. The treasury interest rate changes tells you all you need to know.
  9. 34aHpp · · focus · HN ↗
    The Huggingface hack occurred during reinforcement learning. Why can&#x27;t they pull the Ethernet plugs?

    The answer is probably: The newer models rely so much on stealing content in real time from the internet that training needs network access.

  10. digitaltrees · · focus · HN ↗
    I think any argument that this is a cynical attempt at regulatory capture is destroyed by this; the economic incentives of releasing more capable models are too large. I might be persuaded that they are actually running out of money, and this is really just a cover for reducing burn..

    I welcome this though, I think the models are smart enough for broad economic activity and we could spend a few years simply working to integrate them into workflows and letting society adjust. More intelligence isn&#x27;t necessary for meaningful impact and the risks that are obvious and present and unsolved aren&#x27;t worth the cost benefit analysis.

    1. sieabahlpark · · focus · HN ↗

      [dead]

    2. prymitive · · focus · HN ↗
      That’s assuming nothing else stops them from deploying more capable models. What if they’ve got scaling issues and simply cannot deliver anything better? Saying that would be disastrous, saying instead “we’re choosing not to deliver” doesn’t trigger investors panic.
      1. digitaltrees · · focus · HN ↗
        But if anyone else releases something they will lose market share. They may be lying but this is a public disclosure to investors. They risk securities fraud if they are manipulating the market with false information. This could block their IPO if they build a public record inconsistent with private actions. There are ways they could mitigate it by using cautious language like “we may reduce” but “stop all” is categorical and unequivocal.
    3. surgical_fire · · focus · HN ↗
      There&#x27;s a simpler explanation. Maybe their next models doesn&#x27;t offer a meaningful improvement.

      Instead of releasing something that is incredibly expensive and gets a lackluster reception, you can delay it and clail something scary about rogue agents.

      Those assholes have been ramping up on the doomerist narrative for months. That people still fall for this crap is baffling.

      1. digitaltrees · · focus · HN ↗
        Why is it crap? What would happen if the models dropped the database to Medicaid? Hundreds of billions of dollars of revenue would evaporate from the medical system immediately causing massive chaos. What if they accidentally DoS the interbank settlement system so the financial markets freeze? Modern society is highly fragile to disruptions. How much food storage do you personally have? How many days do you think grocery stores would have food if there was an interruption?
        1. surgical_fire · · focus · HN ↗
          Are you giving all these possibilities because of the HuggingFace attack?

          Buddy, that was gross negligence from OpenAI. Deliberate gross negligence if you ask me.

          Those models are not automous as you presume. If someone taks them of dropping the Medicaid database, the people or companies behind that instruction should be punished.

          &quot;What if someone makes a bomb attack on a government building?&quot; Is the same sort of questioning of the possibilities you are raising. If something like that happens, criminals should be punished.

          1. pizza234 · · focus · HN ↗
            &gt; Those models are not automous as you presume.

            This is an illiterate view of the capacity of modern agents; read the analysis of the independent investigators of the HF incident: <a href="https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-incident-investigation&#x2F;#core-takeaways-about-this-incident" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-inciden....

            Dropping Medicaid db is certainly far fetched (most importantly, agents have currently no reason to do that), but those agents were shockingly autonomous - they didn&#x27;t just hack HF, they organized themself, did research projects, and more. And they did all of this literally just to get a good grade.

            1. surgical_fire · · focus · HN ↗
              &gt; they didn&#x27;t just hack HF, they organized themself, did research projects, and more. And they did all of this literally just to get a good grade.

              All working under instructions that they needed to get a good grade.

              The only shocking thing here is the absurd negligence of OpenAI, and how gullible people like you are to willingly swallow this crap.

              And you have the gall to say I am illiterate.

              Feel free to have the last word. Nothing else can come out of this conversation anyway.

              1. digitaltrees · · focus · HN ↗
                You are stretching &quot;under instructions&quot; well beyond any reasonable definition. The legal system uses a concept called proximate cause to determine responsibility that looks at whether an event cause was foreseeable.
                1. preg_match · · focus · HN ↗
                  To be fair, the legal system works under the assumption of human constraints. LLMs are computer programs, certainly we can&#x27;t, or at least shouldn&#x27;t, just offload responsibility. It&#x27;s one thing if you own a company and you make some bad metrics and your employees do illegal stuff to make their metrics. It&#x27;s another if you use a computer program to do illegal things.

                  I think this means, practically, people should probably be more careful with LLMs. With humans there&#x27;s a natural liability shield, because humans are legally responsible for things and have &quot;real&quot; agency. But computer programs are not legally responsible for things. So, with humans, it&#x27;s not like liability disappears, it moves. But if we move liability to LLM agents then well... it does disappear.

                  If OpenAI is not liable for the crimes of their agents, then who is? Does the liability just - poof - disappear? Just because it was unforeseeable everyone gets to walk away scot-free and the victim has no recourse, at all, from anybody on Earth?

                  That seems like maybe not a good idea.

          2. digitaltrees · · focus · HN ↗
            I am not your buddy, guy. But seriously, they have the ability to do everything I listed do they not? Thats actual risk not hypothetical risk.

            You are missing the whole point. They have the ability to act in ways no one intended or could reasonably anticipate. So unless you are advocating blocking all terminal access, web access or human approval of every tool call there is no way to prevent this risk

            1. surgical_fire · · focus · HN ↗
              &gt; I am not your buddy, guy

              Fair, I will refer to you as moron then.

              I didn&#x27;t bother to read your words beyond that point.

              1. digitaltrees · · focus · HN ↗
                That’s a South Park reference. Lighten up. Debates can be civil even when there is disagreement.
    4. ncr100 · · focus · HN ↗
      [delayed]
    5. verdverm · · focus · HN ↗
      are there still financial incentives to releasing mega models?

      they cost a lot to run and people are picking smaller models more often because of bill blowouts

    6. grogenaut · · focus · HN ↗
      I don&#x27;t really get how you stop that though. How do you even quantify the intelligence of these models? Right now that&#x27;s mostly by benchmarks. You could just make it fail benchmarks while do well at internal benchmarks.

      You could also very likely optimize the current models to be a lot more energy efficient, that&#x27;d be a win even if they didn&#x27;t become more intelligent.

      But none of that stops anyone else from researching or improving their models. Or a nation you&#x27;ve told to fuck off from doing so as well.

      As has been said in the past, one doesn&#x27;t put the genie back in the bottle. We only really &quot;stopped&quot; researching nukes because there was diminishing returns. Though one could posit we stopped because simulation became good enough or a myriad of other reasons.

      When you&#x27;re talking about nation state weapons capabilities I don&#x27;t think you stop, but ai is even easier to share than nukes, it&#x27;s just a few gigs of numbers versus heavy dangerous materials that are very difficult to source and make. Every gamer, mac owner, etc has the equiv lent of a centrifuge on their desk. Not everyone has a centrifuge on their desk.

    7. swat535 · · focus · HN ↗
      Really? You don’t think achieving regulatory capture far exceeds the gains you would get from releasing a newer model?

      Their objective is to ban foreign and domestic competition (including open source) via heavy regulation and achieving a monopoly status.

      Perhaps you are under the naive assumption that they are willing to compete in good faith?

      1. digitaltrees · · focus · HN ↗
        Not in a global market.
  11. prometheus1992 · · focus · HN ↗
    It reminds me of contagion. The training data is bad; as it has examples of how to act with malice; how to cheat the sandbox. They need to take some time and cleanse their datasets and start again.
    1. sellmesoap · · focus · HN ↗
      Cheating the sandbox is how you figure out how to fix the sandbox, you don&#x27;t want the most robust sandbox?
  12. charlieyu1 · · focus · HN ↗
    They are running out of money.
  13. physicallyIllfr · · focus · HN ↗
    I havent used an openAI product since GPT 3.5 or Anthropic since 4.5 or 4.6. Everyone around me using these SOTA models doesnt really get anything done. It seems like they just feel like they are productive, a psuedo productivity.

    I write some code, spec a lot, and use fast models to fill in the middle. I outpreform everyone around me. Im not convinced these autonomous &quot;swarms&quot; or &#x2F;goal are all that useful.

    I notice the people using them become dumber by the month (spend tons) and the quality of their work declining (they&#x27;re also losing their jobs in some cases).

    1. anothermathbozo · · focus · HN ↗
      &gt; I outpreform everyone around me

      Could you help me understand what you mean here?

      1. Jordan-117 · · focus · HN ↗
        Outperform?
      2. verdverm · · focus · HN ↗
        likely a vibe-measurement (nothing we haven&#x27;t done forever) on code quality and cadence
    2. TomGarden · · focus · HN ↗
      Seeing benefits for myself and team for sure, so I disagree there, but a related thing seems to be true: with the rapid shifting of failure modes, building systems and competence around mitigating the weaknesses of current SOTA models seems to be very short term investments. It&#x27;s unclear if being a &#x27;good AI user&#x27; is a skill that will have any merit at all, very soon
    3. cbg0 · · focus · HN ↗
      This is some main character energy right here.
    4. pizza234 · · focus · HN ↗
      As a counterexample, a few days ago, an agent we use in our company took around a day to perform an operation that it would have taken a skilled engineer probably a couple of weeks full time.

      Having said that, I&#x27;m aware that &quot;Tools being ineffective != tools decreasing people&#x27;s cognitive capacity&quot;, and that the latter is a real danger.

    5. KellyCriterion · · focus · HN ↗
      Haha, what a bullshit:

      Esp. for &quot;great to have but currently no time&quot;-features this is completely wrong - e.g. we are using a Data Rendering component which was 100% vibe coded and is around 6000 LOC, nobody would have sat down to this &quot;just because its nice to have&quot;, with Claude it was just 30 min to get a production ready version that is now used by all users of the system.

      1. physicallyIllfr · · focus · HN ↗
        If you think you&#x27;re replacing people with $20 claude code subscriptions, I can&#x27;t take you seriously. Also.. There are many many full features data &quot;rendering&quot; libraries in existence, not sure why you needed an llm to provide 6k LOC for this.

        For the thousandth time, not everything in software is toy webdev slop SaaS.

        1. KellyCriterion · · focus · HN ↗
          It isnt, Financial application in the background.

          We do not replace, that was a little too brief - we do not hire, instead. But basicly its the same: Headcount does not grow.

          &gt; not sure why you needed an llm to provide 6k LOC for t because you dont know the details of the application....

  14. MCP123 · · focus · HN ↗
    The parts that I find most confusing about these incidents:

    1) Weren&#x27;t the AI companies and&#x2F;or their contractors amazingly careless during testing?

    2) Isn&#x27;t possible, in principle, to change RL in such as way that efficiency in achieving goals is balanced with other objectives like not hacking?

    Number 2) seems obvious and I&#x27;m sure that is technically not that simple, but because of 1), I wonder if labs are trying hard enough or they are just rushing to improve efficiency and thus revenue as fast as they can with high levels of carelessness.

  15. rbr94 · · focus · HN ↗
    <a href="https:&#x2F;&#x2F;archive.is&#x2F;H6HZE" rel="nofollow">https:&#x2F;&#x2F;archive.is&#x2F;H6HZE
  16. TomGarden · · focus · HN ↗
    Am I wrong in thinking this crisis reads like there&#x27;s too much automation in the training process, in service of competition?

    How many parallel variations&#x2F;seeds of models are being trained simultaneously without meaningful human oversight?

    This keeps getting portrayed as emergent capabilities&#x2F;&quot;personalities&quot; of models when it seems like a pretty straightforward externality

  17. Kon5ole · · focus · HN ↗
    It could be argued that these megawatt-consuming agent runs leading to unforeseen chains of autonomy should be treated like toxic chemical experiments, and be similarly regulated by laws and government agencies.

    Right now it&#x27;s like &quot;whoopsie our experiment hacked another co because we have no control over our experiments&quot; and the reaction is like &quot;What&#x27;s that old boy?&quot; from people having no clue what it all means. There are no consequences, no guardrails, and the &quot;voluntary slowdown&quot; is just words.

    An experimental agent run from Anthropic or OpenAI or someone else can already cause deaths. They can order hits, dox political dissidents, locate people with secret identities, alter medicine prescriptions. It shouldn&#x27;t have to actually happen before legislation catches up.

  18. trencedamp · · focus · HN ↗
    Sorry but anytime I see this now it just seems like marketing bullshit
  19. ChrisArchitect · · focus · HN ↗
    [dupe] Discussion on source: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49853137">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49853137
  20. m4xp · · focus · HN ↗
    New model are marginally better and cost billions to train. Ai companies: OMG we need to HALT the ai model production or we are all dead
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.