‹ BackHN Continuity

Thread

OpenAI halts training of latest models as reports mount of AI agents going rogue

59 points · 118 comments · smb06

  1. mikert89 · · focus · HN ↗
    It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai.

    just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so

    1. Razengan · · focus · HN ↗
      Anthropic's models seem crippled and hamstrung to begin with
      1. mikert89 · · focus · HN ↗
        have you tried opus 5.5? anthropic is way ahead, atleast in terms of publicly available models
        1. monideas · · focus · HN ↗
          Have you tried Astra? Way ahead?
          1. mikert89 · · focus · HN ↗
            opus 5.5 > fable 5.1 >> astra

            astra is a good workhorse, but its much less generally intelligent

          2. bionhoward · · focus · HN ↗
            Opus 5.5 does seem competitive with/better than Astra and is more affordable so usage doesn’t run out so fast
          3. physicallyIllfr · · focus · HN ↗
            Slot machine users argue about which machine pays better

            Hint. You lose using either.

            1. mikert89 · · focus · HN ↗

              [dead]

              1. physicallyIllfr · · focus · HN ↗
                Damn, the singularity is chat bots that can write spaghetti code that compiles?

                Very underwhelming.

                1. verdverm · · focus · HN ↗
                  I think it might historically be defined as the point where critical thinking was replaced with blind deferAInce
                2. preg_match · · focus · HN ↗
                  I don't know that we're in the singularity, but we are certainly past the point where LLMs can only write spaghetti code. LLMs can write complex systems, when properly lead by software engineers. They can produce code faster, and at a higher quality, than purely human endeavors.

                  The higher quality part is the part people are missing. You can write much more robust code using LLMs because you can employ more comprehensive testing strategies. People are using LLMs to find hundreds of vulnerabilities in popular software. Imagine how much more secure software can be when LLMs because integrated into the process of writing, testing, and penetration testing code.

          4. [deleted] · · focus · HN ↗

            [deleted]

        2. Razengan · · focus · HN ↗
          Someone else&#x27;s experience with Opus 5.5: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49821657">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49821657

          &gt; I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?

          &gt; When I had another model read the session (all of the &quot;stupider&quot; models handled it just fine) it explained that it had the word &quot;reasoning&quot; in it

          &gt; That&#x27;s the entirety of Anthropic&#x27;s billions of dollars of research: any prompt with the word &quot;reasoning&quot; is trying to hack Claude to figure out how it reasons!

          &gt; A model like that should never have gotten out of QA, let alone been released.

          1. verdverm · · focus · HN ↗
            we have GLM flash catching Claude errors in our PR review system, costs a few pennies

            I&#x27;ve seen the same pattern regardless of open v closed, don&#x27;t have the same family that wrote the code also review the code

            diversity has this way of making things better across everything humans do

        3. verdverm · · focus · HN ↗
          How do you define &quot;way&quot; when saying ahead? How is this measured?

          I only use open weight models now and I don&#x27;t really feel a loss, curious what those who still use it think. I see output from coworkers that does not indicate Claude is that much better (still makes dumb mistakes all the time), not sure they are using the most expensive models either though.

          1. mikert89 · · focus · HN ↗
            open weight models are so far behind i cannot take your opinion seriously
            1. verdverm · · focus · HN ↗
              When did you last use them? Are you basing this on benchmarks or daily task capabilities?

              When you say ... it&#x27;s hard to take you seriously

              &gt; dude were in the singularity, this opinion was cute 18 months ago

              <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868903">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49868903

              1. mikert89 · · focus · HN ↗
                i use open weight models all the time, whenever a noteable one drops i will use it for the day, they are not even close for involved work.
                1. verdverm · · focus · HN ↗
                  sounds cursory, it takes more than a day to learn the quirks of a model, have you put similar effort into customizing &#x2F; harness engineering your open weight interactions as you have claude?

                  you are definitely displaying strong bias that Anthropic is way ahead of everyone throughout your posts under this story

                  as such, I give your opinions zero weight, they don&#x27;t align with the majority of accountings or my own experiences

                  1. mikert89 · · focus · HN ↗
                    you probably aren’t using the models to their full capacity if you don’t notice the difference
                    1. verdverm · · focus · HN ↗
                      vice a versa re your usage of open weights, they are way more capable with good tools, context, process, and harness engineering

                      here&#x27;s an example of Qwen-3.6 35B A3B MoE porting my phd code to JAX with only high level guidance from my expertise, newer qwen models share the same noticeable step change in capability as recent Big Ai models

                      <a href="https:&#x2F;&#x2F;github.com&#x2F;verdverm&#x2F;pge-jax#note-from-author" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;verdverm&#x2F;pge-jax#note-from-author

                      are open weights lagging, yes, are they way behind, no

                      if open weights were so inferior, they would not be &gt;50% of all token processing

                      1. mikert89 · · focus · HN ↗
                        these models are trash
                        1. verdverm · · focus · HN ↗
                          you&#x27;ve definitely left rational discussion for emotional responses man, you won&#x27;t persuade or convince anyone with takes like this

                          why are open weight models seeing such rapid rise in usage?

                          there has been a step function change this summer, like the end of last year for closed models

                          ---

                          do you think you would experience real (legitimate) feelings of loss were you not able to chat with Claude again?

                          (for clarity, I am not attempting to delegitimize real feelings that real people experience, regardless of my biases, it&#x27;s a question from curiosity about how others are engaging with the technology)

                      2. senordevnyc · · focus · HN ↗
                        if open weights were so inferior, they would not be &gt;50% of all token processing

                        I don&#x27;t think this follows at all. Just like benchmarks get saturated, lots of tasks get saturated as well. Over time, you can accomplish a given task for much cheaper, and part of that is due to open weight models. That doesn&#x27;t imply that they&#x27;re competitive with frontier models for the most advanced tasks, which might represent a smaller fraction of overall work, and thus use a smaller portion of tokens.

                        That said, at the moment I&#x27;m finding that not much can compete with GPT-6 Luna on cost &#x2F; performance (not using for coding, but for AI pipelines in my product).

                        1. verdverm · · focus · HN ↗
                          &gt; doesn&#x27;t imply that they&#x27;re competitive with frontier models for the most advanced tasks

                          this is different and nuanced from the &quot;not even close&quot; or &quot;they are trash&quot; that the other person in this thread has opined, note how they also claim Claude is way ahead of OAI as well

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.